*We customize the course outline and content to your specific needs and relevant use cases.
Module 1: Scalable dbt modeling and project structure
- Applying a consistent staging, intermediate, and marts layer model
- Establishing clear naming conventions and model responsibilities across layers
- Using ref() and source() as dependency definitions rather than simple SQL substitutions
- Designing model boundaries that support readability, reuse, lineage, and maintainability
Module 2: From dbt model to executed warehouse SQL
- Understanding parsing, Jinja rendering, DAG construction, compilation, and SQL execution
- Using dbt compile to inspect generated SQL and separate Jinja problems from SQL problems
- Comparing target/compiled with target/run and understanding what each output represents
- Understanding how view, table, incremental, and ephemeral materializations transform the same SELECT logic into different executable SQL
Module 3: Incremental models and loading strategies
- Understanding is_incremental(), first runs, subsequent runs, and why full refreshes follow different execution paths
- Using append for immutable event or log data and understanding duplicate risks
- Using merge with unique_key for inserts and updates, including merge_update_columns and merge_exclude_columns
- Comparing delete+insert and insert overwrite approaches for key ranges or complete partitions
Module 4: Advanced incremental operation and performance
- Understanding microbatch, introduced with dbt 1.9, for time based processing of very large histories and controlled backfills
- Choosing strategies based on late changes, business keys, table size, platform support, and full refresh cost
- Designing incremental predicates and lookback windows for late arriving data instead of relying only on max(updated_at)
- Managing schema changes, full refreshes, controlled backfills, idempotency, partition_by, and cluster_by for reliable and cost aware operation
Module 5: Advanced testing and test scoping
- Combining and configuring generic tests such as relationships and accepted_values
- Writing singular SQL tests for business rules that do not fit standard generic tests
- Applying where configurations to individual tests to focus validation on relevant subsets
- Designing test coverage for new data, historical exceptions, and important business rules
Module 6: Model contracts and structured documentation
- Defining explicit column names, data types, and supported constraints with enforced model contracts
- Understanding where contracts strengthen interfaces between dbt models and downstream users
- Maintaining model and column descriptions systematically in YAML
- Generating and using dbt documentation and lineage views as an active development and review resource
Module 7: Snapshots, seeds, and existing packages
- Using snapshots to capture historical changes with a Slowly Changing Dimension Type 2 approach
- Configuring snapshots and understanding when historical change tracking is appropriate
- Using seeds as version controlled reference and configuration tables, including mappings and control data
- Integrating existing packages such as dbt_utils for date spines and surrogate key generation
Module 8: Jinja, macros, and additional productivity packages
- Introducing Jinja through variables, {% set %}, expressions, and simple control structures
- Writing small reusable macros without moving immediately into complex metaprogramming
- Using dbt-audit-helper to compare results between model versions or implementations
- Using codegen to accelerate repetitive source and model YAML generation
Module 9: Team development, environments, and CI/CD
- Using branches, pull requests, and code reviews for everyday dbt development
- Separating development and production environments clearly
- Working with profiles.yml and environment variables for environment specific configuration
- Connecting environment management and automated validation to practical CI/CD workflows
Module 10: Model selection and graph operators
- Using --select and --exclude to control which project resources are executed
- Applying graph operators such as +model and model+ to include upstream or downstream dependencies
- Selecting resources by tags, paths, and combinations of selection criteria
- Building efficient development, testing, and deployment selections instead of rebuilding the entire project
Module 11: dbt artifacts, lineage, and Slim CI concepts
- Reading manifest.json to understand nodes, dependencies, metadata, tests, and project structure
- Using run_results.json to inspect execution status, timing, and node results
- Understanding how artifacts support lineage analysis, project reporting, and automated tooling
- Connecting state, selection, and project artifacts to targeted CI approaches such as Slim CI
Module 12: Exposures and the dbt Semantic Layer
- Defining exposures in YAML to represent dashboards, reports, and other downstream consumers in the DAG
- Using exposures to improve ownership, impact analysis, and visibility beyond transformation models
- Introducing the dbt Semantic Layer and MetricFlow as a way to define governed metrics above analytical models
- Understanding when centralized metric definitions complement marts rather than embedding every metric directly in final tables