Blog

Guardrails for AI-written dbt projects with dbt-bouncer

Pádraic Slattery

Pádraic Slattery

September 28, 2026
9 minutes
‌

dbt 2.0 was released last week. Most of the attention went to the engine: after many months of development, dbt 2.0 ships Fusion as generally available (GA) and the coverage has been about how much faster it parses and compiles.

Over the same period, a growing share of the dbt models arriving in pull requests all over the world have not been written by hand. They are drafted by a coding agent and then adjusted. At Xebia we now see this across a range of client projects, and the pattern is consistent: the cost of producing a dbt model has reduced, while the cost of reviewing one has not moved, and may even have gone up. Authoring used to be the bottleneck; now it is trust.

What agents actually get wrong

Agents rarely get the SQL wrong. Give an agent a source table and a description of what you want, and you will get back a query that runs (eventually) and returns rows.

What it cannot give you is your team's conventions, because those conventions were never written anywhere it can reach. They came out of a discussion eighteen months ago and they live in the heads of the three people who were in the room, one of whom no longer works for the company:

  • Staging models are named stg_[source]__[entity]s.
  • Boolean columns start with is_, except for the two in the finance mart that predate the rule and would break a Power BI dashboard if updated.
  • Timestamps carry their timezone in the name, so updated_at_utc and never updated_at.
  • Anything in marts must have an owner in its meta block, because that is who gets messaged when a pipeline fails.

An agent faced with a gap like this does not stop and ask. It picks something plausible and moves on. Once, that is a nitpick in review. Across forty pull requests it is drift, and drift is expensive to unwind later.

There is a second-order problem that is more worrying. Back when colleagues hand-wrote models, the reviewer read the model properly. When an agent wrote it and the diff looks tidy, the reviewer skims the PR, sometimes only looking at the AI review comment and never at the actual code. The code that gets the least human attention is now the code that had the least human involvement in the first place.

The fix is to write the conventions down in a form that the agent can access.

Why not just put conventions in CLAUDE.md?

I am a fan of CLAUDE.md; I even wrote a blog about structuring a dbt project around it. Write your conventions there and an agent will follow them most of the time.

Most of the time is the problem. A CLAUDE.md is an instruction to a model, and the model weighs it against everything else in its context. As the file grows, and as the context fills with the actual task, individual instructions get less attention. Nothing tells you when one is skipped or over-ruled.

A dbt-bouncer check is an assertion. It runs after the code exists, it returns the same verdict every time, and when it fails the build goes red. It also covers code the agent did not write: a colleague in a hurry, a different agent, a contributor whose editor never loaded your CLAUDE.md at all.

So as a rule of thumb, put the conventions in CLAUDE.md so the agent gets them right most of the time. Put them in dbt-bouncer so nothing that breaks them gets merged.

Step one: make the conventions executable

This is the problem dbt-bouncer was built for, and I wrote about the original version here. It runs checks against dbt's artifacts, so it needs no database connection and works with any adapter. There are 127 (and counting) checks covering models, sources, macros, exposures, seeds, snapshots, tests and run results.

In v4 (released this week!) you no longer need a config file to find out whether any of this is useful to you. Three presets ship with the package:

pip install dbt-bouncer
dbt-bouncer run --preset standard

Running that against our test project gives:

Running dbt-bouncer (4.0.0)...
Using the `standard` preset configuration.
Validating conf...
Assembled 91 checks, running...
Running checks... ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ 100%
`dbt-bouncer` failed. Please see below for more details.
Failed checks
╭──────────────────────────────────────────────────────┬──────────╮
│ Check name                                           │ Severity │
├──────────────────────────────────────────────────────┼──────────┤
│ check_source_not_orphaned:12:sources_that_dont_real… │ ERROR    │
╰──────────────────────────────────────────────────────┴──────────╯
Done. SUCCESS=90 WARN=0 ERROR=1

minimal is a starting point for a project with no conventions at all, standard is what we would put on most projects, and strict is the full set. When you outgrow a preset, dbt-bouncer init will write you a config file to edit.

Once you have a config you are happy with, it belongs in two places.

  1. The first is pre-commit, so it fires before the code leaves a developer's machine:
# .pre-commit-config.yaml
repos:
  - repo: https://github.com/godatadriven/dbt-bouncer
    rev: v4.0.0
    hooks:
      - id: dbt-bouncer
        args: ["--config-file", "dbt-bouncer.yml"]
  1. The second is CI, so a developer who skips the hook (or does not have the hook set up) does not skip the check.

Step two: make the conventions readable by the agent

A failing pre-commit or CI tells you that an agent broke a convention. It tells the agent nothing: it writes a model, the hook or CI fails and a developer feeds the failure back (or you ask the agent to parse the failure logs). That loop works, but it burns a cycle every time, and the agent is no wiser as to the why behind the failing check.

dbt-bouncer v4 closes the loop by letting the agent run the checks itself via an MCP server. This is an optional feature and is of most use when adding dbt-bouncer to a pre-existing dbt project.

pip install 'dbt-bouncer[mcp]'
{
  "mcpServers": {
    "dbt-bouncer": {
      "command": "dbt-bouncer",
      "args": ["mcp"]
    }
  }
}

That registers an MCP server exposing four tools: list_checks, explain_check, read_project_config and run_checks. Before it writes anything, the agent can read the project's config and see which conventions apply. After it writes, it can run the checks and fix its own output. The pull request you eventually read is one that already passes.

Every check in v4 has a stable rule code, and you can ask about one from the command line as well as via the explain_check tool:

dbt-bouncer explain MO020
╭─ check_model_description_contains_regexp_pattern (MO020) ──────────╮
│ Models must have a description that matches the provided pattern.  │
│                                                                    │
│ !!! info "Rationale"                                               │
│     A free-text description field is easy to fill with             │
│     placeholder or low-quality content. Requiring descriptions to  │
│     match a pattern ensures that documentation meets a baseline    │
│     standard of usefulness rather than just being non-empty.       │
╰────────────────────────────────────────────────────────── manifest ─╯

The rationale is the part that changes agent behaviour. Told only that a check failed, an agent will do the smallest thing that makes the message go away, often the wrong thing. Told why the rule exists, it tends to fix the underlying problem.

For Claude Code users, the repository doubles as a plugin and ships a skill that reads an existing dbt project and proposes a config built from the conventions already present in it. It is a pragmatic way to start: adopt what you are already doing, then tighten.

Step three: make the exit code worth believing

The first two steps put the agent in a loop where it runs the checks itself. An agent reads the exit code, records the step as done and moves on to the next task.

In v4 exit codes tell you which kind of failure you have:

A pipeline that treats any non-zero code as failure keeps working unchanged. A pipeline that treats 1 as the only failure needs updating, because a misconfigured run can now be told apart from one that found problems.

Turning this on where you already have violations

If you point --preset strict at a mature dbt project, you will get hundreds, possibly thousands, of failures. Nobody schedules a cleanup sprint for that, so the usual outcome is that the tool gets switched off.

v4 adds a baseline for this:

dbt-bouncer baseline --config-file dbt-bouncer.yml
git add .dbt-bouncer-baseline.json
dbt-bouncer run --baseline .dbt-bouncer-baseline.json

The first command records today's failures. After that, only failures absent from the baseline fail the build. Existing debt is visible but not blocking. There is also a --state flag, which compares against a directory of artifacts from a previous run instead of a file.

This suits agent-written code: the model your agent wrote this morning is held to the full standard. The model somebody wrote in 2022 is not, until someone chooses to touch it.

Running this on dbt 2.0

dbt-bouncer v4 runs against dbt 2.0 with no config change, because dbt 2.0 emits the same artifact schemas as dbt 1.x. Every check that worked before still works.

What did change is how you generate catalog.json:

  • pip install dbt gives you the Fusion CLI, which writes catalog.json. pip install dbt-oss gives you the Apache-2.0 build, which writes its catalog as Parquet instead. If you run catalog checks, use dbt.
  • --write-catalog is no longer available on dbt build. Run dbt compile --write-catalog first, then dbt build, in that order.

If a catalog is missing while catalog checks are configured, v4 tells you so and exits 3 instead of skipping them.

Where this leaves you

Conventions in a dbt project have always lived in documentation that nothing enforced. Agents have made that harder to ignore, because an agent cannot read your Confluence page and will not ask you what you meant.

So write the conventions down as checks, give the agent access to them, and make sure the exit code tells it what actually went wrong.

dbt-bouncer v4 is out now. The migration guide covers the breaking changes, and the repository is the place to start.


Are you part of an organisation looking into implementing best practices around dbt? Our analytics engineer consultants are here to help – just contact us and we'll get back to you soon. Or are you an analyst, analytics engineer or data engineer interested in learning more about dbt? Check out our dbt Learn course at Xebia Academy or have a look at our job openings.

Written by

Pádraic Slattery

Pádraic is a technical-minded engineer passionate about helping organizations derive business value from data. With experience in data engineering, Business Intelligence development, and data analysis, he specializes in data ingestion pipelines and DataOps.

Contact

Let’s discuss how we can support your journey.

‌
‌
‌
‌
‌
‌
‌
‌
‌