Skip to main content
Bird & Bear
Menu
← All articles

BigQuery pipelines can now document themselves

BigQuery pipelines can now document themselves

Two useful BigQuery updates landed in mid-August. The first being a routine model bump, whilst the second is a great example of what “AI-ready data” looks like.

What happened

New Gemini models for BigQuery’s built-in AI functions (10 August 2026): gemini-3.1-flash-lite and gemini-3.5-flash are now generally available for BigQuery’s generative AI functions. As a reminder, these are the AI.GENERATE, AI.GENERATE_TABLE and related SQL functions that let you call a Gemini model directly against a column of data without the need for a separate pipeline. The rollout covers the us, eu, and global multi-regional endpoints. As stated in Google’s own model documentation, 3.5 Flash-Lite is positioned as a low-latency, cost-effective model for high-volume, simple extraction work, such as classifying support tickets or extracting survey results from free-text.

The Data Engineering Agent can now generate its own pipeline metadata and sync it straight to Knowledge Catalog (13 August 2026): BigQuery’s Data Engineering Agent, the natural-language interface for building and modifying Dataform pipelines, can now define semantic metadata directly inside the SQLX pipeline code, whether it’s written by a person or generated by the agent itself from a plain-English prompt. Once that pipeline runs successfully, the metadata syncs automatically into Knowledge Catalog, Google Cloud’s governance and discovery layer (the artist formerly known as Dataplex Universal Catalog). The upside of this is that it means you do not need separate documentation that lives elsewhere and runs the risk of becoming outdated.

What is Knowledge Catalog and why does it matter?

It’s worth being specific here, because “metadata syncs to a catalog” undersells what’s actually on the other end of that sync, and I do think Knowledge Catalog is a highly underused product.

Knowledge Catalog is Google Cloud’s central metadata and governance layer. In Google’s own words, it functions as “a central data governance and agentic access layer” that automatically discovers and indexes BigQuery assets (datasets, tables, views, models, routines), then lets you enrich them with business context and search across all of it, either as an end user, or by an AI agent. A few of the things it does once your pipeline’s metadata is flowing into it:

  • Turns tables into searchable, documented assets: Every BigQuery object becomes a cataloged “entry” that can carry ownership info, plain-language descriptions, and business glossary terms, attached to the actual table rather than living externally.
  • Column-level tagging, including PII markers: Metadata isn’t limited to the table level. Individual columns can carry data-quality scores and sensitivity flags (e.g. “this column contains PII”), which is what makes automated governance and access control possible rather than something a human has to remember to apply.
  • Data quality scorecards: Knowledge Catalog can pull in Dataform assertion results, so a table’s catalog entry shows not just what the data is but whether it’s currently passing its own quality checks.
  • Lineage tracking: It automatically traces how data flows and transforms into and out of BigQuery tables. This is useful for the boring-but-critical question of “if this column is wrong, what else downstream is wrong,” and for auditing where sensitive data actually ends up.
  • Natural-language, permission-aware search: Search is semantic rather than keyword-only, and it respects IAM and VPC Service Controls. This means that if a person (or an agent) is searching for “customer churn data”, it only surfaces assets they’re actually allowed to see, and gets pointed at the right table without needing to know its literal name.

That last point is the one that matters most for the “AI-ready data” argument specifically. Google’s own framing (from the Knowledge Catalog announcement) is blunt about why this exists: AI agents given direct access to a warehouse without this layer tend to hallucinate joins, misread ambiguous column names, and generate confidently wrong SQL. A table schema alone doesn’t tell an agent what a column means, only what type it is. Knowledge Catalog is explicitly built to close that gap, giving an agent (or a person) verified context in the form of descriptions, relationships, even pre-verified query patterns.

What you can actually do with this

Instead of the standard flow of a data engineer writing a pipeline, or analyst implementing new tracking and then, if there’s time, going back to document it in a separate tool, the documentation gets defined as part of the pipeline itself. It can also be generated by an agent when it builds the pipeline in the first place. This means it’s immediately live in a governed, searchable catalog the moment the pipeline succeeds rather than being a task to come back to or bake into processes. A few things that unlocks in practice:

  • A business user or analyst can search Knowledge Catalog in plain language (“which table has verified monthly active users”) and land on the right, documented table instead of guessing between five similarly-named ones or pinging the data team with questions.
  • An AI agent, whether that’s Gemini, a custom agent, or increasingly a client’s own internal tooling, can be pointed at Knowledge Catalog to ground itself before writing a query, rather than working blind off a raw schema and quietly producing plausible-looking wrong answers.
  • PII and sensitivity tagging happens at the column level as part of the pipeline definition, not as a retrofit exercise someone runs once a year (or never). This is is the difference between governance that’s actually enforced and governance that’s aspirational.
  • Data quality becomes visible at the point of discovery: someone finding a table via search can see whether it’s currently passing its own assertions, not just that it exists.

What’s new here is that generating and maintaining documentation is shifting from a manual task to something the pipeline (or the agent building the pipeline) does as a byproduct of doing its actual job.

What to actually do about it

Neither update needs urgent action now. The Gemini model bump is a drop-in upgrade for anything already using BigQuery’s AI functions, and the metadata/Knowledge Catalog feature is still in Preview at the time of writing. But it’s worth two things: if you’re already running Dataform pipelines in BigQuery, it’s worth checking whether your team has Knowledge Catalog switched on and whether pipeline metadata is actually being defined (the Knowledge Catalog FAQ and the Data Engineering Agent pipelines guide are the two places to start). Also, if a client or brand is evaluating whether their data is genuinely ready for AI tooling, “does an agent have grounded context to query against, or is it working off raw table names” is now a specific, answerable question, not a vague one. This is a good concrete example to use when that conversation comes up and helps close the gap between “AI-ready” as a claim and “AI-ready” as an actual, checkable state.

Newsletter

Occasional notes on data, tracking & AI-readiness

No spam, no growth-hacking. Just an email when there's something worth reading.