Skip to main content

Easy ways to prepare your BigQuery warehouse for AI

Katie Kaczmarek•23 April 2025•4 min read
Easy ways to prepare your BigQuery warehouse for AI

You’ve probably heard that AI is coming to make our lives easier, especially in tools like BigQuery. But here’s the thing: AI isn't magic. If you want it to be accurate and useful, you need to set it up for success.

One of the best ways to do that? Improve the metadata in your BigQuery warehouse.

Metadata is like the index or contents page in a book, it quickly tells you exactly what’s inside and where to find it. Creating clear metadata means AI can more easily understand your data warehouse and give you more accurate, relevant results.

Here’s how you can easily start improving your warehouse metadata today:

  1. Clearly describe your tables and fields (what they contain and why they're useful).
  2. Link related tables together (like connecting dots) so it's obvious how they're related.
  3. Tag your tables with keywords to help you and AI find them quickly.
  4. Keep everything neat and consistent, using the same approach everywhere.

1. Add descriptions to your tables and fields

Currently you either have to do this manually (yes, typing away) or write a short script that loops through your tables and sets descriptions programmatically. There's good news however: BigQuery's new Data Insights feature should be able to automatically generate table and column descriptions, simplifying the documentation process. One thing worth knowing: generating a description doesn't put it on the table by itself. On the Insights tab, underneath the generated table description, click "View column descriptions". That opens a panel with a "Save to details" button for the table and a "Save to schema" button for each column, and clicking those is what actually writes the description onto the table, where it shows up everywhere, including to any AI reading it through a normal BigQuery connection.

Why bother?

Imagine AI trying to figure out what "cust_id" or "rev_tot" stands for. Field and table descriptions are like name tags at a party; the clearer they are, the smoother the introductions go. Good descriptions mean AI spends less time guessing and more time providing accurate insights.

Quick wins (A term I hate in SQL requests):

  • Add simple, clear field and table descriptions
  • Include example values where it helps
  • Use Data Insights to scale it across your warehouse
  • Link related tables clearly

Think of your datasets like a family tree, everything is connected. Clearly defining relationships between your tables makes it easier for AI (and your team!) to understand how everything fits together.

Simple steps:

Need help with your data platform?

We build intelligence platforms on BigQuery, Dataform and Google Cloud - from setup to ongoing optimisation.

  • Clearly document relationships in your table descriptions.
  • Identify joining keys in your field descriptions

3. Tag your tables for quick discovery

Tags are your shortcut for finding data fast, think of them like labels on folders or bookmarks in your browser. They’re great for quick filtering and navigation, helping you (and AI) pinpoint exactly what you need without wasting time.

My advice?

  • Think practically. Use BigQuery's built-in labels (simple key-value tags you can set on any dataset or table) to reflect actual usage, like finance, marketing or PII. They're worth using for a second reason too: labels are one of the things that travel through to an AI reading your data via a BigQuery connection, so they're doing double duty.
  • Be strategic and consistent in tagging, making it easy to sort and access your tables quickly.

4. Use clear, consistent naming conventions

Good naming conventions aren’t just nice, they're essential. Clear names prevent future confusion, save time, and make your warehouse easier to navigate for everyone, especially AI.

Imagine naming your files "final_final_3.csv" in your Google Drive. Finding anything later would be a nightmare! Clear, consistent names like "sales_data_jan2025" help AI understand exactly what’s in each table.

Bonus Tip: Create report-specific tables

If you want to go one step further, consider building core reporting tables designed to make analysis easier.

Think of these as simplified, centralised “AI-friendly” versions of your data. They can also act as your single source of truth for dashboards and recurring reports.

Examples of useful AI-ready tables:

  • core_metrics_summary: daily or weekly KPI snapshots
  • user_engagement_core: simplified GA4-style user data
  • product_performance: clean sales data by product
  • customer_lifetime_value: key user value metrics
  • data_dictionary_ai: a table AI can refer to for definitions and aliases

These tables make it easier for AI tools to produce reliable outputs and for people to report from the same base data.

Quick recap (because we love simplicity):

  • Add clear field and table descriptions using BigQuery's Data Insights.
  • Clearly link related tables together.
  • Use strategic tags for easy navigation.
  • Implement consistent and understandable naming conventions.
  • Consider creating core reporting tables to support AI and your team.

Small efforts now will make AI a more helpful, accurate assistant for everyone in your team, coders and non-coders alike.

What are you doing today to get ready for AI in BigQuery?


We've since gone deeper on this, specifically on Google's Knowledge Catalog and what's changed around AI and BigQuery since this was written: What's actually in Google's Knowledge Catalog & How to make an AI agent accurate in BigQuery

Need help with your data platform?

We build intelligence platforms on BigQuery, Dataform and Google Cloud - from setup to ongoing optimisation.

How ready is your data?

Take our short assessment to find out where your data stack stands and what to prioritise next.


Suggested content

How to make an AI agent accurate in BigQuery

A year ago, I wrote about getting a BigQuery warehouse ready for AI ( Easy ways to prepare your BigQuery warehouse for AI) , but didn't look properly into Knowledge Catalog (known as Dataplex at the time). We decided to go back and look through Google's Knowledge Catalog properly: what's actually in there, how it works and what it relates to. What Knowledge Catalog is Knowledge Catalog is Google's metadata layer for BigQuery data (and a few other sources). It sits alongside your tables rather

Katie Kaczmarek•15 Sept 2026

What's actually in Google's Knowledge Catalog

Google's Knowledge Catalog has been renamed four times. Data Catalog, then Dataplex Catalog, then BigQuery universal catalog, then Dataplex Universal Catalog, and now Knowledge Catalog, as of 10 April 2026. The API, gcloud and IAM roles still all say "dataplex". If you land on a page that mentions Dataplex and wonder whether you're reading something out of date, you're probably not. That's just the product's fifth name in four years. I spent some time looking through the different sections of K

Katie Kaczmarek•15 Sept 2026

BigQuery Tips: When your query is technically correct but BigQuery won't run it

There is a particular kind of frustration that comes from staring at a query you know is correct and watching it fail. No syntax error. No logic problem. Just a wall. We hit two of them on the same project. What we were building The job was to migrate ga4_daily_snapshot for a large enterprise client from a BigQuery scheduled query into a proper Dataform pipeline. The scheduled query had been added to over time until it was too large to maintain with any confidence. Moving it to Dataform w

Katie Kaczmarek•17 Aug 2026