6 min read
If you’re a data person building with AI, you’ve probably been questioning the role of the semantic layer. Do you build a thorough semantic layer to give context? Or do you allow the model to build context for itself? My position has been that a semantic layer is a prerequisite for any AI analytics work. But with recent testing, that view has softened.
Before I explain why, I want to take a step back and define for you what a semantic layer is in today’s AI driven environment, how that definition has changed, and what it looks like in practice.
A semantic layer is where business meaning lives in a form machines can read.
It translates raw tables into business language, defining “revenue,” “active user,” or “customer” once so every tool, person, and AI model pulls the same number.
The idea is thirty years old, but how it gets built has changed.
Semantic layers used to require months of upfront modeling by a specialized team. Today, definitions can be drafted from existing queries, written in plain language, and revised in hours.
AI raised the cost of skipping it, and lowered the cost of building it.
A model that writes SQL still has to guess what “revenue” means if it’s not already defined, but correcting that definition is quick, and documentation can be generated.
It defines metrics, dimensions, and business logic once, instead of reimplementing them in every dashboard, spreadsheet, and query.
In practice, a semantic layer holds three things:
A metrics layer is a narrower version that holds only metric definitions.
What used to be a large upfront project is increasingly an ongoing, iterative practice.
The traditional build was deliberate and slow. A central data team interviewed stakeholders, modeled metrics before anyone asked a question, and controlled every change. Critics argued this created bottlenecks: the business moved faster than the layer could be updated (Ben Bausili), and the process created new gatekeepers instead of widening access (MotherDuck).
Three shifts have changed that:
The stakes are higher, even as the work gets lighter.
In a Futurum Group survey of 818 enterprise decision makers (1H 2026, companies with $100M+ revenue), about 59% plan to increase or newly adopt semantic layer spend over the next 24 months, with accuracy and hallucination risk as the top concern (Futurum Group). MIT CISR found organizations with well developed data curation practices are 3x more likely to succeed with data and AI initiatives, yet only 21% rate their own practices that way (MIT CISR). In one benchmark, a curated semantic layer lifted AI accuracy from 93% to 100% (MotherDuck).
Below is how the traditional approach compares to what a semantic layer looks like today:
| Traditional semantic layer | Semantic layer in an AI driven environment | |
|---|---|---|
| How it’s built | Modeled upfront, before questions are asked | Drafted from existing queries and usage, then reviewed |
| Who maintains it | A central data or IT team | The data team plus domain owners, often with AI assistance |
| Primary reader | People and BI tools | People, BI tools, and AI models |
| Time to update | Typically weeks to months | Often hours to days |
| Format | Proprietary modeling language | Code or plain language descriptions, version controlled |
| Scope | One universal model for the whole org | A shared core, with room for team specific views |
Most definitions are small, specific, and surprisingly contested.
Each row below is a common business term, a question an AI assistant might receive, and what happens with and without a governed definition.
| Business term | Question asked | Without a definition | With definition |
|---|---|---|---|
| Revenue | “What was revenue last quarter?” | The model picks a table and may include refunds, tax, or test accounts | Paid invoices, net of refunds, excluding internal accounts |
| Active user | “How many active users do we have?” | The model may count logins, any event, or every account | A user who ran at least one query in the last 30 days |
| Customer | “How many customers churned?” | The model may count individual users or free trials | An organization on a paid plan; churn means the plan ended without renewal within 30 days |
| Region | “Show revenue by region” | One query uses billing address, another uses shipping | Billing country, mapped to four sales regions |
In the traditional model, each of these took a meeting to settle and a sprint to implement. Today, a model can propose a definition based on how queries already use these fields, and a person approves or edits it. If you’re using AI to build dashboards, I recommend asking for an output of the business logic if you didn’t already provide a strict metric definition.
On the plus side, if you’re someone starting a new team and need insights fast, Metabase can build your semantic layer for you upfront and load it all into Data Studio for you to reference. Changing a definition is as easy as asking Metabase to do it.
Defining the semantic layer is the easy part. Deciding how much of it to build is not.
Every implementation on the market is a set of answers to these questions:
If you’re building out a semantic layer for your team, what’s your approach? Are you modeling metrics deliberately across a pieced together data stack, or simplifying with an all in one tool? Is a central semantic layer still crucial, or are narrow, model specific definitions a better bet for accurate output?
If you’re weighing these tradeoffs for your own stack, explore what it looks like to bring semantic modeling and analytics into one place.