Ideas Made to Matter

Data

Why a semantic layer is pivotal to your AI strategy

Kristin Burnham
4 minute read

What you’ll learn: 

  • A semantic layer creates and maintains a consistent representation of data from different sources. It sits between an organization’s data and the people or machines using it, explaining what the data represents, how different data assets relate to one another, and which rules govern the data’s use. 
  • Semantic layers have become increasingly important as organizations expand their use of AI. AI models and other agents, especially those that use natural language, need to know where data is stored, what business context is attached to it, whether it’s reliable, and how it can be used. 

As companies deploy generative AI tools and agents, many are discovering that giving an artificial intelligence system access to enterprise data is not equivalent to giving it the context needed to understand that data.

When organizations pull information from applications, documents, data platforms, and outside sources, they often strip away the business context that gives it meaning. Different systems may label the same customers, products, or transactions in different ways and apply different definitions and rules. Organizations may also have different standards for data quality, access, permitted uses, and regulatory compliance. 

Without a consistent way to interpret those differences, AI systems can produce answers that are technically plausible but also incomplete, misleading, or wrong, according to a new research briefing from the MIT Center for Information Systems Research. In “The Case for a Semantic Layer,” researchers Hippolyte Lefebvre, Christine Legner, and Cynthia M. Beath describe the elements needed for semantic layers that can make data understandable and usable by both people and machines. 

Understanding the semantic layer

A semantic layer is a system of technologies and techniques that creates and maintains a consistent, unified representation of data from different sources, according to the researchers. It sits between an organization’s data and the people or machines using it, explaining what the data represents, how different pieces relate to one another, and which rules govern its use. 

Semantic layers can draw on technologies such as data dictionaries, taxonomies, knowledge graphs, and ontologies, which define concepts and the relationships among them. Together, they preserve the business context that can disappear when information is removed from the application where it was created. 

That context has become increasingly important as organizations expand their use of generative AI, the researchers write. AI models and agents need to know where data is stored, what business context is attached to it, whether it’s reliable, how it might be combined with other data, how it’s permitted to be used, and which regulatory or privacy requirements apply.

A robust semantic layer captures those details in a form that machines can interpret. It can also capture organizational expertise, including judgments about data quality and exceptions to standard business rules. Companies can then expand their use of AI without rebuilding data definitions and governance controls for every new initiative.

Why businesses need a semantic layer

Many organizations already maintain metadata in data dictionaries, data models, and individual systems. But those resources are often fragmented, incomplete, or dependent on knowledge held by subject matter experts, the researchers write.

The practices required to build a strong semantic layer also remain relatively immature. In a 2024 survey of 349 executives, just 21% rated their organizations’ data curation practices as somewhat or very well developed. Organizations with more developed practices, however, were more than three times as likely to report that they were effective at implementing data and AI initiatives that generated value. They were also twice as likely to say that those initiatives provided a meaningful competitive advantage.

The researchers used the example of Healthcare IQ, a data and analytics company that helps hospitals improve purchasing, pricing, and reimbursement decisions, to illustrate that value. 

In 2022, Healthcare IQ maintained two key data assets: hospital supply chain records and a catalog of nearly 6 million medical products from more than 25,000 manufacturers. The underlying information came from hospital invoices, purchase orders, contracts, electronic health records, reimbursement data, manufacturer websites, and other sources. The same medical product might appear under a description in one hospital system and an internal code in another, making it difficult to compare information across hospitals.

The company built a semantic layer using technologies such as custom data dictionaries, taxonomies, ontologies, data models, and access-control databases. That layer helped it identify equivalent products, standardize descriptions, flag data quality problems, and apply privacy and regulatory controls. It also automated 80% of the work required to onboard new hospital customers.

Healthcare IQ’s experience shows how a semantic layer can turn fragmented data into standardized, reusable data assets. A semantic layer can also help AI models and agents interpret enterprise data with context, leading to more accurate and less risky outputs.  

3 keys to building a semantic layer

Creating a comprehensive semantic layer is a long-term effort. The researchers recommend that organizations take three actions to ensure that their investments generate value.

  • Start with priority data assets. Rather than attempting to describe and organize every piece of enterprise data at once, leaders should identify the data needed for their highest-priority AI initiatives. They can then invest in tools and techniques that address the contextual gaps holding those initiatives back. 
  • Govern the semantic layer itself. Because the semantic layer contains important information about the meaning and acceptable use of data, it needs a clearly designated owner. That person should be accountable for its quality and results, work closely with enterprise data platform owners, and participate in decisions about when data and definitions should be updated or retired. 
  • Use AI to help create and maintain metadata. Organizations should also look for opportunities to use AI to manage the metadata that describes their data assets. AI can support tasks such as generating metadata, cleaning and classifying data, recommending access controls, and identifying connections among data assets. AI will become increasingly important for these tasks as organizations manage growing numbers of AI tools and volumes of unstructured content. 

As generative AI tools and agents proliferate, competitive advantage will increasingly depend on how well organizations make their proprietary data accessible and understandable to people and machines, the researchers write. Organizations that do this well can scale AI faster, at lower cost, and with greater confidence that their data is being used in appropriate ways.



Hippolyte Lefebvre is a research collaborator with the MIT Center for Information Systems Research, an assistant professor in management information systems at University College Dublin, and an affiliated researcher with the Competence Center Corporate Data Quality, an industry-funded research consortium and expert community. His research focuses on improving data and AI management in multinational firms.

Barbara H. Wixom is a principal research scientist at MIT CISR. Since 1994, her research has explored how organizations generate business value from data assets. Her methods include large-scale surveys, meta-analyses, lab experiments, and in-depth case studies. She teaches the MIT Sloan Executive Education course Data Monetization Strategy: Creating Value Through Data.

Christine Legner is a research collaborator with MIT CISR and a professor of information systems at the University of Lausanne. She is also the academic director of the Competence Centers Corporate Data Quality. She develops concepts, tools, and methods for data management. 

Nick van der Meulen is a research scientist at MIT CISR. He conducts academic research that targets the challenges of senior-level executives, with a specific interest in how companies need to organize themselves differently in the face of continuous technological change. He is one of the faculty members who teaches the MIT Sloan Executive Education course Global Executive Academy

Cynthia M. Beath is an academic research fellow at MIT CISR and professor emerita at the University of Texas at Austin. Her research interests include organization redesign for the digital era, the management of data assets, and the organizational impacts of AI.

For more info Sara Brown Senior News Editor and Writer