A practical framework for building data systems that AI can find, understand, trust, and use.

Conversations around enterprise AI are moving rapidly from models to applications. Foundational models are becoming increasingly capable, and organisations are experimenting with retrieval augmented generation (RAG), tool calling, AI agents and multi-agent systems to bring intelligence into their business processes. Yet, as these systems move closer to production, a familiar constraint emerges: the data they need was rarely designed for AI consumption. 

Enterprise data has traditionally been organized around applications, reporting, transactions and human decision-making. AI introduces a different set of requirements on data. It needs to find relevant information, understand meaning, establish connections across data sources, assess whether it can trust the information at hand, and consume information in a form suitable for reasoning and action. This leads to a fundamental question, what does it actually mean for enterprise data to be AI-ready? 

A layered view of AI-ready data

AI readiness is not a binary state. Connecting a large language model (LLM) to an enterprise database, or adding a vector database to an existing data platform, does not make the underlying data AI-ready. Readiness develops through a set of capabilities, each addressing the requirements for AI systems. AI-ready enterprise data can be understood as a layered set of capabilities, as shown in the figure.

The layered framework: each layer makes enterprise data more usable by AI.

The foundation is an enterprise data organization layer, where structured and unstructured data are maintained at appropriate granularity, with consistency, quality, temporal fidelity and governance. This is followed by machine-consumable metadata and semantic context layer, that allows AI systems to understand what data represents, where it comes from and whether it is appropriate for a given task. The unstructured enterprise data including documents, reports, policies, etc. are contained in the retrieval-ready knowledge layer so that this content can be found by AI using their meaning rather than by exact filenames or identifiers.

Data governance and control layer provide the policies, lineage, access controls, consent, and compliance context that allow AI systems to use information responsibly. Enterprise knowledge integration layer connects shared entities and relationships across structured data, unstructured content, metadata, and governance. Finally, insights and signals layer makes frequently used business logic, metrics, features, and predictions directly reusable by AI applications.

Together, these layers make data progressively more machine-processable, machine-understandable, discoverable, trustworthy, connected, and ready for AI consumption.

Four systems that operationalize the layered framework

The layered framework becomes actionable through four complementary systems: the System of Record, the System of Understanding, the System of Discovery, and the System of Action. A unified discovery plane provides an interface for AI systems to consume capabilities exposed by these four systems.

The four systems: governed data, understanding, discovery, and action work together through a unified discovery plane.

System of Record

The System of Record provides the governed foundation for structured and unstructured enterprise data. A lakehouse can support progressive refinement, schema evolution, transactional guarantees, versioning, and time travel for structured data. Additionally, governed 'object and document storage' provide equivalent foundations for unstructured content. This system can also host feature stores, metric stores, and derived signals. That lets an organization compute valuable business insights once and reuse it across AI applications. A time-tagged training dataset, a governed customer risk score, or a derived business metric becomes a reusable enterprise asset instead of something that every application has to recompute each time.

System of Understanding

Data is useful to AI only when its meaning is associated with it. The System of Understanding provides machine-consumable descriptions of schemas, business definitions, ownership, quality, freshness, lineage, and policies. It may include a metadata catalog and an enterprise knowledge graph that captures the entities and relationships connecting information across domains. For example, an enterprise knowledge graph can connect a customer record in a database, a mention in a contract document, related transactions in a lakehouse, and current state available through an API. Organizations can start with semantic linking and entity resolution, then evolve toward richer graph representations. GraphRAG can complement semantic retrieval when an AI system must reason across explicit relationships.

System of Discovery

Traditional discovery assumes that the consumer already knows which dataset, schema, or column contains the answer. AI systems cannot make that assumption. The System of Discovery provides an AI-facing retrieval surface over enterprise knowledge. It uses vector indexes and semantic retrieval to find relevant documents, metadata, reports, knowledge-graph context, and other information based on their meanings.

RAG is especially effective when knowledge is distributed across a large collection of unstructured content. Graph traversal provides a complementary path when the question depends on explicit relationships. The discovery layer is not the authoritative source of enterprise data; it is a retrieval and access surface over governed assets maintained by the Systems of Record and Understanding.

At these three systems, data quality becomes actionable for AI consumption. Experienced teams often know, through tribal knowledge, which data sources are complete, current, or reliable enough for a particular decision. AI systems cannot infer that reliably from a table name or a retrieved document. Quality metrics, validation status, freshness, lineage, and ownership must be available as governed metadata or retrieval-ready information, so that the AI can assess whether a source is fit for the task. Otherwise, weak data is likely to produce weak responses.

System of Action

The System of Action is where AI systems turn the capabilities from data into decisions and actions. An AI agent should be able to query structured data, retrieve documents, traverse enterprise knowledge graphs, access reusable signals, and invoke real-time APIs without needing to understand the implementation details of every backend system.

The unified discovery plane can bring these capabilities together through a consistent, AI-oriented interface. It should also carry the context needed for responsible use: access controls, policies, quality signals, and provenance should travel with the data and tools that an agent uses.

Data for AI and AI for Data

The architecture has a second direction. AI does not only consume AI-ready data; it can help create and maintain it. This is the AI for Data dimension, which complements the Data for AI foundation.

AI can help classify sensitive information, enrich metadata, process documents, extract and resolve entities, evolve knowledge graphs, perform advanced data-quality checks, and generate reusable predictive signals. Feedback from agent usage, failures, and quality signals can then inform upstream improvements to data preparation and curation.

That does not mean every data-preparation task should invoke an LLM. Smaller domain-specific models, traditional machine learning, deterministic rules, asynchronous processing, batching, and event-driven workflows are often better suited to high-volume or well-defined tasks. LLMs should be reserved for work where semantic reasoning adds meaningful value.

The framework in practice

Consider an out-of-home media-planning AI assistant asked to recommend high-performing advertising frames in Chicago for young professionals, subject to inventory availability in Q3 of the year, and to compare the recommendation with similar campaigns from the previous year.

A practical example: the System of Action combines governed data, connected context, and retrieval-ready knowledge to produce a grounded recommendation.

The System of Record supplies inventory, audience scores, frame-performance metrics, and historical campaign results. The System of Understanding connects audiences, locations, frames, campaigns, and the business meaning of performance measures such as reach, coverage, and frequency. It also exposes the quality and freshness information the AI assistant needs to judge whether the data is suitable for the recommendation. The System of Discovery retrieves historical campaign briefs, planning documents, audience research, and market reports. The System of Action brings that evidence together: it applies the current inventory constraints, identifies comparable campaigns, and generates a recommendation with supporting rationale and historical context.

The quality of the answer depends on more than the model. It depends on whether the enterprise has prepared the data, context, governance, and reusable signals that the model needs to reason well.

Preparing for durable enterprise AI

AI readiness is a data-engineering challenge as much as an AI challenge. More sophisticated agents can improve an application only to the extent that they can work with data that is discoverable, interpretable, trustworthy, connected, and usable.

The durable path is to build those capabilities progressively. Start by organizing and governing enterprise data. Make its meaning, quality, freshness, lineage, and policies visible to machines. Prepare knowledge for retrieval, connect entities across domains, and materialize high-value signals for reuse. Then let AI systems consume those foundations and help improve them over time.

As models become more capable, the differentiator will increasingly be the quality of the enterprise knowledge they can safely use. AI-ready data turns that knowledge from a passive asset into an active participant in reasoning and action.