OPEN GEOSPATIAL CONSORTIUM

Declared Meaning: Why AI-Readiness Is a Data Architecture Problem

A background article prepared in the context of the Metadata Summit 2026 and the OGC Testbed on Trusted Data and Systems, by Ingo Simonis, Robert Atkinson, and Piotr Zaborowski.

This background article is one of three related documents that set out the same idea at different levels. It makes the case for why AI-readiness is a matter of data architecture rather than of any single standard. The companion framework "Building Blocks for Interoperable, AI-Ready Geospatial Data" shows how that architecture is built and applied to real data, with worked examples. The OGC research paper on a machine-interpretable standards ecosystem, OGC document 26-021, provides the formal foundation on which both rest. In short, this is the why, the framework is the how, and the research paper is the underlying machinery.

OGC Document #: 26-051

Cite as: Simonis, I., Atkinson, R., Zaborowski, P., Villar, A., Toscano, M., Noardo, F. (2026) Declared Meaning: Why AI-Readiness Is a Data Architecture Problem. OGC Document 26-051. Open Geospatial Consortium, doi: https://doi.org/10.62973/26-051

ABSTRACT

Artificial intelligence does two useful things with data. It finds signals that would otherwise be missed, and it carries out at scale work that was previously possible only one case at a time. Neither ability creates new data; both multiply the value of data that has already been collected and paid for. The opportunity this opens is economic, yet what has to be built to capture it is technical, because a machine arriving at a dataset today cannot reliably establish what it is, how good it is, how it was produced, or whether it may be combined with the dataset beside it. The property that changes this is AI-readiness. It is a capability rather than a badge. A dataset or an endpoint is AI-ready when it is linked to further information richly enough that an agent can follow those links to an authoritative source instead of guessing, and when the links resolve, its behavior approaches the deterministic and the room for confident invention shrinks to almost nothing. This is where metadata matters, because the links an agent follows are the declared meaning that description provides. The mistake is to expect a single standard to supply it. The dominant habit, which ISO 19115 exemplifies, treats description as one fixed set of attributes, and that is too narrow, because a description must serve many purposes at once and no one schema serves them all. What is needed instead is a data architecture: a coordinated set of standards and well-defined patterns that state how meaning is declared, where it lives, how one source points to the next, and who maintains each piece. The reason to build it is economic. In the environmental domain the value of AI at scale takes the form of prevention rather than production: the flood anticipated while there is still time to act, the evacuation ordered early enough, the driver routed around a wildfire rather than into it. The cost of going without shows up as insurance claims, destroyed infrastructure, and lives, so the return is measured in the disasters avoided and the money never spent. That return depends entirely on trust, because a warning worth acting on is one whose basis can be traced back to the observations behind it. An architecture that makes meaning explicit and every link resolvable is what makes the result auditable, and building it is what turns a large public data holding into something an agent can be trusted to act on.

What this is for

Let's begin with something concrete and not at all technical. Suppose you could issue an instruction of roughly this kind and expect a dependable answer: Take the current weather forecast and the current wetness of the ground, look at every valley in the country, and tell me which ones could flood the way the Ahr valley flooded, while there is still time to act. Nothing in that request lies beyond today's science. The forecasts exist, the terrain models exist, the soil and runoff data exist, and the hydrology is well understood. It is nonetheless not how the work is done. Today a person has to find each dataset, obtain access to it, work out what is inside it, connect it to a model, and run the case by hand, valley by valley. There are thousands of valleys and a limited number of hydrologists, so the analysis is carried out for only a handful of places, usually those that have already flooded.

The Ahr valley is worth dwelling on, because it shows how narrow the gap between available data and useful warning can be. The July 2021 flood killed 190 people in Germany, 134 of them in that one valley, making it the deadliest flood in recent German history and the country's most expensive natural disaster. The subsequent analysis is the part that matters here. Roughly three quarters of the fatalities occurred outside the officially designated flood hazard zones, and the forecasts issued twenty-four hours ahead underestimated what was coming. The hazard maps existed. The forecasts existed. What failed was the step from data that was already in hand to a warning specific enough, early enough, and trusted enough to move people. That step is the subject of this article.

This is exactly the kind of work that artificial intelligence changes, and it is worth being precise about how. AI does two useful things here and not much else. It finds patterns and signals in data that nobody had the time to look for, and it performs at scale a task that was previously feasible only one instance at a time. Neither ability creates new data. Both multiply the value of data that has already been collected and paid for. In the environmental field, that value usually takes the form of something that does not happen. The flood that was anticipated, the evacuation ordered early enough, the dam opened onto farmland instead of a city, the driver routed around a wildfire rather than into the smoke. The economic argument is a prevention argument, which must be framed in terms of the disaster avoided and the cost never incurred.

Prevention of that kind only works if the result can be trusted, and trust here is a practical matter rather than a sentiment. Ordering twelve thousand people out of their homes is expensive and carries a heavy price if nothing then happens. The first question afterward is always the same. Where did that decision come from? If it cannot be answered by following a trail from the warning back to the observations behind it, then either the warning is not issued, or it is issued and ignored. Predict twice with nothing to show for it and on the third occasion nobody leaves. Being able to answer the question is therefore not an administrative nicety. It is the condition on which the whole capability rests.

What stands in the way is neither the science nor the availability of the data. What stands in the way is that a machine arriving at a dataset cannot reliably establish what it is looking at, how good it is, how it was produced, or whether it may be combined with the dataset beside it. A person bridges those gaps by knowing the field and telephoning a colleague. A machine cannot, so it either stops or, far more dangerously, fills the gap with something that looks right. This article is about what must change for an instruction of the first kind to become answerable, and its argument is that the required change is architectural.

Why this is a data problem rather than a metadata problem

Calling all of this a metadata problem is not incorrect, yet does real damage. The word suggests something small, separate, and secondary, a form completed after the real work is finished. The reality is that whether a fact counts as description or as substance depends entirely on where you are standing. To the utility that owns a gas pipeline, the fact that its position has been confirmed by survey rather than merely planned is a descriptive detail attached to a record. To the contractor about to open a trench beside it, that same fact is the most important information in the job, because it decides whether the line can be relied upon to be where the drawing says it is. Nothing about the fact changed. The role it plays changed. The same holds for a calibration date, a coordinate reference system, the definition of a measured quantity, or the identity of the model that produced a forecast. Each is background to one party and primary evidence to another.

It follows that the thing to be designed is a data architecture. When someone has a decision to make, what they need is to understand the data itself, where it came from, what it means, and how far it can be trusted. That understanding is exactly the material we have been calling metadata, which is what makes the label so misleading, because it is not a secondary layer sitting beside the data but part of knowing what the data is at all. The description and the thing described belong to a single estate and require a single set of mechanisms, a single way of declaring meaning, and a governance model. The word metadata remains useful where the point is specifically about the descriptive role, and it is used in that narrow sense in the pages that follow. The architecture, though, is a data architecture, and there is a practical reason to insist on the point. Label a proposal a metadata initiative and the people whose problems it would solve stop listening, because they do not have a metadata problem. They have a problem about data they cannot find, cannot interpret, or cannot defend a decision on.

What AI-readiness actually means

AI-readiness is generally used as though it were a badge awarded to a dataset, usually for adding a few more fields to a record. It is better understood as a capability. A dataset or an endpoint is AI-ready when it is linked to further information richly enough that an agent arriving at it can follow those links to an authoritative source rather than filling the gap itself. Readiness in that sense is what makes scale attainable. An agent able to resolve what a quantity means, how accurate it is, how it was produced, and whether it may be combined with a neighboring dataset can be set to work on ten thousand valleys or an entire national forest inventory, and its conclusions can be audited afterward, because every step it took leads back to something a person recorded deliberately. When the links are present and resolvable, and the agent follows them all, its behavior approaches the deterministic. Put the same question twice, and it takes the same route to the same answer, for reasons that can be inspected.

Set against that ideal is the failure mode the architecture exists to prevent. Where a needed piece of meaning is missing or ambiguous, an autonomous agent does not stop and report an obstacle. It infers, it approximates, and it may return an answer that reads exactly like a correct one. This deserves stating carefully, because the risk is not the colloquial idea of a machine producing obvious nonsense. The risk is good-looking nonsense, a confident and plausible result with no trail behind it, which then feeds a decision about an evacuation or a construction site and goes unquestioned until it is too late to matter. An agent that fails loudly is an inconvenience. An agent that fails invisibly is a liability, and the difference between the two is almost entirely a matter of what was declared.

Correctness is only half of what declared meaning buys; the other half is cost. The price of running an AI application scales with how much an agent must hold in its context and work out for itself as it runs. Each time it has to rediscover what a field means, infer an undocumented structure, or reconstruct how two datasets relate, it pulls more material into its context window and reasons over it, and that consumed context is the largest part of what a query costs. Declared, linked meaning removes most of that work, because the agent resolves a link and reads the answer instead of deriving it, a short and cheap lookup in place of a long and expensive inference. The effect compounds. A lower cost per query means the same budget buys many more queries, so analyses that were never worth running become worth running, and questions no one could afford to ask at scale become askable. Explicit knowledge is, in this sense, a set of cheap shortcuts, and each shortcut widens the range of problems worth attempting on a fixed investment. Readiness makes results not only trustworthy but affordable, and affordability is what turns a single demonstration into a portfolio of applications.

There is a second economic mechanism, and it is the one that scales fastest. Today, integration is arranged pair by pair. Each time two systems must exchange data, somebody works out how one side's terms, units, and structures correspond to the other's, and writes that correspondence into code that serves only those two. The effort therefore grows with the number of pairs, which grows with the square of the number of participants; a hundred systems that all need to talk to one another imply thousands of bilateral arrangements, each maintained separately and each breaking when either side changes. When meaning is declared once against shared registers, the arithmetic inverts. Every participant declares its own terms and structures once, against the common declarations, and interoperates with everyone else through them. Cost then grows roughly in proportion to the number of participants rather than the number of pairs, and each additional participant becomes cheaper to add rather than more expensive. OGC's research on this describes the effect as integration cost collapsing at the margin. It is the most readily quantifiable part of the business case, and it is what makes the architecture worth funding even before any AI agent appears.

So the question is not how much metadata to produce. It is how much meaning has been declared explicitly and connected to its source, rather than left for the next reader to reconstruct. Nearly everything an agent needs in order to avoid guessing falls under that heading: the data model including the attribute and value space characteristics, the units, the reference system, the classification scheme, the provenance, the accuracy, the conditions of use, the identity of the producing model, and above all the relationships among these and the links that make them resolvable. Where meaning is written down, validated, and connected to the source that justifies it, an agent can check its own interpretation before acting. Where it is implicit or the links are missing, the agent has no option but to fill the gap on its own. Why declaring and connecting meaning turns out to be so much harder than it sounds is the subject of the rest of this article. The difficulty is not carelessness and it is not a shortage of standards. It is structural, and it becomes visible as soon as one follows real data through a real production chain.

The problem, rolled out in detail

Vocabularies that overlap but refuse to align

The part of the problem that is closest to solved is also a warning about the rest. In mature domains the controlled vocabularies are well maintained, globally published, and available in several languages. This is a real achievement. Yet a number of those vocabularies cover the same ground while defining it differently, so they overlap without mapping cleanly onto one another. A single physical quantity such as skin temperature can carry one definition in one encoding and a subtly different definition in another, because the term was coined inside a community that never had to reconcile it with anyone else. Two records can therefore use the same word and mean two different measurements. A taxonomy cannot resolve this, because a taxonomy only arranges terms. What is required is an ontology, meaning a description of what each term actually asserts and how it relates to its neighbors. The lesson generalizes far beyond any one field. Agreement on labels is not agreement on meaning, and a machine that assumes otherwise will silently conflate quantities that a human expert would never confuse.

Description is not one thing, and it is not attached in one place

The dominant habit is to treat description as a set of attributes attached to a dataset, the model that ISO 19115 codifies and that a generation of catalogs implemented faithfully. As an account of a single dataset for the purpose of finding it, that is reasonable, and it has been produced at enormous scale. Its weakness is the quiet assumption that there is one description to attach, when the reality is a chain. Description is not a single artifact created once. It is generated at every stage of a long production chain, and the different stages serve different readers. Consider the path that observational data travels. A measurement begins at a satellite, a radar, a ship, an aircraft, a ground station, or a drone, and that raw observation is archived. It is then calibrated so that it becomes analysis ready, and the calibrated form is archived again. It is fed into a model, and the model output is archived. Quality control and long term monitoring follow, drawing on records from locations that have barely changed in a century or two, and those too are archived. Finally the chain terminates in a product, which may be a forecast, a warning, a map, or a broadcast graphic, and which is again described and stored.

Description exists at each of these stages, but it describes a different thing each time and is consumed by different people for different reasons. Discovery information helps someone establish that a collection exists. Evaluation information tells them whether it contains what they need and whether they are permitted to use it. Operational information tells them how to combine it correctly with something else. Provenance information tells them how it came to be. A record that answers one of these questions rarely answers the others, and the records produced at different stages are typically not linked to one another. The consequence is that no consumer, human or machine, can traverse from a final product back to the observation that produced it without leaving the descriptions behind and resorting to email. That is a limitation of current practice rather than an unavoidable one, and the next section sets out how such a traversal can be made to work.

Provenance is conceptually solved, operationally unsolved, and better assembled than authored

Provenance illustrates the gap between a good model and a working system. A standard way of expressing provenance already exists, the W3C provenance ontology PROV-O, and it does the right thing conceptually, modeling the world as entities produced by activities carried out by agents. What does not exist is the tooling to capture provenance continuously and to join it across the chain, because the people producing the data are occupied writing the code that produces it, not filling in lineage records. The stakes are easy to underestimate. Whether an instrument was last calibrated one year ago or ten years ago changes whether its readings can be trusted, yet that single fact is often known only informally. Some consumers depend on it absolutely, and others do not care, which means it can neither be dropped nor assumed. When the product at the end of the chain is a public warning that triggers an evacuation or an insurance claim, the legal weight rises sharply, and the line between raw data and finished product blurs, because both now carry consequences. Provenance is therefore not a scholarly nicety. It is the difference between a conclusion that can be defended and one that cannot.

The way out is to stop treating a lineage record as a document somebody has to author. Nobody positioned along a long production chain can describe the whole of it, and every attempt to make them try has failed for exactly that reason. What each participant can do is record the one thing they genuinely know, which is the activity they just performed, the inputs it consumed, and the outputs it produced. Records of that kind are small, local, and cheap enough for a processing step to emit as a side effect of running, which is the difference between a task people skip and a task that simply happens. The full lineage is then written by nobody. It assembles itself, because the output of one step is recognizably the input of the next, and tracing a warning back to the instrument that measured the first observation becomes a matter of traversal rather than reconstruction.

Self-assembly of that kind only works if the recording is classified consistently, and this is where the architecture becomes load-bearing rather than decorative. If one step reports that it calibrated something while the next reports that it corrected something, and one calls its output a radiance product while another calls the same thing calibrated radiance, the records will not join and the chain will not assemble. What has to be shared across everyone contributing is therefore twofold: a common approach to classifying the kinds of activity that occur, and a common approach to classifying the kinds of data that go in and come out. Identifiers must also stay stable long enough for the join to be made later, sometimes years later. None of this is a documentation exercise. It is a small set of agreed mechanisms together with maintained vocabularies that every participant points at instead of inventing their own wording, which is precisely why the patterns and registers described later in this article are the practical part of the proposal rather than the theoretical part.

Automation has failed for decades, from both directions

The obvious response to unfilled records is to automate their creation, and that response has been attempted repeatedly over roughly two decades without success. Manual capture fails because the people responsible have no time and little incentive, and because the person nominally in charge of a finished product is frequently a manager for whom the description means nothing and whose honest request is to be handed the product. Automated capture fails because the systems are not yet good enough to distinguish the things that need distinguishing, so fields either stay empty or are filled with something wrong. Neither manual entry nor naive automation has produced dependable descriptions at scale. It is worth being precise about what failed, though, because the two cases are not the same. What has repeatedly failed is asking somebody to author a rich description of a dataset once the work is over, at a moment when they have moved on and the details are already fading. Recording that a particular activity has just been carried out, on named inputs, producing named outputs, is a far narrower obligation, and it falls at the one moment when the facts are unambiguous, which is while the step is running. That distinction is why the self-assembly route described above is more promising than another attempt at form-filling. What cannot be assumed is that the missing meaning will appear later of its own accord. If it is not designed into the process, it will not be there when an agent needs it.

There is now considerable optimism that this is exactly where the current generation of language and multimodal models will help, and it is not misplaced. Such a model can read a dataset together with its column headers, its sample values, and whatever documentation surrounds it, and draft a description far more fluently than any rule-based tool that came before. The caution is that it inherits the failure mode described throughout this article. A model asked to describe data it cannot fully interpret does not decline. It produces a confident and plausible description that may be wrong, and a wrong description is more dangerous than a missing one, because it is believed. AI-assisted description is therefore best treated as a powerful drafting aid whose output is checked against something authoritative, rather than as a replacement for the declarations themselves. It also grows far more reliable the moment there are maintained vocabularies and resolvable sources for it to draw on and be measured against. AI helps most with description once the architecture exists, and the architecture is what lets AI help safely. The dependency runs in both directions.

Silos and the absence of cross-border connection

Even where descriptions are captured well, they tend to stay local. One country's observation repository is not connected to its neighbor's, so a system in one place cannot see what exists in the next, and this repeats across the entire set of national holdings. A town that straddles a border is described twice, in two schemes, with no statement of how the two relate, which is why the practical business of matching data along a shared boundary remains largely manual. The problem is not that any single repository is deficient. The problem is that the connective tissue between them was never built, so discovery stops at the institutional boundary. An agent that could reason across borders in principle is confined in practice to whichever silo it happens to be pointed at.

Precision, accuracy, and the missing bound on meaning

A further difficulty concerns how precise and how accurate a value really is. One source may offer global coverage at a resolution of fifty kilometers for each pixel while another resolves detail down to a meter, and there is no clean way to reconcile the two when they are combined. Coordinates are routinely published to nine decimal places, a precision of fractions of a millimeter, on data whose true positional error is a great deal larger, which invites every downstream consumer to infer an accuracy that was never claimed. The geospatial world has long had a coarse way of bounding data in space and time through bounding boxes and date ranges, and that supports a first pass at discovery. A bounding box lets an agent cheaply discard anything on the wrong side of the planet before it inspects a single dataset in detail, and a date range does the same for time. What is missing is the equivalent bound on the measured quantity itself. Suppose an agent is looking for temperature data to assess heat stress on a city. A search for temperature returns air temperature measured two meters above the ground, sea surface temperature, the radiative skin temperature of the land, soil temperature at several depths, and the brightness temperature a satellite sensor records, which is not an air temperature at all. Every one of them matches the word, and nothing coarse distinguishes them, so the agent has no choice but to open and interpret each dataset in full before it can reject the ones that are the wrong kind of temperature.

The same gap applies to the conditions under which a value holds and the limits within which it is valid. One dataset may be meaningful only over the ocean, another only under a cloud-free sky, another only within the range its sensor was calibrated for, and none of that can be posed as a coarse opening question either. What is needed is a bound on the parameter that works the way a bounding box works for location, a high-level declaration of what kind of quantity a dataset concerns, under what conditions it holds, and within what limits it is valid. Without it, an agent cannot ask whether a dataset is even the right kind of thing before it begins consuming the detail, which is the cheap filtering step that makes discovery at scale possible in the first place.

Within a domain it is tractable, across domains it breaks

The final and decisive complication is that these problems are manageable inside a single community and unmanageable between communities. Inside a domain people share enough tacit context to paper over the gaps. Across domains that shared context disappears, and the same failure recurs in every new initiative. A recurring pattern is that everyone agrees on the big picture, a new format should carry geospatial meaning by reusing an existing convention, and then everyone disagrees on the detail and drifts in separate directions. The cross-domain interfaces are precisely where meaning is most implicit and least visible, which is also precisely where an autonomous agent, which does not respect domain boundaries, is most likely to be working.

Why this is difficult, stated plainly

Taken together, these threads show that the problem is not one of missing standards. It is the opposite. The ground is already crowded with them. For describing a dataset so that it can be found and judged, there is ISO 19115. For stating how good the data is, there is ISO 19157. For recording how something came to be, there is the W3C provenance ontology PROV-O. For describing the instruments and their observations, there is the joint W3C and OGC sensor ontology known as SOSA and SSN. For cataloging holdings, there is the W3C Data Catalog Vocabulary, DCAT, together with its European and geospatial relatives, and there is STAC, the SpatioTemporal Asset Catalog, which OGC published as a Community Standard in 2026. Each of these is real, each is maintained, and each does a genuine job. The difficulty is that there are many partial standards, each applied partially, with no shared way of stating how they relate. Meaning is real but implicit. It is distributed along a chain rather than located in one record. It lives in separate systems that do not reference each other. And it degrades at exactly the domain boundaries where it is needed most. A human expert compensates by supplying context from experience. An agent has no such reservoir, so wherever the declared meaning runs out and the links stop resolving, it substitutes invention. That substitution is the mechanism behind good-looking nonsense, and it is why the description problem and the AI-readiness problem are the same problem. It also explains why another standard on its own cannot be the answer. Each new standard adds one more partial view. What is missing is the layer above them that states how the views relate, and that layer is an architecture rather than a standard.

A new case arriving now: when the producer is itself a learned model

A learned model, meaning an AI-powered model in contrast to a physics-based one, that produces a weather forecast looks at first like a new problem. To be usable, its output has to be accompanied by a description of how it was made: which model, which version, which training data, how it was evaluated, and the conditions under which the result is claimed to hold. On reflection, this is not new. It is provenance, the same information any product needs in order to be trusted, and a physics-based product requires the equivalent in its physical basis, resolution, and schemes. The difference is only that the physical basis has usually been safe to leave implicit, because it is standardized and widely understood, so records could stay thin without loss, whereas a learned model's basis is specific to that model and its training and cannot be reconstructed by inspecting the output. At the level of description, the two are indistinguishable, since the grid, the units, the parameter names, and the time structure are identical, so nothing warns a reader that the basis behind them differs. The requirement to describe how a product was made, always present, simply becomes unavoidable here. Retraining makes the same point, since a new version can change results for reasons that have to be recorded rather than inferred, and so does the spread of learned representations, whose values mean nothing without a record of the model that produced them.

Part of this ground is already covered. The OGC Training Data Markup Language for Artificial Intelligence, adopted as an OGC Standard, defines a model for describing geospatial training data, including how that data was prepared, its provenance and its quality, and how labeling and classification schemes are declared. That addresses the training corpus, which is one of the harder parts of the problem. What no adopted standard yet covers in the same way is the trained model as a published product, carrying its own identity, version, evaluation record, and declared envelope of validity alongside the fields it emits. The gap is narrower than it first appears, and it is a gap in coverage rather than in concept.

For the architecture, the lesson is not that a new kind of metadata is needed, since provenance is already among the roles it must serve. The lesson is a requirement on the architecture itself. A learned model is a new category of product, and new categories will keep arriving, each needing a descriptive pattern that no one anticipated. What an architecture must provide is extensibility: the means to add a pattern for a new category as it appears, without reopening and revising the architecture itself. This is exactly what a fixed schema cannot do, since anything unforeseen forces it to be rewritten. Read this way, learned producers are not a special case to be handled once but a standing test that any durable data architecture has to pass, because the supply of new product categories does not end.

The same conclusion is being reached elsewhere

It would be a mistake to present this diagnosis as a geospatial insight. The machine learning community has arrived at the same conclusion independently, and rather faster, which is worth knowing both as corroboration and as a warning about where the geospatial community now sits relative to it.

The clearest instance is Croissant, a metadata format for machine learning datasets developed through MLCommons and built as an extension of schema.org, so that datasets described with it are discoverable by ordinary web search. Version 1.0 established a machine-readable structure for dataset description. Version 1.1, announced in February 2026, went considerably further and in a direction this article will recognize. It added machine-actionable provenance for complete data lineage, adopting the W3C PROV-O model rather than inventing its own; vocabulary interoperability, so that metadata can reference domain-specific ontologies instead of restating them; and structured usage policies that allow consent and licensing conditions to be enforced automatically. Its stated purpose is to make datasets interpretable and reusable by autonomous systems, which is the same definition of readiness proposed here. Adoption is not speculative: the format covers hundreds of thousands of datasets across the major repositories, and NeurIPS now requires Croissant metadata for submissions to its dataset tracks.

The second instance is the layer through which agents now actually reach data. The Model Context Protocol, introduced in late 2024 and since placed under neutral foundation governance, has become the de facto way for an agent to discover and call external tools and data sources, with adoption measured in the tens of millions of monthly downloads. What matters here is not the protocol's popularity but its shape: a server declares what it offers, and a client discovers those declarations before invoking anything. That is self-description at the endpoint, arrived at from the agent side rather than the data side. MLCommons has already demonstrated Croissant and the Model Context Protocol working together, joining dataset description to agent access. In the geospatial community the equivalent bridge is being built too, with work under way to expose the OGC API family through the Model Context Protocol so that models interact with geospatial services through declared tools rather than improvised calls. The same body of work also confirms the cost argument made earlier, since a recognized problem in practice is precisely that injecting every available schema into a model's context is expensive, which is what drives the search for resolvable references instead.

The corroboration is genuine and it should be welcomed. The risk is equally real, and it is the risk this article has already described in another form. Croissant, STAC, ISO 19115, and DCAT now overlap substantially in what they describe, each with its own communities, tooling, and governance. Without a declared statement of how they correspond, and of where they deliberately do not, the outcome will not be one interoperable ecosystem but two or three well-built ones that cannot be joined, which is the original problem reproduced at a larger scale. Declaring those relationships is not a matter of choosing a winner. It is the third element of the architecture set out below, applied to the most consequential boundary now open.

From the problem to a data architecture

The instinct to answer all of this with one grand standard should be resisted, because it is neither achievable nor appropriate. The communities involved do not start from the same place, they carry incompatible histories, and forcing them under a single schema would either fail outright or flatten distinctions that carry real meaning. Consider tasking a satellite. One schema for tasking can be written, but it will never accommodate the genuine differences among all the satellites that might be tasked. The workable answer is an architecture. An architecture in this sense is not another standard. It is the layer that coordinates a set of standards and a body of well-defined patterns, stating what each is for, organizing them into roles, saying how they join up, and naming who is accountable for each part.

The word pattern deserves a plain definition, because it carries much of the weight of the proposal. A pattern is a mechanism, an agreed way of doing a recurring thing. How do you link from your dataset to someone else's, and what should a reader expect to find on arriving at the other end? How do you extend a standard to cover a case it was not written for? How do you state what your extension adds, so that someone else can see the difference without reading your documentation? How do you point at the definition of a term rather than merely using the term? How do you record what you have just done, in a form that will join up with the records made by everyone upstream and downstream of you? How do you update an entry in a register, and how do you reconcile two registers when one calls a property precision, and the other calls it accuracy? Each of those is a mechanism. Limiting the degrees of freedom for each mechanism is exactly what the architecture with its clearly defined patterns needs to achieve.

The cost of improvising them is easy to see in practice, and STAC is the instructive case precisely because it succeeded. It was designed as a lightweight catalog format for collections of satellite imagery and became popular because it was simple and did that job well. It offered an extension mechanism, so communities began extending it to describe data it was never intended for, which was entirely reasonable, and its coverage now runs from radar and lidar to point clouds and machine learning labels. The result is a large and growing collection of extensions that may be published anywhere, with no systematic way to tell whether any two of them overlap, contradict one another, or could have reused one another. To find out whether one suits your need, you read them all. Very little of the interoperability the format promised survives that. What is missing is not effort or goodwill but agreed mechanisms: how an extension is derived from the core, how it declares what it does differently, and how a term such as traffic camera is bound to a definition maintained somewhere dependable rather than explained in prose on page sixty-nine of a specification, where no machine will ever find it. Since OGC published STAC as a Community Standard in 2026, this is no longer a problem to be observed from outside. It is a governance question OGC now owns, and an opportunity to demonstrate the mechanisms on a specification people actually use.

The same shape appears whenever a quality statement is needed. A general standard for describing data quality exists in ISO 19157. It is written to apply across every kind of geographic data rather than being tailored to any one of them, and it describes the quality of the data itself rather than the records that describe the data. To say something useful about a particular dataset you therefore do not use all of it. You take the part that matters, and it differs completely from one kind of data to the next. For a map that classifies each patch of ground as forest, water, or built-up land, the question that matters is how often the label is correct, which calls for a measure of classification accuracy. For a surveyed property boundary, the label is not in doubt, and the question that matters is how close each point lies to its true position, which calls for a measure of positional accuracy in centimeters. Classification accuracy is meaningless for the boundary, and centimeter positioning is beside the point for the land-cover map, so each draws on a different part of the standard and may extend it. What each ends up with is a self-contained, reusable piece of the larger standard, tailored to one purpose, and a piece of that kind is called a building block.

This is not a coinage of the present article, and the direction is not speculative either. The 2023 revision of ISO 19157 made the move itself: the catalogue of quality measures that the previous edition carried inside it was taken out of the standard and placed in a separate part of the series, to be maintained as a reusable set rather than as text inside one document. That is refinement and registration in the sense used here, performed by a standards committee on its own material. The idea has been worked out under exactly that name in OGC's research on a machine-interpretable standards ecosystem, document 26-021, which defines a building block as a self-contained, independently testable package and sets out aggregation, refinement, extension, and substitution as the formal operations for composing and adapting them. The definition is in section 3.3, the anatomy and the composition operations in section 4.3, the profiling mechanism in section 4.2, and the register family in section 4.4. A practical framework that applies these building blocks to real data, with worked examples, a governed check-in workflow, and a full standards stack, is set out in the companion document "Building Blocks for Interoperable, AI-Ready Geospatial Data." Where the present article makes the architectural case, that framework shows, on a concrete example, how it is built.

Multiply that across every domain and every kind of quality statement, and the picture looks less like one large standard than like a set of composable building blocks with declared relationships between them.

A third case is the most telling, because it comes from one of the most carefully coordinated efforts in the field. The generic vocabulary for describing catalogs, known as DCAT, is a good citizen of the kind this article argues for. It is a W3C standard, and rather than reinventing terms for titles, licenses, themes, and provenance, it deliberately reuses established vocabularies for each, among them Dublin Core, SKOS, FOAF, and the W3C provenance ontology PROV-O. Built on it is DCAT-AP, a European application profile that constrains the generic vocabulary for use across European data portals, tightening what is mandatory and which code lists are allowed without adding anything foreign to the base. So far, this is the model working as intended.

The complication appears one step further out, with the geospatial version known as GeoDCAT-AP, which began as a European Commission initiative. It is often described loosely as a profile of a profile, and it is neither. It is an extension rather than a profile, and the distinction matters. A profile narrows a base and stays within it. An extension adds elements the base does not contain, and GeoDCAT-AP does the latter, binding into the same framework the union of the ISO 19115 core profile and the INSPIRE metadata elements, so that it carries geospatial content DCAT-AP has no notion of. Because it adds rather than only constrains, the tidy picture of ever-tighter profiles nested one inside another simply does not describe what is there.

Worse, the base and the extension do not sit together cleanly. They overlap, and in places they conflict. The documented example concerns the controlled list of themes a record must carry. The European profile mandates one such list while the geospatial extension requires another, so a validator built faithfully for the base rejects records that are valid under the extension, and the reverse. This is not an exotic corner case. It is the everyday consequence of two standards sharing a boundary that neither fully owns. The revealing part is the response to it. The people maintaining these specifications have openly posed, as an unresolved question, how the relationship between such profiles should even be expressed and checked, and whether the existing means of describing a profile is adequate to the task. Each time the underlying vocabulary is revised, the profile and then the extension have to be realigned in turn, and that realignment is done by hand.

The lesson is not that any of these efforts is deficient. Each is careful, well resourced, and centrally coordinated, which is exactly why the case instructs. If a coordinated European programme with dedicated institutions behind it has no settled, machine-checkable way to state how its own catalog standards relate, then the missing thing is plainly not another standard. It is the mechanism for declaring the relationships between standards, and supplying that mechanism is what an architecture does.

Against that background the architecture has five things to deliver, none of which any individual standard delivers on its own.

Running through all five is governance, which is not a separate concern to be bolted on at the end but the other half of every technical element. A register implies an interface for writing to it, reading from it, and changing an entry, and each of those implies a decision about who may do so and by what process. A vocabulary is only worth linking to if it will still resolve in ten years, which means somebody has to be accountable for keeping it available and for publishing how it is governed. A profile that anyone may quietly edit is not a foundation to build on. In practice the technical and the governance questions arrive together: not only how a term is defined but who defines it, not only how a mapping is expressed but who may change it and what becomes of everything downstream when they do. An architecture that specifies the mechanisms and leaves the accountability unstated has solved the easier half of the problem.

How the architecture delivers AI-readiness

Once the roles are named, their contents declared, their relationships written down as testable artifacts, and their custodians identified, AI-readiness stops being a badge and becomes a property of the system. A machine can traverse the chain instead of guessing across it. It can establish that a collection exists, read the declared meaning of its contents, follow the provenance back to the observation and the calibration that produced it, judge whether the accuracy and the reference system suit the task, check the conditions of use, determine whether the product came from a physical model or a learned one and what that model was trained on, and verify its own interpretation against the declared mappings before it acts. Following links replaces guessing with lookup, and the more completely the links resolve, the closer the behavior comes to deterministic. At every step it is reading meaning that somebody put there on purpose rather than reconstructing meaning that nobody recorded.

Return now to the instruction this article began with. Every element it requires already exists somewhere: the forecast, the wetness of the ground, the shape of the terrain, the hydrological model, the record of which places have flooded before. What the architecture adds is the ability to find each one, establish what it means, judge whether it fits, and combine them without a person in the loop for every valley, ten thousand times over, at a cost that makes prevention practical rather than aspirational. Just as importantly, it leaves the trail from the warning back to the observation intact, so that when somebody asks where the number came from there is an answer, and the next warning is still believed. In the Ahr valley the hazard maps did not describe where people would actually die. An architecture cannot supply a better forecast, but it can make the difference between a model run that nobody can check and one whose every input, assumption, and limit is stated and can be argued with before the water arrives.

That is the whole argument. Where the architecture is present, an agent reaching the edge of what it knows finds a declaration telling it what to do, or an explicit statement that no safe correspondence exists, which is itself an instruction to stop. Where the architecture is absent, an agent reaching the same edge finds nothing and improvises, and the result is confident, plausible, and unaccountable. The architecture does not make the machine cleverer. It makes the machine's work checkable, which is the only basis on which anyone will act on it. Trust does not become deliverable. It becomes auditable.

How this architecture is put into practice, step by step and on real data, is the subject of the companion framework "Building Blocks for Interoperable, AI-Ready Geospatial Data."

ANNEX A

Architecture Overview

This annex collects four diagrams that present the architecture from four angles: its overall structure, the anatomy of a single building block, federation across governance tiers, and the practical view of a publisher and a consumer of data.

The first diagram gives the overall structure and reads from top to bottom. A consumer poses a task. The registers are the connective tissue through which that task is searched and resolved. Those registers point in turn to the composable specifications and to the vocabularies they draw on. Everything ultimately describes the data and the API endpoints, which sit across the governance tiers.

Figure 1: Architecture overview

The second diagram opens a single building block. Each block is a self-contained, independently testable package, and the six cells correspond to the artifacts it carries in the OGC framework: an identity and dependency declaration, a schema defining its structure, a semantic context in which its terms are bound to shared vocabularies, a set of executable validation shapes, transformation specifications, and examples that are tested continuously. Two of these carry most of the architectural weight. The semantic context is where terms bind to shared vocabularies, and the validation shapes are where a written requirement becomes a check that runs.

Figure 2: Anatomy of a building block

The third diagram reads from left to right and shows federation across governance tiers. On the left are five tiers, from international standards bodies down to communities of practice, each autonomous internally, with its internal systems out of scope. In the centre is the federation boundary, which every tier connects to independently; the double-headed connectors indicate that the only interaction between tiers is the boundary crossing itself. On the right are the six elements that must be declared for a crossing to work, each with the established standards that already serve it: identity through persistent identifiers, semantics through SKOS and OWL, structure through JSON Schema and XML Schema, access through the OGC API family and OpenAPI, provenance through PROV and ISO 19115 lineage, and quality through ISO 19157 and the Data Quality Vocabulary. The consequence that matters is that data need not climb the tiers. Any tier can federate directly with any other as soon as those six elements are present.

Figure 3: Federation across governance tiers

The fourth diagram takes the view of someone publishing data and someone consuming it, and reads from top to bottom in three bands. In the upper band, two organizations publish through an OGC API — Features endpoint, and each declares five things: the endpoint and its collection, the self-description surface of conformance, queryables and schema, the feature type as a registered identifier, the observed-property terms bound to a vocabulary, and the sampling strategy. The two are deliberately unlike each other. One operates an automatic sensor recording every fifteen minutes at a gauging station and uses the term water level, while the other performs a monthly manual staff-gauge reading at a measurement site and uses the term stage height.

The middle band is the shared infrastructure: registers for feature types, for vocabularies with their crosswalks and their declared non-correspondences, for building blocks with their schemas, contexts and shapes, for transformations covering units, coordinate reference systems, term crosswalks and temporal aggregation, for profiles, and for validation.

The lower band is the consuming agent. It discovers both collections, reads their self-descriptions, resolves the feature types and terms in the registers, and then reconciles three mismatches. The terminology differs, and a crosswalk resolves it. The units and reference systems differ, and a registered transformation applies. The sampling strategies differ, and here the honest options are to aggregate the finer series or to declare the two not comparable. That last case is the practical form of the declared non-correspondence described in the third architectural element above.

Two features of the drawing carry meaning. The dashed channels running down the outer margins are the data itself, travelling directly from the endpoints to the agent and bypassing the registers, which makes the point that registers carry interpretation rather than payload. The closing bar states the outcome: a combined series with declared provenance, stated limits, and a trail back to each observation.

Figure 4: The builder's view

ANNEX B

A Worked Example: The Life of a Weather Forecast and Where Its Description Lives

This annex is the concrete companion to the article. It lays out the physics-based production chain in full, and the section on learned producers builds directly on the chain described here.

The argument of this article is easier to follow with a single domain held in view from beginning to end. Weather and climate data is a good choice, because almost every difficulty described above appears in it at once, and because the stakes at the end of the chain are high enough that the cost of getting description wrong is easy to see. What follows is one worked example. It follows data from the moment it is observed to the moment a warning reaches the public, and at each step it points out the piece of meaning that someone knew and no record captured.

The part that already works, and why it is still not enough

Start with the good news. This is a field that has taken metadata seriously for decades, and in one respect it has succeeded. Its parameter vocabularies are well maintained, published openly, and available in many languages. If you want to know what a named quantity is called, the answer exists and is easy to find. The trouble is that there is more than one such vocabulary, and they do not line up. Parameter names come from several encodings that grew up separately, among them the GRIB and BUFR formats used to move operational data and the CF conventions used with netCDF files in research. Each covers roughly the same physical world, and each describes it a little differently. A quantity as ordinary as skin temperature, the temperature of the very surface of the land or the sea, can carry one definition in one system and a subtly different definition in another. Two datasets can therefore use the same word and mean two measurements that are not the same. A simple list of approved terms cannot fix this, because the problem is not the spelling of the term. The problem is the meaning behind it, and pinning meaning down precisely enough for a machine needs an ontology, which is a description of what each term asserts and how it relates to the others.

The journey of a single measurement

Now follow the data itself. A measurement begins at an instrument. It might be a satellite, a ground radar, a sensor on a ship or an aircraft, an automatic weather station that has stood in the same field for a century, or a drone flying through the air or diving under the sea. That first raw observation is placed in an archive. Before it can be used it is calibrated, which means it is corrected and adjusted so that its readings can be trusted and compared with others, and the calibrated version is archived in turn. It then feeds a forecast model, whose output is archived again. Along the way the data used for long-term climate monitoring is treated with particular care, because a station that has not moved in one or two hundred years is precious, and its readings are cross-checked against reference measurements whose audit trail leads back to the international measurement standards maintained near Paris. Each of these steps produces a description, and each step's description describes a different thing for a different reader. The record that helps a modeler is not the record that helps a climate scientist, and neither is the record that a downstream system needs in order to combine two datasets correctly. The catch is that these records are rarely linked to one another. You can hold the final product in your hand and still be unable to trace it back to the observation that started it without leaving the descriptions behind and emailing a colleague.

Provenance, and a small story about a calibrated instrument

This is where provenance matters, and a small concrete story makes the point. Consider an instrument owned by a meteorological service in a country with very little money. To calibrate it properly, someone carries it to a training course in Europe, has it calibrated once, and takes it home. Good practice would have that instrument recalibrated every couple of years. In reality it may be relied upon for ten. Whether the instrument was last checked two years ago or ten years ago, and whether it was ever checked against the international standards at all, changes how far its readings can be trusted. For some users that single fact is decisive. For others it does not matter in the slightest. This is why the fact cannot simply be dropped, and why it cannot simply be assumed. It has to be recorded and carried with the data, and today it usually is not, because the tools to capture it as a matter of routine do not exist. A good model for writing provenance down has existed for years. What has never existed is the everyday tooling that would let the people doing the real work, who are busy writing code in Fortran, Java, or Python, record it without extra effort.

Data, products, and twenty years of failed automation

There is also a distinction that the field draws sharply and that outsiders often miss, between data and a product. The data is the measured or modeled field. A product is the thing made from it for someone to use, such as a map, a video, a broadcast graphic, or a written bulletin. The description is supposed to be captured at the moment the product is created. In practice the person responsible at that moment is often an account manager or a member of a marketing team, for whom description means nothing and whose honest request is simply to be handed the finished product. Attempts to take the human out of this step and generate the description automatically have been made repeatedly for about twenty years, and they have not succeeded, because the systems are not yet good enough to tell apart the things that need telling apart. So the description is either missing or wrong, and neither manual effort nor automation has closed the gap.

Step back and look at the archives as a whole and a different failure appears. The great international climate efforts, the successive rounds of coordinated model experiments known as CMIP5, CMIP6, and now CMIP7, produce data on a scale measured in exabytes. Every institute that takes part follows standards. The difficulty is that they follow different ones. These are not dead archives that no one visits. They are living collections that scientists really need to read, and searching them is part of the daily job. Yet the data is often simply too hard to find, so it is not found, and effort is duplicated or results go unused. The information needed to locate the right dataset exists somewhere in someone's record. It is just not expressed in a way that a search can use.

The same problem, repeated across borders

The disconnection repeats across borders. The people who hold one country's observations do not know what sits in the repository next door. A service in one country cannot see into the archive of another, and this holds not for two or three countries but for the whole set of nations that observe the weather, which numbers well over a hundred and ninety. Each archive may be perfectly good on its own. What was never built is the connective tissue between them.

How good is the number, really?

Now a subtler problem, about how good a number really is. One dataset offers global coverage at a resolution of tens of kilometers for each pixel. Another, from a radar or a high-resolution satellite, resolves detail down to a meter. When someone wants to combine them, there is no clean way to reconcile the mismatch, and no agreed way even to state it plainly. A common habit makes this worse. Coordinates are routinely written to nine decimal places, which implies an accuracy of a fraction of a millimeter, when the true error is very much larger. The precision on the page is not the accuracy in the world, and there is no standard, high-level way to declare the difference so that a consumer, whether a person or a machine, is warned.

A bounding box, but for the parameter

From this comes one of the most useful ideas in the whole example. In the geospatial world there is a familiar, coarse way to begin a search. You draw a box on a map and pick a range of dates, and that bounding box in space and time narrows the field before you look closely. There is no equivalent for the quantity being measured. If you are interested in skin temperature, you cannot ask the coarse opening question of what kind of thing it even is, whether it is a physical quantity or a chemical one, before diving into the detail. What is missing is the equivalent of a bounding box, but for parameters, a simple high-level way to bound the measurement of interest so that discovery by subject can work the way discovery by place and time already does. Building it, once again, leads back to ontologies, because bounding meaning requires first describing meaning.

Describing the instrument itself

There is a mature way to describe the instruments themselves that shows what good practice looks like. Instead of trying to list every possible property of every possible sensor in advance, the characteristics of a sensor are themselves recorded as observations about that sensor. Its behavior under ideal conditions, under extreme conditions, and under the conditions it can merely survive can all be stated, and the different facts can come from different authorities, some from the manufacturer and some from the body that calibrated it. The model is rich and it works. What is missing is a standard way to attach these descriptions to the data they belong to, so that the two travel together. A linked-data encoding is the natural mechanism for making that attachment, but the pattern for doing it consistently is not yet settled.

Why the end of the chain makes all of this matter

Finally, follow the chain to its end, because the end is what makes all of this matter. The last products are forecasts and warnings, including warnings of thunderstorms and tornadoes, some of them written by hand as text. These feed public alerting systems, among them the European service that gathers national warnings together, and such systems are governed strictly, because the consequences are real. A warning can trigger an evacuation, move money, and decide an insurance claim. Parts of the data and the products therefore carry genuine legal weight, and at that point the tidy line between data and product blurs completely, because both now have consequences. The account managers and marketing teams at the very end still only want the finished product, and they are not wrong to, but everything that lets a machine or a careful person decide whether that product can be trusted lives in the description that was, or was not, captured all the way back down the chain.

What the example shows

That is the whole example, and its lesson is simple. At every stage something was known to someone. The definition behind the parameter, the date of the last calibration, the true accuracy of the number, the model that produced the field, the conditions under which it holds. None of these facts was secret. Each was simply left unstated, or stated in a place and a form that the next reader could not use. The failures here are not failures of measurement or of science. They are failures to declare, and to connect, meaning that already existed. That is exactly the gap a data architecture is meant to close, and it is why the abstract argument and this very concrete story are, in the end, the same argument.