Building Blocks for Interoperable, AI-Ready Geospatial Data: The Next Chapter of Spatial Data Infrastructures and Geospatial Ecosystems

A practical framework for digital twins and data spaces

This document reflects the current state of an ongoing conversation on interoperability and AI-readiness among OGC members and partners. It draws on work that OGC has carried out with its members across pilots and testbeds, and in EU-funded research projects, and on the wider debate about the future of spatial data infrastructures and geospatial ecosystems. It is offered as a proposal and a basis for discussion, a shared and evolving reference rather than a normative position or the view of any single organization.

OGC Document #: 26-050

Version 1.0, 12 August 2026

Cite as: Simonis, I., Atkinson, R., Zaborowski, P., Villar, A., Toscano, M., Noardo, F. (2026) Building Blocks for Interoperable, AI-Ready Geospatial Data: The Next Chapter of Spatial Data Infrastructures and Geospatial Ecosystems. OGC Document 26-050. Open Geospatial Consortium, doi: https://doi.org/10.62973/26-050

Table of Contents

Acknowledgements

Executive summary

1 Introduction

1.1 What this document is and what it does

1.2 Who should read this, and what you will get from it

1.3 Scope of this version

1.4 What the framework is, and what it is not

2 The problem: why data is hard to combine

3 The core idea: building blocks and four mechanisms

3.1 What a building block is

3.2 Four mechanisms, four questions

3.3 The interoperability layers

4 How data flows through the framework

4.1 What a digital twin execution produces

4.2 Linking inputs, outputs and the model

5 Two worked examples: why separate mechanisms beat one schema

5.1 Example one: the same real-world quantity, two structures, one meaning

5.2 Example two: one structure, two unrelated meanings

5.3 The general pattern

5.4 The payoff: full provenance and easy reuse

6 Data profiles and canonical formats

6.1 Multidimensional spatio-temporal data

6.2 Tabular and statistical data

6.3 Graph and knowledge data

6.4 Workflows and applications

6.5 From Standards to Building Blocks

7 Controlled vocabularies

7.1 Reusing authoritative vocabularies

7.2 How building blocks attach meaning to the data

8 Catalog profiles and provenance

8.1 Catalog hierarchy

8.2 The catalog metadata model

8.3 The provenance model

8.4 Metadata profiles

8.5 Building blocks behind the catalog and provenance

9 The check-in workflow

9.1 Governance

9.2 The stages

9.3 The three building-block roles

9.4 The validation contract

9.5 Handling gaps

10 Building blocks in practice

10.1 What is inside a building block

10.2 The convention for semantic uplift

10.3 Demonstrator mapping

11 Reading a building block in the viewer: a worked example

11.1 The About tab

11.2 The Dependencies graph

11.3 Following one dependency chain from top to bottom

11.4 The Examples tab

11.5 The Data Structure tab

11.6 The JSON Schema tab

11.7 The Semantic Uplift tab

11.8 The Validation tab

12 Current status and next steps

13 Getting started

14 Conclusions

References

Appendix A. Standards stack

Appendix B. Building-block index

B.1 Reused from ILIAD

B.2 Defined by SeaDOTs

Appendix C. Vocabulary registry

Appendix D. Namespace registry

Appendix E. Tools inventory

Appendix F. Assisted tooling inventory

Appendix G. Data flow architecture

Appendix H. Definitions, acronyms and abbreviations

Acknowledgements

First and foremost, the research behind this document was made possible by funding from OGC strategic members and the European Commission. The framework presented here grew out of several EU-funded research projects as a collaborative effort between OGC and OGC member organizations, and was further explored and refined in OGC Pilot and Testbed activities. The EC-funded SeaDOTs1 project provides the worked examples used throughout this document, and SeaDOTs itself reused many of the building blocks resulting from other OGC standardization efforts as well as the ILIAD Digital Twin of the Ocean project. All these projects were carried out with many partners, and what is presented here reflects their combined work rather than any single contribution. Nevertheless, the content of this document reflects only the authors' view. It is not an official OGC standard or position.

A development of this kind cannot be produced by any one group working alone. It depends on continuous discussion and exchange among the many parties involved: the users who state what the data must answer, the researchers who define the concepts, the developers who build the tools, and the portal and platform operators who publish and serve the results. That sustained conversation across all of these perspectives is what turns a collection of standards into a working, shared framework like the one documented here.

Individual parts of the approach were advanced and tested through OGC Pilot and Testbed activities. In that context, OGC gratefully acknowledges Natural Resources Canada, the US National Oceanic and Atmospheric Administration, and the US National Geospatial-Intelligence Agency, the United Kingdom Hydrographic Office, and above all the European Space Agency (ESA) and the US National Aeronautics and Space Administration (NASA). Through their strong support for open-science research, ESA and NASA in particular have contributed substantially to the present understanding of what building blocks are, how they should be shaped, and what functions they can serve.

OGC also acknowledges the United States Geological Survey, which commissioned work on the future of spatial data infrastructures that has helped frame the questions this document addresses.

OGC also gratefully acknowledges the United States Geological Survey (USGS) and the General Authority for Survey and Geospatial Information of Saudi Arabia (GEOSA). Both have commissioned OGC to explore and help define the future of spatial data infrastructures, a future now increasingly described as the geospatial ecosystem and one of the central forces driving innovation across the field. OGC works closely with both organizations on this agenda, and their leadership in commissioning and guiding the work has been instrumental in advancing the shared understanding that this document builds on.

Finally, OGC thanks its members and sponsors, whose sustained engagement and support have carried this work forward and brought the community a good deal closer on the shared journey toward optimal interoperability and AI-readiness.

Executive summary

Across the geospatial community, more and more value depends on data being usable not only by the team that produced it, but by other people, other organizations, and increasingly by software and AI agents. Environmental observations, model outputs, statistics, and survey data arrive from different communities in different structures, described differently and named with different vocabularies. Left as they are, they cannot easily be combined, compared, or trusted, and a machine arriving at them has to guess what they mean. This document sets out a practical way to close that gap.

The central design choice is deliberate. Rather than forcing every dataset into one common schema, the framework keeps each community's structure and connects those structures through shared discovery and provenance mechanisms. Four mechanisms do the work. Data profiles describe how data is structured. Catalog profiles describe how data is found and linked. Controlled vocabularies describe what terms mean. OGC building blocks package all of this into reusable, testable components. The building block is the piece that ties the others together, and much of this document explains what it is and why it makes data interoperable, verifiable, and ready for automated use.

To keep the framework concrete rather than abstract, it is illustrated throughout with a running example: SeaDOTs, a research project building digital twins of the ocean, which has to combine exactly these kinds of data and whose work is public. The examples show how the approach applies to the data that goes into a model run, the products that come out of it, and the record of how those products were made. The framework itself is general, and nothing in it is specific to the ocean domain or to any one project. Data spaces and digital twins are two settings where the need is especially visible, but the same approach serves any organization that wants its data to be found, understood, trusted, and reused beyond its own walls.

A word on how this document should be read is important, because it goes to the heart of why it was written. This is a proposal and a basis for discussion, not a normative OGC position or a specification that anyone is required to follow. Nothing in it is new or independent research surfacing for the first time. Its purpose is the opposite. It draws together the many initiatives, pilots, testbeds, and community conversations of recent years and brings the current state of that shared discussion to a single, coherent point. It is offered to be read, challenged, and improved by the community it draws on, and it is expected to evolve as that conversation does.

The framework is also honest about its limits. It shows how building blocks apply to the inputs, outputs, and lineage around a model run, and it does not yet address the running, coupled system in full. That boundary, and the work still open, are set out in the status section, so that the reader knows exactly what is proposed, what is already demonstrated, and what remains to be worked out together.

Figure 1: From fragmented data to a linked, AI-ready ecosystem. Data arrives from different communities in different structures, with different metadata and vocabularies, and no single common schema can hold it. OGC building blocks solve this by binding structure (data profiles), discovery (catalog profiles), and meaning (controlled vocabularies) into one versioned, testable artifact, which a governed check-in workflow then validates and publishes. The result is data whose provenance is complete and whose meaning is explicit enough for people and software agents to use without guessing. The two panels below show the two cases that make the approach powerful: two sources with different structures but a shared meaning, and two unrelated datasets that share only a structure.

1 Introduction

1.1 What this document is and what it does

Across the geospatial community, the same problem keeps returning in new forms. Environmental observations, numerical model outputs, biodiversity information, socio-economic statistics, indicator models and digital twin workflows come from different scientific communities. They use different data structures, different metadata conventions and different vocabularies. Left as they are, they cannot easily be combined, compared or reused, and increasingly they cannot be consumed reliably by automated and AI-driven tools either.

This document sets out a practical framework for making such data interoperable and machine-actionable without forcing every dataset into a single grand schema. Its purpose is to define a small set of shared, reusable interoperability contracts, and to explain how any dataset can meet the contract that fits it. The approach is built on OGC building blocks, and it applies as much to a single dataset that someone wants to publish as to a large collection that has to be integrated across organizations.

Seen this way, the document is best read as a practical roadmap for anyone who wants to make data genuinely usable by others, and increasingly by software and AI agents, rather than only by its original producer. Digital twins are one prominent setting for this, because they consume and produce many kinds of data and depend on clear provenance. Data spaces are another, because they exist precisely to let independent participants share and reuse data under common rules. The need, however, is far more general. Any organization that wants its data to be found, understood, trusted, and acted upon by parties and systems it does not control faces the same problem, and the building-block approach set out here is a way to solve it once and reuse the solution wherever the data has to travel.

The framework is not the work of one project or one author. It reflects the current state of an ongoing conversation among OGC members and partners, developed over many years across pilots, testbeds and collaborative innovation initiatives, and in a series of research projects. The wider debate on the state and future of spatial data infrastructures and geospatial ecosystems has fed directly into it. What is written here is best read as a shared, evolving reference that many organizations have helped to shape, not as the position of any single institution.

This document is the practical member of a set of three. It shows how interoperable, AI-ready data is built in practice. The background article "Declared Meaning: Why AI-Readiness Is a Data Architecture Problem" makes the case for why this is a matter of data architecture rather than of any single standard, and the OGC research paper on a machine-interpretable standards ecosystem, document 26-021, provides the formal foundation on which both rest. The chapters that follow are best understood as one worked instantiation of that architecture, made concrete with a public example.

To keep the framework concrete rather than abstract, it is illustrated throughout with a running example: SeaDOTs, a research project building digital twins of the ocean. SeaDOTs is a useful setting because it has to combine exactly the kinds of data described above, and because its work is public. Its three demonstrators, in Germany, Norway and Sweden, and its target platforms, the European Digital Twin Ocean known as EDITO and, optionally, the European Open Science Cloud known as EOSC, supply the worked examples used in later chapters. The framework itself, however, is general, and nothing in it is specific to the ocean domain. SeaDOTs integrates environmental observations, numerical model outputs, biodiversity information, socio-economic statistics, indicator models and digital twin workflows. These assets come from different scientific communities. They use different data structures, different metadata conventions and different vocabularies. Left as they are, they cannot easily be combined, compared or reused.

This document uses the interoperability framework developed for the project to explain how data of different types and natures, with different origins and governance regimes, can be integrated into a digital twin and made AI-ready. The purpose of the framework is not to invent a single grand schema that everything must fit. Its purpose is to define a small set of shared, reusable interoperability contracts, and to explain how each dataset can meet the contract that fits it.

All three demonstrators study offshore wind farms. The Norwegian case assesses the impact of an offshore wind farm on the community and environment of Utsira Island. The Swedish case explores how an offshore wind farm affects the fishing community and economy on the Baltic coast north of Gotland. The German case looks at a multi-use offshore wind farm that also hosts low-trophic aquaculture.

1.2 Who should read this, and what you will get from it

This document is written for two kinds of readers at once, and it is arranged so that each can take from it what they need.

The first is the strategic reader: the policy lead, the agency director, the program manager, or the planner who is shaping how an organization, a nation, or a community will share and use data in the years ahead. Readers who are charting the future of spatial data infrastructures and the emerging geospatial ecosystem will find here a clear account of where interoperability is heading and why, expressed in plain language and without requiring any prior familiarity with the underlying standards. For this reader, the opening chapters are the heart of the document. They explain the problem that makes data hard to combine, the idea of the building block as the unit that solves it, and, through two concrete worked examples, why separating structure, meaning, and provenance is the design choice that pays off. A strategic reader can stop there and still come away with a firm grasp of the approach and what it implies for the future of spatial data infrastructures.

The second is the practitioner: the data manager, the architect, the standards specialist, or the developer who has to put the approach into practice. For this reader, the later chapters and the appendices carry the detail: the canonical formats, the controlled vocabularies, the catalog and provenance model, the governed check-in workflow, and a full worked walk-through of a real building block, backed by the standards stack, the building-block index, and the vocabulary, namespace, and tool registries.

No prior knowledge of building blocks, of any particular domain, or of the example project used throughout is assumed. A reader who works through the document should come away understanding what a building block is, what problem it solves, how the pieces fit together, and how to begin applying the approach to their own data. The examples happen to come from digital twins of the ocean, because that setting has to combine unusually diverse data and because its work is public, but nothing in the approach is specific to that domain.

The audience is therefore deliberately wide. If you produce, steward, or publish data of any kind, and you want it to be found, understood, trusted, and reused by people, systems, and AI agents beyond your own team, this document is written for you. You do not need to work on digital twins or data spaces to benefit from it. Those are simply two settings where the need is especially visible, and the same building-block approach applies equally to a national mapping agency, a statistical office, a research group, or a single team publishing one dataset well.

1.3 Scope of this version

It is worth being explicit about what this version of the framework covers, so that no reader forms the wrong expectation.

This version shows how building blocks apply to three things. First, the inputs to a digital twin, meaning the observations, forcing data, parameters, and reference datasets a model consumes. Second, the outputs, meaning the products a model run generates. Third, the lineage, meaning the record that connects an output back to the model, the parameters, and the input data that produced it.

Figure 2: In scope (left): inputs, outputs, and lineage; out of scope is the full digital twin, with lots of models interacting with each other and the physical environment

This version does not yet demonstrate building blocks for the twin as a living system. A digital twin, in its fullest sense, is a running, coupled set of models that stays connected to the real system it represents and feeds information back and forth. The framework does not standardize how the models run, how they are coupled, or how they exchange state at runtime. What it standardizes is the data and provenance envelope around each model run.

This boundary is deliberate and appropriate for the current stage of the project. Getting the inputs, outputs, and lineage right is the foundation on which the running twin can later be described. The framework is a living document. A future version is expected to apply the same building block approach to the running twin itself, and Section 12 records this as a planned direction rather than a gap to be apologized for.

1.4 What the framework is, and what it is not

The framework sits between the raw demonstrators and the target platforms. It captures the principles and workflows that make the project’s data reusable. It covers data inputs and outputs, the processes and workflows that matter for reuse, the vocabularies of model properties and indicators, and the relationships between indicators.

The framework defines interfaces, not operational pipelines. It says how data should be shaped and described so that systems can exchange it, but it does not itself run the connections between those systems. For example, Norwegian biodiversity observations are already published into the global OBIS2 system. If a scenario needs richer detail than OBIS carries, the framework defines the contract for acquiring and transforming that richer source, and it captures the acquisition and harmonization steps as reusable tools. The project cannot guarantee that live data will keep streaming into EDITO after the project ends, but the tools that put data into the required shape will remain available.

2 The problem: why data is hard to combine

The same difficulty appears in almost every field. Suppose you are building a decision-support tool that has to bring together several datasets produced independently by different groups. Consider, for example, a tool for planning an offshore wind farm. It might need ocean temperature and salinity from a numerical model, catch statistics from a fisheries agency, seabed biodiversity from a national survey, employment figures from a statistics office, and the results of a simulation that ties these together. Whatever the domain, each of these arrives in its own form.

The model output is a multidimensional grid with dimensions, variables, and a coordinate system. The biodiversity data needs taxonomic identifiers to say which species were observed. The statistics come with aggregation dimensions and code lists. The social survey follows a complex internal data model with its own vocabulary. The simulation carries provenance describing how it was run. Each of these is reasonable on its own, and each reflects how its own community actually works.

The tempting solution is to define one common schema that everything is converted into. In practice, this fails in one of two ways. Either the single schema forces each dataset into a shape it does not naturally have, which loses information the source community depends on, or the schema grows into a giant superset in which most fields are meaningless for any given dataset, which no one can validate in a meaningful way. Section 5 shows both failure modes with concrete examples.

This framework takes the other path. It keeps multiple data profiles, one family per kind of asset, and it maintains interoperability through shared discovery and provenance mechanisms rather than through a single schema. The rest of this document explains the pieces of that approach.

3 The core idea: building blocks and four mechanisms

3.1 What a building block is

An OGC building block is a small, reusable, self-contained package that captures one interoperability contract3 completely. A single building block can contain a schema that defines the JSON structure of some data, a context that maps each field to a precise meaning on the web of linked data, a set of example records, validation rules that check whether a dataset actually complies, and, where relevant, transformation rules for converting a source into the required shape. It also carries human-readable documentation.

The value of packaging all of this together is reuse. A building block can be published once and then referenced by many datasets and by other building blocks. When one part of the design changes, only the relevant building block changes, and everything that depends on it can be revalidated automatically. A building block is versioned, so improvements can be made without silently breaking the datasets that rely on earlier behavior. In short, a building block turns an interoperability agreement into a testable software artifact rather than a paragraph in a specification that everyone interprets slightly differently.

A project rarely needs to invent all of its building blocks from scratch. Many are already published by OGC, by the ILIAD digital twin of the ocean project, and by the OGC Open Science work, and a project defines its own only where a genuine gap exists. Blocks are not only reused but composed. A block can aggregate others, refine or extend one to cover a case it was not written for, or substitute one for another, and these are the same formal operations used across the wider building-block ecosystem.

3.2 Four mechanisms, four questions

The framework achieves interoperability through four complementary mechanisms. Each answers a different question, and keeping them separate is what makes the whole approach flexible.

Data profiles define how information is structured. They answer the question: how is this data shaped and in what format is it exchanged? Catalog profiles define how an asset is described, found, linked, and reused. They answer the question: how is this asset discovered? Controlled vocabularies define what the terms in the data and metadata actually mean. They answer the question: how is this information interpreted consistently by different people and systems? Building blocks provide the reusable implementation artifacts, including the schemas, examples, validation rules, and transformations, that put the other three into practice. They answer the question: how is interoperability implemented and validated?

Figure 1. For each digital asset, the framework defines how the data is structured (data profile) and how it is discovered (catalog profile). The terms used are defined in vocabularies, and building blocks bind all of this together and make it testable.

Figure 3: For each digital asset, the framework defines how the data is structured (data profile) and how it is discovered (catalog profile). The terms used are defined in vocabularies, and building blocks bind all of this together and make it testable.

A single dataset can take part in all four at once. It can conform to a data profile, be described through a catalog profile, reuse terms from several vocabularies, and be implemented using one or more building blocks. Because the four are separate, a change in storage format, in metadata model, or in vocabulary can be handled on its own without disturbing the others.

3.3 The interoperability layers

The same idea can be seen as a set of layers, from the business question at the top to the technical protocols at the bottom. Each layer has its own concern and its own SeaDOTs mechanism.

Layer Concern SeaDOTs mechanism
Business Policy questions, management scenarios, indicators Demonstrator methodologies
Semantic Meaning of concepts and variables Controlled vocabularies
Structural Data and metadata structures Data and catalog profiles
Technical APIs, storage and exchange protocols OGC APIs, STAC, DCAT, GeoZarr

The framework does not try to standardize all project data into one schema. It defines reusable interoperability contracts suited to specific classes of assets. Data profiles describe structure, vocabularies describe meaning, catalog profiles describe discovery, and building blocks bind them together and can also describe implementation patterns such as transformations. Together they provide interoperability across very different SeaDOTs’ assets.

4 How data flows through the framework

Before looking at the mechanisms in detail, it helps to see the whole path a piece of data takes. Raw data enters the framework, is profiled and harmonized, is described in the catalog, is validated against the relevant building blocks, and is finally published to a target platform where it can be discovered and reused. The sections that follow make this concrete with the example of a digital twin run, but the same path applies to any data that has to be prepared for sharing and reuse.

Figure 2. The high-level flow of digital assets. Raw demonstrator data is ingested, profiled, harmonized, cataloged, validated and published to EDITO, with profiles that can also be served to EOSC.

Figure 4: The high-level flow of digital assets. Raw data is ingested, profiled, harmonized, cataloged, validated, and published to EDITO, with profiles that can also be served to EOSC.

4.1 What a digital twin execution produces

Each demonstrator runs a different kind of model. The German case couples a numerical ocean model with aquaculture yield in a multi-use scenario. The Swedish case uses agent-based models of fisheries on the Swedish coast. The Norwegian case maps fuzzy-cognitive relationships for wind-farm investment decisions. Despite the differences, every run can be understood in the same simple way. A simulator takes a set of input data and a set of control parameters, and it produces a set of output variables.

Figure 3. A digital twin run seen as a function. For a given demonstrator, the simulator turns input data and control parameters into output variables such as predicted yield, employment or catch.

Figure 5: A digital twin run seen as a function. For a given demonstrator, the simulator (S) turns input data (I) and control parameters (C) into output variables (O) such as predicted yield, employment, or catch.

The outputs, such as a predicted yield time series, local employment, tons of fish caught, or a biodiversity measure, are the variables that feed the project’s indicators. From many runs, the project derives generalized or equilibrium relationships that support decision-making. This document focuses on the interpretable data at the level of a single run, meaning its inputs, its outputs, and the way they are linked.

4.2 Linking inputs, outputs and the model

The connection between an input, an output, and the model that ties them together is the execution record. One execution is one run of one workflow. It records which model was used, which inputs and control parameters went in, and which products came out. This record is what makes a result reproducible and auditable, because a reader can trace any product back to the exact run and inputs that produced it.

Figure 4. The digital assets of one digital twin run. The execution record executes the workflow and records the input data and control parameters it used and the products it produced. The products then feed the generalized indicators model.

Figure 6: The digital assets of one digital twin run. An experiment executes a workflow, uses input and configuration data, and produces products that feed the indicators model.

5 Two worked examples: why separate mechanisms beat one schema

The separation of vocabulary, schema, and building block is not just tidy in theory. Two real examples show concretely why collapsing them into a single artifact would fail. These two examples are the heart of the framework, and a reader who understands them understands the whole approach.

5.1 Example one: the same real-world quantity, two structures, one meaning

The Norwegian reef-effect calculation needs a value for benthic biomass4 density for each affected species. Two independent monitoring programs can supply this value, MAREANO and IMR, but they cannot share a single structure.

The MAREANO program reports values per species inline, as an array over a mapped area and a sampling period, because that is how a MAREANO survey record is built. The IMR program reports a single density value with an explicit uncertainty and an ICES5 area annotation, because that is how the IMR regional baseline series works. A single shared schema would either force one program’s data into a shape it does not have, or become a bloated superset in which most fields are meaningless for any given record. This is exactly the failure mode described in Section 2.

What makes the two sources interchangeable as inputs to the reef-effect calculation is not a shared structure. It is the vocabulary layer. Both programs are bound to one abstract concept, baseline benthic biomass density, defined once. Each program provides a narrower, more specific version of that same concept. Because each program’s fields are mapped to the same abstract meaning, the reef-effect calculation can consume either source without needing to know which program supplied the number.

Figure 7: The same quantity supplied by two programs. A single abstract concept in the vocabulary layer connects two differently structured schemas, so a consumer can use either as the same kind of input.

Three separate building blocks make this arrangement reusable and independently maintainable. One is a vocabulary building block that defines the shared abstract concept, baseline benthic biomass density, together with its two narrower bindings; i.e., the MAREANO concept and the IMR concept are both narrower kinds of that abstract concept. The other two are schema building blocks, one for MAREANO and one for IMR. Each carries its program's data structure, the shape of a MAREANO survey record or an IMR baseline record, together with a small mapping, called the context, that says which concept each field corresponds to. This field-to-concept mapping lives inside the schema building block itself, not in the vocabulary building block.

This division is exactly what buys the independence. MAREANO's schema block can change its survey fields without touching IMR's, because IMR's fields are described entirely inside IMR's own block. A new monitoring program simply adds one more schema building block, which carries that program's field mapping, and one more narrower binding in the vocabulary block, and none of the existing blocks change. The reef-effect calculation, which only ever refers to the shared abstract concept, does not change either. That independence would be impossible if benthic biomass density were locked into one hard-coded schema. The two schemas do share a common envelope, the general observation shape, but their detailed payloads cannot be merged without either dropping IMR's mandatory uncertainty or inventing an empty field on IMR records that was never captured.

5.2 Example two: one structure, two unrelated meanings

Example 1 showed two data sources that were built differently but meant the same thing. Example 2 shows the mirror image. Here, two datasets are built the same way, yet they mean completely different things. Reading the two examples side by side is the quickest way to see why structure and meaning have to be handled as separate concerns.

The two datasets come from different demonstrators and different scientific worlds. One is the output of the Swedish digital twin, a simulation of the herring and sprat fishery. The other is a German harvest scenario, a time series of modeled harvest values. In terms of content, they have nothing to do with each other.

What they do have in common is how their files are physically laid out. Both are stored as GeoParquet, and both describe that layout in the same way, through a single shared building block called the GeoParquet header envelope. This envelope is purely structural. It records things like the file name, where the data came from, the list of columns and their data types, and how the geometry and its coordinate reference system are declared. It says nothing about what any particular column actually means. Because both datasets point to this one shared building block instead of each defining its own layout, any standard GeoParquet tool can open either file without special handling.

Figure 8: The same structural envelope, two unrelated meanings. Two datasets reuse one shared GeoParquet header schema, while each keeps its own vocabulary mapping.

The meaning of the two datasets lives somewhere else entirely, inside each dataset's own vocabulary. The Swedish output uses the SeaDOTs fishery indicator terms. The German scenario uses its own harvest terms. A column in one dataset has no equivalent in the other, and nothing in the shared envelope pretends otherwise. In short, the structure is shared while the meaning is kept apart.

This separation is not just tidy; it pays off in practice. At one point, the shared envelope had a defect in the way it recorded the coordinate reference system, which could make a file unreadable by real tools. Because the rule that guards against this lives in the shared building block, fixing it once corrected both datasets at the same time. Neither dataset's own vocabulary had to change, and the repair for one did not disturb the other.

Taken together, the two examples show that structure and meaning are independent of each other. Two datasets can share a meaning without sharing a structure, as in Example 1, and they can share a structure without sharing a meaning, as here. This is exactly why the framework keeps schemas, vocabularies, and building blocks as separate, reusable pieces. A single common schema would fuse structure and meaning into one, and it would break the moment a new dataset needed one of them but not the other.

5.3 The general pattern

The two examples point to a single lesson, which can be summarized as follows.

If you only had You would lose
One schema for benthic biomass density Either MAREANO’s or IMR’s native structure, or a bloated superset that nobody validates meaningfully
One vocabulary term instead of broader and narrower bindings The ability to tell a downstream consumer that the MAREANO and IMR values are the same kind of thing
One building block that folds a dataset's meaning into the shared file structure Reuse of that structure for the next dataset, because it would now carry the first dataset's meaning; a single fix such as the coordinate reference system correction would then have to be repeated in every dataset instead of made once
One building block that assumes a shared file structure also means shared content The boundary between structure and meaning, so that two unrelated datasets sharing only a file layout would wrongly appear comparable in content

Figure 9: Possible combinations of meaning and structure

The two questions are independent. Whether two datasets share a file structure and whether they share a meaning can be answered separately, which gives four combinations. Example 1 sits in the top-left, where the meaning is shared but the structure differs. Example 2 sits in the bottom-right, where the structure is shared but the meaning differs. Building blocks let a project reuse whichever axis is shared: a vocabulary building block for meaning and a schema building block for structure, which a single common schema could not do.

5.4 The payoff: full provenance and easy reuse

Keeping schemas, vocabularies and building blocks as separate, reusable pieces takes more care up front than writing a single schema. As a close to this chapter, it is worth stating plainly why that care pays off. The benefits are easiest to see from two perspectives: what building blocks give you when you describe data with them, and what they give you when a new dataset needs to reuse them.

The first perspective is description and provenance. When a dataset is described through building blocks and registered in the way this chapter and the check-in workflow in Section 9 set out, the result is a complete provenance profile, not just a file. The record captures which source the data came from, which building blocks governed its description, which transformation turned the source into its compliant form, which fields were mapped to which vocabulary terms, and what was finally produced. Because all of this is written down against the building blocks that governed it, anyone can later trace exactly what was done at each step, how it was done, and on what basis. A building block is therefore more than a schema. Together with the catalog and provenance records, it makes the whole description auditable and reproducible.

The second perspective is reuse when a new dataset arrives. If the building blocks that other groups have already built are public, they serve as worked examples that show how to describe your own data, and the two examples in this chapter map onto the two situations you are most likely to meet.

In the first situation, your dataset measures the same real-world quantity as an existing one. You can then declare that your data also maps to the shared abstract concept, and you can take the existing building blocks as a template. Seeing how others mapped their fields to that concept is the quickest way to work out how to map your own data to it, rather than modeling it from scratch.

In the second situation, your dataset shares a structure rather than a meaning. If it is another GeoParquet output, you reuse the same GeoParquet header building block, which already describes the complete form of your file. You then look at how other groups described their content in their own blocks, and you define an analogous block for your content, reusing their approach as a pattern.

In both situations, the two perspectives come together. The moment your new building block is used, the provenance benefit from the first perspective applies to your data as well, so that everyone who later encounters it can again trace exactly what was done, when, and how. Reuse lowers the effort of describing a new dataset, and it does so without giving up any of the transparency. That combination is what a single common schema cannot offer, and it is the reason the framework is built from building blocks in the first place.

6 Data profiles and canonical formats

SeaDOTs data falls into a few broad categories. Each has a preferred storage format for efficiency and a canonical exchange format for interoperability. For every category, the project registers a catalog record, based on the OGC API Records model and profiled for STAC for data and for DCAT-AP for open-science discovery.

Because the categories are genuinely different in structure, a single profile would either lose detail or become unworkable. Multidimensional model outputs need dimensions, variables, and coordinate systems. Biodiversity observations need taxonomic information. Statistical datasets need aggregation dimensions and code lists. Social survey data follows complex internal models. Workflows need provenance and execution metadata. The project therefore uses several profile families while keeping common discovery mechanisms through STAC, OGC API Records and DCAT.

6.1 Multidimensional spatio-temporal data

This category covers gridded and array data such as model outputs and reanalysis, where variable values are mapped onto a predefined spatial and temporal scheme. NetCDF with the Climate and Forecast conventions is the common legacy and source format, and it is well understood. The preferred cloud-native storage format on EDITO is GeoZarr, which is chunked for efficient access and stores its metadata as JSON, which makes extensions easier.

Concern Format
Preferred storage GeoZarr, cloud-native and chunked
Legacy or source NetCDF with CF conventions and NCEI templates (National Centers for Environmental Information)
Served via S3 bucket, ERDDAP, THREDDS, OGC WMS, WCS, and EDR
Canonical exchange CoverageJSON, an OGC standard
Metadata STAC datacube extension and CF global attributes
Vocabularies NERC, EMODnet, SeaDataNet and CF standard names

Typical examples include Baltic Sea physics from Copernicus Marine, Utsira wind measurements (see Utsira Nord Offshore Wind Area and Utsira Nord Metocean and Wind Dataset), and DWD reanalysis. The catalog profile for this category is named catalog-data-multidim.

6.2 Tabular and statistical data

Tabular data usually arrives as CSV or plain text, sometimes following a column convention such as SeaDataNet. SeaDOTs prefers semantically rich formats in which the meaning of each column is explicit and linkable, such as GeoParquet, GeoJSON, and CSV on the Web. EDITO supports GeoParquet and CSV, so these are the preferred targets.

Concern Format
Preferred storage GeoParquet, cloud-native and spatially indexed
Source formats CSV, SDMX-ML and SDMX-JSON
Canonical exchange GeoJSON Feature and FeatureCollection
Metadata STAC Item, DCAT 3.0 Dataset, and OGC API Records
Vocabularies SDMX code lists such as NUTS and ICES codes, and Darwin Core for species

Typical examples include Eurostat fishery and tourism statistics, ICES catch data, HELCOM marine protected areas, Utsira observation tables and social survey data. The catalog profile for this category is named catalog-data-tabular. Dataset-specific GeoParquet profiles should reference the shared GeoParquet header building block rather than redefining the envelope, as explained in Section 5.2.

A note of practical caution is worth recording. The coordinate reference system (CRS) stored in a GeoParquet file should be encoded as a valid and complete PROJJSON object, as defined by the GeoParquet and PROJ specifications, rather than as a shorthand representation that merely resembles JSON. Using incomplete CRS definitions can lead to interoperability issues and inconsistent coordinate interpretation across software environments. The shared header building block previously discussed now constrains this, and the conversion tools emit correct values, but a produced file should always be sanity-checked by opening it with a real reader before it is trusted.

6.3 Graph and knowledge data

Some artifacts, such as the fuzzy-cognitive models that express relationships between variables and indicators, are naturally graph data. When they arrive as plain text or CSV, their rich version is converted into canonical RDF.

Concern Format
Storage RDF Turtle and RDF-star
Served via Apache Fuseki SPARQL endpoint
Canonical exchange JSON-LD and Turtle
Metadata DCAT 3.0, PROV-O provenance, Ocean Information Model and GeoSPARQL
Vocabularies Ostrom social-ecological system ontology, PROV-O and SKOS

Typical examples include the Ostrom social-ecological system model for offshore wind governance, composite indicators and Norwegian relationship matrices. The catalog profile for this category is named catalog-data-graph.

6.4 Workflows and applications

In an open-science setting, data is linked to the workflow that generated it and to the input data and control parameters that fed it. The connecting point is the experiment, meaning one execution. SeaDOTs represents these relationships at the catalog level. The project does not standardize how models are executed, although some run on EDITO and comply with its requirements.

From an interoperability and reuse point of view, two aspects matter for a workflow. First, its metadata, which describes data requirements, owner, license, and provenance. Second, its technology stack, which determines whether the application can run on EDITO or only elsewhere. These assets are represented at the catalog level with references to the code base and to any published papers.

6.5 From Standards to Building Blocks

This chapter makes the scale of the problem visible. Running a digital twin means agreeing on storage formats, exchange formats, metadata models and controlled vocabularies, and doing so separately for gridded data, tabular data, graph data and workflows. On their own, these standards are only a menu of possibilities. Each one can be read, interpreted and combined in more than one way, and two teams working from the same list can still end up with data that does not fit together.

This is where building blocks earn their place. For each category, a building block fixes one specific, tested combination of these standards and writes it down as a concrete artifact. It records, for instance, that multidimensional data is stored as GeoZarr, described with the STAC datacube extension and CF attributes, and named with CF standard names from the NERC server, and it captures that decision as a schema, a set of examples and validation rules rather than as prose that each reader interprets anew.

In this way building blocks turn the long list of standards in this chapter from something every project has to reassemble by hand into something that can be referenced, validated, reused and versioned. They are the layer that makes the standards operational, so that the many agreements a digital twin depends on are applied the same way every time and can be checked automatically rather than taken on trust.

7 Controlled vocabularies

7.1 Reusing authoritative vocabularies

Profiles ensure that data is structured consistently. Vocabularies ensure that the terms in the data mean the same thing to everyone. Because SeaDOTs combines environmental, ecological, social and economic data, several domain vocabularies are needed. In each case, the project reuses an authoritative external vocabulary wherever one exists.

Environmental variables such as temperature, salinity, wave height, and wind speed are described with the CF standard names, published through the NERC Vocabulary Server, and harmonized with SeaDataNet vocabularies. Biodiversity and ecosystem data uses Darwin Core for observation terms, WoRMS6 for taxonomic identifiers, and OBIS7 for exchange. Statistical and socio-economic data uses SDMX for statistical structures, Eurostat code lists for regional and economic classifications, and ICES code lists for fisheries reporting. Provenance and workflow metadata uses PROV-O, DCAT and the OGC Open Science building blocks.

Some social-ecological concepts that SeaDOTs needs are not covered by any existing vocabulary. Examples include social acceptance indicators, fisheries displacement indicators and socio-ecological relationship models. These are maintained in project-specific vocabularies and indicator registries, with mappings to external standards kept wherever possible, so that the project does not create isolated terms with no connection to the wider web of data.

7.2 How building blocks attach meaning to the data

As with formats, vocabularies deliver their value only once they are attached to real data, and that is the function building blocks perform here. A vocabulary on its own is a shared list of agreed terms. It says what temperature, a taxon identifier or a statistical region means, but it does not by itself say which field of which dataset carries that meaning. The building block is where the connection is made. Inside each building block a small mapping, the JSON-LD context introduced in Section 5.1, binds the dataset's own fields to the authoritative terms described in this chapter, so that a column named in a project's private way is tied unambiguously to a CF standard name, a Darwin Core term or an SDMX code. Because this mapping lives inside a building block alongside examples and validation rules, the meaning is not merely asserted in prose but can be reused by the next dataset, checked automatically, and read back as linked data once the dataset is published. The same building block also carries the honest exceptions. Where SeaDOTs needs a concept that no external vocabulary yet covers, the building block records it as a candidate definition with links to the closest existing terms, rather than inventing an isolated label. In short, the vocabularies in this chapter supply the meaning, and the building blocks are what fasten that meaning to the data in a form that can be reused and verified.

One further point completes the picture, and it matters as much as the mappings themselves. Attaching meaning is not only about declaring that two fields mean the same thing. It is sometimes about declaring, explicitly, that two things which look alike are not the same and must not be combined. A building block can therefore record a declared non-correspondence, for example that a surface skin temperature and an air temperature two meters above the ground are different quantities despite sharing a label, so that a consumer, and an automated agent in particular, is told to keep them apart rather than left to average them by accident. A stated incompatibility is worth as much as a clean mapping, because a machine can act safely on it, whereas an unstated one is exactly the trap that produces confident, wrong answers. The companion article "Declared Meaning" develops this point in full.

8 Catalog profiles and provenance

This chapter builds up in four steps, from the shape of the catalog to a concrete specification. Section 8.1 shows how the catalog is organized, by demonstrator and by the role each record plays. Section 8.2 introduces the basic elements of a single run and how they relate, in plain terms. Section 8.3 shows how those elements are encoded in real standards, what each link means, and where the activity fits. Section 8.4 is the level below that, the concrete profile for each record type, namely which standard it is built on, which fields a compliant record must carry, and how it is published for EOSC. Taken together, the earlier sections explain the model and the later ones turn it into the checklist that the check-in workflow in Section 9 enforces, with the DCAT and GeoDCAT-AP serialization as the bridge that makes SeaDOTs records discoverable across the wider open-science landscape.

8.1 Catalog hierarchy

Data profiles describe the content of an asset. Catalog profiles describe the asset itself, so that it can be discovered, its provenance tracked, and its records exchanged between EDITO, EOSC, and external catalogs. The catalog is STAC-first for data, because that is what EDITO expects, and it is aligned to OGC API Records for applications, following the structure of the OGC Open Science building blocks. For EOSC, these records are linked to the DCAT and GeoDCAT vocabularies.

Catalog records can represent data, an application or workflow, and an execution. Within the SeaDOTs catalog, these are organized per demonstrator.

/Catalogs/SeaDOTs/
├── germany/
│   ├── workflows/{workflow_id}      OGC Records item
│   ├── experiments/{experiment_id}  STAC item
│   ├── inputs/{input_id}            STAC item, per type
│   └── products/{output_id}         STAC item, per type
├── norway/
│   └── ...
└── sweden/
    └── ...

The tree is read from the top down. The catalog is divided first by demonstrator, with a branch for Germany, Norway, and Sweden, and then by the role each record plays inside that branch. The same four roles appear everywhere. Workflows describe the reusable models and processing steps and are registered as OGC API Records items. Experiments record the individual runs of those workflows. Inputs and products hold the concrete datasets that a run consumed and produced, and both are registered as STAC items so that EDITO and other spatial catalogs can index them directly.

The choice of standard follows the kind of thing being described: a workflow is in effect a piece of software, which OGC API Records captures well, whereas an input or a product is a spatial dataset, which STAC is built for. Because the same roles recur in the same order under every demonstrator, a reader or a tool can move through the catalog predictably, from a country to a particular run, and from that run outward to the data it used and the data it produced.

8.2 The catalog metadata model

The catalog represents each digital twin run as a small set of linked records rather than as one large document. Before looking at the technical detail, it helps to see the basic shape, which the simplified diagram below captures.

Figure 10: The four basic elements of a catalog entry and how they relate. An execution is one run of a workflow; it uses its inputs and generates its outputs, and each output is derived from the inputs it came from.

There are four elements. A workflow is the reusable model, transformer, or processing service. An execution is one concrete run of that workflow. An input is a dataset, parameter file or configuration that the run consumes, and an output is a product that the run produces.

They relate in only a few ways. An execution is an instance of a workflow, so the workflow acts as the plan that the run follows. When the run happens, it uses its inputs and generates its outputs, and each output is linked back to the inputs it was derived from. That handful of elements and links is the entire model. The next section shows how these same records and links are written down in real standards, and it introduces one further element, the activity, that belongs to that more detailed level.

8.3 The provenance model

Section 8.2 introduced the elements and how they relate. This section shows how each element is expressed in a concrete standard, what each record actually carries, and how the whole pattern travels between platforms. The diagram below is the same model as before, now drawn in full detail that includes the new element Activity.

Figure 7. How catalog entries relate. An execution instantiates a workflow, uses input data and generates output data, all discoverable as catalog records.

Figure 11: How catalog entries relate. An execution instantiates a workflow, uses input data and generates output data, all discoverable as catalog records.

Read the diagram by its three groupings, because each one gathers the records that come from a single standard. The execution and the workflow are OGC API Records items, so each carries record fields such as id, type, conformsTo, a title and a description; the execution adds its start and end time, while the workflow adds its application category, method and version. The activity sits in the OGC Open Science and PROV-O group and carries the plain provenance fields. The input and the output are STAC items, so each carries the spatial fields a catalog needs, a geometry, a bounding box, a datetime and its assets.

What makes this more than a picture is that every link carries two things at once, a plain meaning and the exact term used to encode it. The line from the execution to the workflow is the plan relation, written prov:hadPlan. The line to the input is the used relation, written prov:used and, in STAC, as a link with relation input. The line to the output is the generated relation, written prov:generated and as a link with relation output. The line from the output back to the input is prov:wasDerivedFrom. Because each relation has both a human label and a machine term, the same record can be read by a person and resolved by a tool.

One element in this diagram did not appear in the simpler picture, the activity, and it is worth explaining on its own, because it sits very close to the execution. Recall that the workflow is the reusable plan, and the execution is one concrete run of that plan, the event that happened at a particular time with particular inputs and outputs. The activity is the operation that the plan prescribes, in other words what is actually done, and in the diagram it hangs off the workflow rather than off the execution. A simple analogy keeps the two apart. The workflow is a recipe. The activity is a step the recipe prescribes, such as bake at 200 degrees, which is what is to be done. The execution is the concrete cooking on Sunday, with these ingredients and this result. So the execution answers when and with what a run took place, while the activity answers which operation is carried out. The two are kept separate so that the catalog-facing run record, the execution, stays distinct from the operation concept, the activity, which SeaDOTs reuses directly from PROV-O rather than defining its own, and which lets a machine reason about what a workflow does each time it runs.

The main entities are the following. A workflow is the catalog record for a reusable digital twin application, model or processing service, and it also serves as the plan that executions instantiate. An application package is the optional executable profile of a workflow, capturing its computational contract, including declared inputs and outputs, runtime requirements and software version. An execution is the record of one concrete run, with its specific inputs, parameters, time boundaries and output products. An activity is the operation that a run performs, reused from PROV-O rather than redefined locally. An input is a concrete dataset, parameter file or configuration consumed by a run. An output is a concrete product generated by a run, linked back to the inputs it derives from and to the execution that produced it.

Finally, the same pattern can be written in more than one serialization at once, which is what lets it move between platforms. It appears as OGC API Records for discovery, as application-package metadata for executable packages, as PROV-O relations for the execution graph, as STAC items for the datasets, and as DCAT and GeoDCAT-AP for harvesting into EOSC.

8.4 Metadata profiles

Sections 8.2 and 8.3 described the elements of a run and how they are encoded. This section turns that into a concrete specification. For each of the four record types, it states the standard the record is built on, the fields a compliant record is expected to carry, and the form the record takes when it is published for the European Open Science Cloud (EOSC). In short, 8.2 and 8.3 explain the model, and this table is the checklist you use when you actually build or check one of these records. It is also the specification that the validation step of the check-in workflow in Section 9 enforces.

Four profile types cover the different asset classes, each built on a standard basis and each with a path to EOSC through a DCAT serialization.

Profile Standard basis Key fields EOSC / DCAT serialization
Workflow OGC API Records, PROV-O, application-package link type, applicationCategory, method, version, inputs, outputs, applicationPackage DCAT 3.0 data service or software record with plan semantics
Execution OGC API Records, PROV-O startTime, endTime, workflow, used inputs, generated outputs, parameter log DCAT 3.0 dataset or provenance-qualified record
Input data STAC 1.1, OGC API Records datacube, EO, table and temporal extension fields, assets GeoDCAT-AP 3.0
Output data STAC 1.1, PROV-O derived-from links, assets, quality metrics, temporal extent GeoDCAT-AP 3.0

The table has one row per record type, matching the four elements from Section 8.2, and three things are worth reading out of it.

First, the standard basis column shows that each record is grounded in the standard that suits its nature, the same principle seen in the catalog hierarchy. Workflow and execution records are OGC API Records items enriched with PROV-O, because they describe software and the history of a run. Input and output records are STAC items, because they describe spatial datasets.

Second, the key fields column is the practical part. It names what a reader, or a validator, should expect to find in each record. A workflow carries its type, application category, method, version, its declared inputs and outputs, and a link to its application package. An execution carries its start and end time, the workflow it ran, the inputs it used, the outputs it generated, and a log of its parameters. An input carries the STAC extensions that fit its shape, the datacube, EO or table extension, together with its temporal extent and its assets. An output additionally carries its derived-from links, quality metrics, and temporal extent. These fields are what lets a later user discover the record, judge its quality, and trace it back to its origin.

Third, the last column is the one genuinely new point in this section, the bridge to EOSC. A record that lives in the SeaDOTs catalog is also serialized as DCAT so that EOSC and other European data portals can harvest it. The flavor of DCAT follows the kind of record. Workflow and execution records become plain DCAT 3.0 entries, a data service or software record for the workflow and a provenance-qualified record for the execution, because that is how EOSC represents software and process history. Input and output records become GeoDCAT-AP 3.0, the geospatial profile of DCAT, because they are spatial datasets. This second serialization does not replace the STAC and OGC API Records forms. It sits alongside them, so the same record is discoverable both inside EDITO and across the wider open-science landscape.

8.5 Building blocks behind the catalog and provenance

As in the previous chapters, it is worth ending by placing the building blocks in the picture, because everything in this chapter is delivered as building blocks rather than as prose. The catalog hierarchy, the elements of a run, the way they are encoded, and the field-level profiles are all defined by building blocks: catalog-workflow, catalog-execution, and the generic catalog-data block, together with its specializations for multidimensional and tabular data. These are the concrete, versioned artifacts that pin down each record type, so that the model described in this chapter is something a machine can produce and check, not only something a person can read.

It helps to be precise about what one of these building blocks actually is. A building block is not itself a STAC item, nor an OGC API Records record, nor a DCAT entry. It is a single definition. It consists of a JSON schema, which fixes the structure and the fields a record must have, and a JSON-LD context, which is a small mapping that states what each field means. Around these, it also gathers examples and validation rules. Everything else in this section follows from one fact: these several familiar representations are not written separately, but are all obtained from that one definition.

They are obtained on two levels. The first is the level of the plain JSON structure. STAC and OGC API Records are two closely aligned ways of writing a catalog record as JSON, and the building block's schema is written so that a single record satisfies both at once. There are not two documents here, but the same record read through two compatible lenses. So when we say that a record is both a STAC item and an OGC API Records record, we mean literally the same bytes, valid against both specifications.

The second level is the level of linked data, and this is where PROV-O and DCAT come in. The JSON-LD context attached to the record states, field by field, what each value means in terms of shared vocabularies. Applying that context turns the plain JSON into a graph of statements. Because the context maps the record's links to PROV-O terms, for example the used and generated relations from Section 8.3, the same record can be read as a provenance graph. Because it also maps the record's descriptive fields to DCAT and GeoDCAT-AP terms, the same record can be serialized as a DCAT entry for harvesting into EOSC. PROV-O and DCAT are therefore not extra files written by hand. They are what the one record looks like when it is read as linked data through its context.

Figure 12: One definition, several views. A building block is a single definition, from which one catalog record is produced. Read as plain JSON, that record is at once a STAC item and an OGC API Records record. Read as linked data through its JSON-LD context, the same record is a PROV-O provenance graph and a DCAT or GeoDCAT-AP serialization. Because every view derives from the one definition, they cannot drift apart.

The benefit of this arrangement is consistency. Because the STAC view, the OGC API Records view, the PROV-O graph and the DCAT serialization all derive from a single authoritative definition rather than being maintained independently, they cannot drift out of step with one another. Fix or extend the building block once, and every view moves with it. The same definition also carries the field checklist of Section 8.4 as validation rules, so a record can be checked automatically instead of being reviewed by eye. To be accurate about the mechanics, the DCAT serialization is in practice produced by a generation step that reads the record, rather than appearing entirely for free, so single source here means a single source of truth from which the other forms are generated, not one file that is all of them at once.

Seen across the whole document, this is the same mechanism appearing at a new layer. Chapter 6 used building blocks to pin down the shape of the data, and Chapter 7 to pin down its meaning. Here they apply to the records that describe the data and connect the runs to one another. That is what allows the check-in workflow in the next chapter to treat an entire catalog entry, with its structure, its provenance and its EOSC serialization, as a single thing that it can generate, validate and publish.

9 The check-in workflow

The check-in workflow is the governed path for bringing any new data, metadata, model output, workflow or external reference dataset into the framework. Its guiding principle is simple. Data is not registered first and interpreted later. It is checked against a curated building-block contract first, and then registered with explicit meaning, validation evidence, provenance and transformation lineage.

9.1 Governance

Human-curated building blocks are the authority for registration. They define the schemas, contexts, examples, validation rules, accepted vocabularies, provenance requirements and known transformations that a dataset must satisfy before it is treated as an interoperable SeaDOTs asset. As it currently stands, software agents can accelerate profiling, metadata generation, gap detection and the drafting of transformations, but they do not replace the human-curated contracts. Agent output is accepted only when it can be traced to the selected building blocks and passes validation.

The outcome of a check-in is therefore not only a published dataset. It is a small, reviewable interoperability package that includes, where relevant, a registered building block for the original source shape, a selected or extended target building block for the compliant profile, a catalog record linking source, target, workflow, provenance and publication, a tested transformation from source to target, a STAC item or collection for EDITO, and proposed vocabulary entries for any meaning not already covered.

Figure 13: The check-in control plane. The human-curated building blocks at the top are the reference contract. Incoming data passes down through the five steps of the control plane, from intake to strict validation, and a successful check-in produces the four outputs at the bottom, which together form one reviewable interoperability package.

The diagram above reads from top to bottom in three bands, and it is worth following it in that order. At the top is the authority. The human-curated building blocks, the SeaDOTs blocks together with the imported ILIAD and OGC ones, are the reference contract. This is the principle stated at the start of the chapter made visible: nothing enters the framework on its own terms. Everything is checked against these curated blocks first. In the diagram, an arrow runs from the building blocks down into the control plane, labeled "is the reference for", and it makes this visible: every step below takes these blocks as its reference and measures its results against them.

The middle band is the control plane itself, the governed sequence a dataset passes through, in five steps. In step 1, intake and profiling, the incoming data is inspected to determine its format, structure, spatial and temporal coverage, and whatever metadata it already carries. In step 2, building-block matching, the profiled data is compared against the register to find the blocks that fit it, namely a block for the source shape, a block for the compliant target profile, and a block for the catalog record, the three roles that Section 9.3 describes. In step 3, gap and semantic review, the data is measured against those blocks and the accepted vocabularies, and anything missing or unmatched- an absent license, an unmapped term, a missing coordinate system - is flagged rather than hidden. In step 4, transform proposal and generation, a transformation from the source shape into the chosen target profile is built and tested, so that data which is not already compliant can be brought into line. In step 5, strict validation, the result is checked against the schema, the context, the transformation, and the way the data renders through a live service, and only a result that passes, or that carries documented blockers, is allowed through.

The bottom band shows what a successful check-in produces. The single arrow labeled produces makes the point that these are not four separate exercises but one package created together. The package has four parts. The first is the registered building blocks, meaning the source, target, and catalog blocks together with their transforms, examples, and tests. The second is the published data itself on EDITO, as GeoZarr, GeoParquet or RDF in the storage bucket. The third is the catalog records that make it discoverable, as a STAC item or collection with its OGC Records and DCAT or GeoDCAT serializations. The fourth, needed only where a required term did not yet exist, is a proposal for the definition server, carrying the candidate term, its mappings, its provenance and its review status.

One note to avoid confusion. This diagram shows the control plane at a coarse level, as five steps. The next section breaks the same path into finer stages, numbered from intake through registration, so the five steps here and the more detailed stage list there describe the same workflow at two levels of zoom.

9.2 The stages

The workflow moves through a fixed sequence of stages. Each stage has a clear purpose and produces defined artifacts. The following table describes how these stages numbered 1-8 map to the more coarse-grained steps 1-5 in the control-plane diagram above.

Control-plane step (Section 9.1) Detailed stages below
1 Intake and profiling 0 Intake and 1 Profile
2 Building-block matching 2 Match
3 Gap and semantic review 3 Assess gaps and 4 Define semantics
4 Transform proposal and generation 5 Transform
5 Strict validation 6 Validate
The outputs band of the diagram 7 Publish and 8 Register legacy

The table below expands each of these stages in turn. For every stage, it states its specific purpose (what the stage is meant to achieve) and its main artifacts (the concrete outputs it leaves behind). Reading down the artifacts column also shows how the interoperability package from Section 9.1 is assembled piece by piece, because each stage contributes one part of it, from the first acquisition notes at intake to the entries finally committed to the building-block repository.

Stage Purpose Main artifacts
0. Intake Capture source endpoint, files, owner, license, intended use and demonstrator context Source notes and a reproducible acquisition command or script
1. Profile Detect format, schema, dimensions, properties, spatial and temporal coverage, units and available metadata Property inventory, sample records and an initial catalog draft
2. Match Select candidate building blocks from the SeaDOTs and ILIAD registers Source, target and catalog building-block candidates
3. Assess gaps Compare source fields and metadata against the selected blocks and vocabularies Quality assessment, property coverage table and missing-metadata list
4. Define semantics Reuse known vocabularies first, then isolate any uncovered terms as candidate definitions An input package for the OGC Definition Server
5. Transform If the source is not already compliant, generate a tested transformation to the target Transformation code, mapping table and a rationale for any excluded fields
6. Validate Run strict schema, context, semantic, transformation and service checks Validation report, passing and failing examples, and service evidence
7. Publish Write compliant assets to EDITO storage and register the discovery metadata EDITO bucket asset, STAC item or collection, and OGC Records and DCAT serializations
8. Register legacy Commit the reusable source profile, target extension, catalog record, transformation and tests New or updated entries in the project building-block repository

9.3 The three building-block roles

You may have noticed that the stages kept selecting building blocks in three recurring roles. The matching stage looked for a source block, a target block, and a catalog block, and the outputs registered blocks in exactly those three roles. The word role matters here, and it is the key to the whole section. A check-in can involve more building blocks than these three, and the transformation is one we will come to shortly, but three of them do the defining work. This section steps back to explain those three roles, because they are the backbone of every check-in. It is important to note that these three are roles a building block plays inside one check-in. They are not the same as the four roles a description itself can play across the wider architecture, namely establishing that something exists, judging whether it fits and may be used, showing how to combine it correctly, and recording how it came to be. The companion article "Declared Meaning" sets out those description roles; the three roles here are their practical counterpart at the moment a single dataset is brought in.

There are three defining roles because a single check-in has to answer three different questions about one dataset.

  1. What did the data look like when it arrived?

  2. What must it look like to be compliant with SeaDOTs and EDITO?

  3. And how is it described so that others can find it and trust it?

Rather than fold these into one artifact, the framework gives each question its own building block, and each such block plays one of the three roles. This is the same separation of concerns introduced in Chapter 3, now applied to a single incoming dataset.

The source block answers the first question. It preserves the original shape of the data, with its real fields and its real example records, so that nothing about the source is lost or quietly rewritten. It is the faithful record of the data as its own community produced it.

The target block answers the second question. It is the compliant model that SeaDOTs and EDITO expect, and it is usually not new. It is one of the data profiles from Chapter 6, or a catalog profile from Chapter 8, reused here as the destination into which the incoming data must be brought.

The catalog block answers the third question. It is the discovery and provenance record that links the source, the transformation, the target, the publication, and the quality evidence together. It is the catalog-data, catalog-workflow, or catalog-execution profile from Chapter 8, now doing its job for one concrete dataset.

It is worth being explicit about what each of these three blocks actually contains. Each of them is a building block in the sense used throughout this document, so each one bundles a JSON schema, which fixes the structure and the fields, and a JSON-LD context, which maps those fields to vocabulary terms. In other words, a single block describes both the format and the meaning of the record it governs. The vocabulary terms themselves need not live inside the block. The authoritative definitions can sit in a separate vocabulary building block that the source, target, and catalog blocks simply reference, which is exactly the split introduced in Section 5.1 between a concept and the fields that point to it. What differs between the source and the target block is how complete this mapping is. The source block maps its fields to known terms only as far as the incoming data allows and leaves the rest marked as gaps, which is what the gap and semantics stages then work on. The target block, being the compliant model, maps fully onto accepted vocabularies.

This raises a natural question. If matching selects a source block and a target block, and the two usually differ, where is the building block for the transformation between them? The answer is that the transformation is deliberately not a fourth role. The three roles are three descriptions of one dataset: how it arrived, how it must look, and how it is found. A transformation is not another description of the dataset. It is the operation that turns the source shape into the target shape, so it belongs between two of the roles rather than beside them. In the workflow, it is produced at the transform stage as tested transformation code, together with a mapping table and a rationale for any field that is deliberately dropped. It is then committed as part of the check-in package, normally as transformation rules carried with the source block, which is why the diagram in Section 9.1 shows transforms in the outputs band next to the three blocks rather than as a block of their own. When a transformation is general enough to be reused beyond a single dataset, it can be registered as a building block in its own right, and the inventory in Appendix B lists a few such transform blocks. These dataset-specific transforms should not be confused with the general conversion tools, such as the CSV-to-GeoParquet helper, which are reusable tooling rather than the record of one dataset's check-in.

It helps to connect this back to the provenance model of Chapter 8, because the two fit together exactly, and it answers the question a reader is likely to ask: is there, after all, a building block that carries the move from source to target? There is. When that step is packaged, whether as transformation rules carried with the source block or as a standalone transform block, the package plays the role of a plan, in the same sense that a workflow is a plan in Section 8.3. Applying it to a concrete dataset is then an activity, the operation that uses the source as its input and generates the target as its output. So the transformation is a building block; it simply belongs to the operation side of the model, as a plan whose application is an activity, rather than to the descriptive side, where the three roles live. Put in one line, transformation rules are to a single transform run what a workflow is to a single execution.

Across the stages, these three roles are the thread that runs through the workflow. Profiling characterizes the source that the source block preserves; matching proposes candidates for all three; the transform stage connects the source to the target; and the publish and register stages commit them. The strong preference throughout is to reuse an existing block and to extend one only when the gap is real, so that each check-in adds as little new material as possible to the shared registry.

9.4 The validation contract

A check-in is complete only when a clear contract is satisfied. The examples in the source block must come from real data and preserve the source fields. Every source property must be either mapped to the target, intentionally excluded with a stated reason, or listed as an unresolved gap. Every meaningful property must be linked to an accepted vocabulary or reported as a candidate term. The required catalog, provenance, license, temporal, spatial, and asset fields must be present or explicitly flagged. Non-compliant source data must have a documented and tested transformation. The schema, context, validation rules, transformation tests and service rendering checks must pass or have documented blockers. Finally, the published assets, records, source records, transformations and building-block identifiers must all be cross-linked.

One clarification completes the picture: these checks run at two levels. The static level validates the building-block package on its own, its schema, its context, its examples and its tests. The runtime level checks that the published asset actually renders correctly when it is served through a live OGC API, such as Features, Records or Environmental Data Retrieval. A check-in is trusted only when both levels pass, or when any failure is recorded as a documented blocker.

9.5 Handling gaps

Missing meaning and missing metadata are treated as first-class outputs of a check-in, not as errors to be hidden.

Handling a missing term relies on the OGC Definition Server, so it is worth saying briefly what that is. The OGC Definition Server is not a single central website but a kind of service that a project runs as its own instance, in order to publish its controlled terms and give each one a stable, resolvable identifier together with its authoritative definition. SeaDOTs operates its own instance for the project's building blocks and vocabularies, and that instance is linked into the wider OGC ecosystem, so that its identifiers resolve and its terms connect to the broader web of definitions rather than standing alone. In this role a definition server is to vocabulary what the building-block register is to structure, namely the place where a term is defined once and then referenced everywhere, instead of being reinvented in each dataset. That is why the project's own definition server is the destination for any genuinely new SeaDOTs term. The concrete SeaDOTs instance, and the OGC register used to mint resolvable identifiers, are described in Chapter 11 and listed in Appendix C.

When a field, unit, indicator or relationship cannot be expressed with an available term, the check-in produces a proposal for the OGC Definition Server rather than silently minting an opaque local term. The proposal includes the suggested identifier and label, a definition and scope note, source field names and examples, unit and datatype evidence, candidate matches from known vocabularies, and the source provenance. A new term is used only after human review, or held as a provisional mapping with an explicit marker.

Missing metadata is handled by severity. A blocking gap, such as an absent license, prevents publication as compliant and requires a correction or a curator decision. A medium gap allows publication only with an explicit caveat recorded in the metadata. A low gap can be filled by reliable inference, with the source of the inference recorded. Any value that was inferred rather than supplied by the provider is labeled as such, so that a reader can always tell the difference.

Running through all of this is a point worth stating plainly: the human curator, not the automated agent, has the final say. The agents are good at detecting gaps, drafting proposals and suggesting inferences, but they are poor judges of how much a given gap actually matters. An agent will almost always produce some inference to fill a blank, yet it can rarely say, with justified confidence, whether that blank was a harmless detail or a fact that changes how the data may legitimately be used. Deciding how serious a gap is, whether a proposed term is really the right one, and whether an inferred value is safe to trust is domain judgment, and it belongs to the curator. This is why the severity levels above end in a curator decision rather than an automated one, why a new term becomes official only after human review, and why every inferred value is labeled as such. The automation exists to bring the gap, its evidence and its options to the curator quickly, not to close the gap on its own. Treating a confident-sounding inference as if it were a settled fact is precisely the failure mode this part of the framework is designed to prevent.

10 Building blocks in practice

By this point the building block has already done a great deal of work in this document, from the worked examples in Chapter 5 to the catalog and check-in machinery in Chapters 8 and 9. What has not yet been shown is a plain look at the artifact itself. This chapter is that practical view. It opens the box to show what a building block actually contains, it sets out the simple authoring convention that every SeaDOTs block follows so that any JSON payload can also be read as linked data, and it closes with a table showing how these blocks are reused across the three demonstrators. The next section therefore begins by recalling what is inside a block, now as a concrete inventory rather than as a concept, so that a reader who wants to build one or read one has a firm starting point.

10.1 What is inside a building block

A building block gathers the artifacts that operationalize an interoperability contract. Each block is a self-contained, independently testable package, and it carries up to six machine-actionable artifacts: an identity and dependency declaration, a JSON Schema that fixes its structure, a JSON-LD context in which its terms are bound to shared vocabularies, a set of executable validation shapes, transformation specifications where a source must be converted, and examples that are tested continuously. Human-readable documentation accompanies these, and this is the same anatomy set out in the companion article "Declared Meaning." SeaDOTs reuses building blocks from the OGC Building Blocks collection, from ILIAD, from OGC Open Science and from its own repositories, and it maintains detailed inventories in the corresponding repositories and in the appendices to this document.

The main categories of building block used by SeaDOTs are catalog blocks, data blocks, a shared envelope block, provenance blocks, indicator blocks and relationship blocks, together with a family of blocks specific to the Norwegian reef-effect calculation.

10.2 The convention for semantic uplift

All SeaDOTs building blocks follow the same convention, which is what allows any JSON payload to be read as linked data. A schema file defines the JSON structure. A context file maps each JSON key to a precise predicate on the web of data. Any JSON payload can then be interpreted as RDF by attaching the context. Where present, validation-rule files check the resulting graph.

A new data type follows a standard layout. It has a definition file with an identifier, name, abstract, and dependencies; a schema; a context; human-readable documentation; a set of examples; example files; and a set of positive and negative test cases. Schemas can reference other building blocks by identifier, which is what lets the shared envelopes and abstract concepts of Section 5 be reused rather than copied. A single command runs the standard OGC post-processing that validates the package and builds its published form.

10.3 Demonstrator mapping

The following table shows, for each demonstrator, how a data type maps from its source format to a canonical format and to the building blocks that formalize it. It illustrates how much reuse comes from the shared ILIAD blocks, with project-specific blocks added only where needed. Each block name links to its page in the building-block register, so an interested reader can jump straight to its schema, context, and examples. The reused blocks live in the ILIAD features register and the project-specific blocks in the SeaDOTs register.

Demonstrator Data type Source format Canonical format Reused ILIAD block SeaDOTs block
Germany Sea surface temperature NetCDF-CF, ERDDAP GeoZarr stac_multidim_data, coverageJSON
Germany Salinity, nutrients, chlorophyll NetCDF-CF GeoZarr stac_multidim_data
Germany Mussel growth NetCDF-CF CoverageJSON grid coverageJSONFisheries
Germany Energy production CSV GeoParquet, GeoJSON oim-obs-cs
Common Fishing activity GFW API JSON GeoJSON, GeoParquet gfw-fishing-events
Germany Harvest time series scenario GeoJSON GeoJSON and GeoParquet harvest-timeseries-scen-m3-source, harvest-timeseries-scen-m3-geoparquet
Common Wind and wave fields NetCDF-CF GeoZarr, CoverageJSON stac_multidim_data, coverageJSON
Common Wind turbine configuration CSV, GIS GeoParquet oim-obs, stac_multidim_data
Norway Indicator relationships (Utsira) Cross-impact matrix (CSV) PropertyRelationship (JSON-LD to RDF) oim-variables, expressing the crossImpact-Utsira-OWF relationship
Sweden Macroalgae observations HELCOM API GeoJSON and RDF macroobservation
Sweden Simulation output CSV GeoParquet swedish-DT-simulations-output
Sweden Species abundance and biomass CSV, HELCOM GeoParquet oim-bio-tdwg
Sweden Economic indicators CSV GeoParquet oim-obs-cs
All Workflows and DT applications Excel inventory STAC Collection apkg, tst-inventory-seadots
All Experiments (DT runs) Excel inventory STAC Collection bblocks-openscience
All Marine protected areas EMODnet WFS GeoJSON marine-protected-area-emodnet

The crossImpact-Utsira-OWF entry is deliberately not a link. It is not a standalone building block and has no page of its own. It is the concrete cross-impact relationship for the Utsira case, expressed through the oim-variables register and the PropertyRelationship pattern, so only oim-variables carries a link.

11 Reading a building block in the viewer: a worked example

Every building block in the register has its own page in the OGC building block viewer, and that page is organized as a set of tabs. Each tab is a different view of the same underlying definition. The definition itself is a small package of files that a build tool assembles: a JSON Schema that fixes the structure, a JSON-LD context that fixes the meaning, a set of examples, a set of validation rules, and sometimes even a set of transformers that transform one representation to another (not shown here). The tabs simply present that one package from different angles. The About tab reads it as prose, the Data Structure tab reads it as a tree, the JSON Schema tab shows the raw schema, the Semantic Uplift tab shows the context, the Validation tab shows the rules, and the Dependencies graph shows how the block sits on top of others.

In the following sections, we walk through these tabs one at a time. The running example is the block named ogc.hosted.seadots.swedish-DT-simulations-output, titled Swedish DT Simulations Output. It is a SeaDOTs profile for one row of the Swedish Digital Twin herring and sprat fishery simulation. It carries the status labels Schema and Under development, and its one-line summary reads: a profile for Swedish Digital Twin herring and sprat fishery simulation output, with examples for the raw tabular artifact, a SensorThings Observation view, and a GeoParquet representation header, where 17 of the 63 source columns are reserved-for-future-use placeholders rather than region indicators. That single sentence already tells you what the block is for and warns you about the one data trap it contains. The tabs below fill in the detail.

11.1 The About tab

Figure 14: The About tab of the Swedish DT Simulations Output block.

The About tab is the human-readable front page of the block. At the top, it repeats the title, the machine identifier, the status labels, and the summary sentence, so a reader always knows which block they are looking at and how mature it is. Directly below sits a green banner stating that all examples and tests for this building block pass validation. That banner is a live signal that the examples shipped with the block still satisfy the block’s own rules, so what you are reading is internally consistent.

The Description section underneath is the authoritative narrative for the block. For the Swedish block, it explains that the rows describe herring and sprat fishery state produced by an agent-based model. It records that the source artifact is preserved exactly as supplied, a whitespace-delimited table of 60,000 rows and 63 columns. It states plainly that 46 of those columns are populated fishery, market, and management-scenario indicators, while 17 columns named fu_01 through fu_17 are reserved-for-future-use placeholders that are always zero in the current data. It then makes the key clarification: the fu_ prefix is a naming artifact of the source model, not a resolved region or NUTS code, and those columns must not be treated as region indicators until the model owner defines their meaning. This is the single most important caveat in the whole block, and the About tab is where it is stated.

The Description also introduces the two interoperable views the block offers over the same artifact. The first is a SensorThings Observation view that treats one selected simulation row as an observation of fishery state. The second is a GeoParquet representation header that declares the tabular columns and the GeoParquet metadata needed for a geometry-joined file. The text is careful to say that the GeoParquet header is intentionally metadata-only, because the supplied source carries no row-level geometry, and that the example therefore includes an approximate Swedish case-region footprint as a placeholder that a production file should later replace with authoritative geometry. The section closes with reproducible instructions for regenerating the GeoParquet example, including the required Python packages and the two converter scripts. The point to take away is that the About tab documents not only what the block is but also how its derived artifacts were produced, which is what makes the block trustworthy to a stranger.

11.2 The Dependencies graph

Figure 15: The Dependencies graph in the Full view, with nodes colored by register.

The About tab does not end with the description. Scroll to the bottom of the same page and it shows a Dependencies panel. The Dependencies panel draws the block and everything it stands on as a graph, with a toggle between a Simplified view and a Full view. The Simplified view shows only the nearest, most important neighbors, which is the right level for a first look. The Full view expands the complete transitive set of blocks that are pulled in, directly or indirectly. Nodes are colored by the register they belong to, and the legend names those registers: the SeaDOTs project blocks, the Marine API Profiles, the Observations family drawn from ISO 19156 and OGC and W3C work, and the OGC Main building blocks.

Reading the graph, the Swedish DT Simulations Output node sits in the middle. It points to the two blocks it uses directly, the OIM Observations profile and the GeoParquet Header. From OIM Observations, the arrows cascade down through SOSA Observation Feature and the JSON-FG feature blocks to the OGC Main foundations such as Feature, Feature Collection, and GeoJSON. The value of the graph is that it turns an abstract claim, that this block reuses established standards, into a picture you can trace with a finger. It also explains the Validation tab, because the shapes listed there are exactly the rules attached to the blocks shown here. The graph can be busy in the Full view, so the next section takes a single path through it and follows that path from top to bottom, which is the clearest way to understand what the reuse actually buys you.

11.3 Following one dependency chain from top to bottom

The dependency graph shows every block at once. To explain the benefit of the building block approach, it helps to isolate one straight path through that graph and read it as a stack. The diagram below does that. It follows the Swedish block down one branch of the graph, from the fishery-specific profile at the top to the base geometry encoding at the bottom.

Figure 16: One dependency chain, from Swedish DT Simulations Output down to GeoJSON. Each block “builds on” the one below it.

Read from the top, the chain is this. The Swedish DT Simulations Output block is our project block, and it contributes only the fishery-specific fields. It builds on OIM Observations, the marine profile of an observation. That builds on the SOSA Observation Feature, which fixes what was observed, when it was observed, and what the result was. That builds on the JSON-FG Feature in its lenient form, which adds a time stamp and a coordinate reference system to a feature. That builds on Feature, the general notion of a geospatial feature with geometry, properties, and links. And Feature builds on GeoJSON, the base geometry encoding that essentially every geospatial tool can already read.

The first benefit of arranging things this way is economy of authorship. The Swedish block only had to define the thin top layer, the fields that are genuinely specific to the herring and sprat simulation. Everything underneath, the notion of an observation, of time, of a coordinate reference system, of geometry, of a readable encoding, was inherited rather than reinvented. The author of the Swedish block only needs to write a small profile, not a large standard.

The second benefit is graceful degradation, which is the interoperability payoff. Because each layer is a valid thing in its own right, a consumer can enter the stack at whatever level it understands. A plain GIS tool that only knows GeoJSON still opens the geometry and reads it. A client that understands JSON-FG additionally gets correct time and coordinate reference information. A client that understands SOSA gets the full observation semantics, meaning it knows this is an observation with a phenomenon time, a result time, and a result. A marine client that understands OIM gets the domain profile on top. No single consumer has to understand the whole stack to get value, and more capable consumers are rewarded with more meaning. Nothing breaks for the simple reader while the sophisticated reader gets everything.

The third benefit is inherited validation. As the Validation tab will show in chapter 11.8, the SHACL shapes that check a Swedish row come from these very layers. Because the block sits on SOSA and Feature and their relatives, it is automatically checked against their community-agreed rules without the SeaDOTs team writing a single one of those rules. Correctness is inherited along the same edges as structure.

The fourth benefit is precise, traceable provenance, and this is the point that is easy to conflate with reuse but is genuinely separate. Every field in the Swedish block resolves through the JSON-LD context to a stable URI in a known vocabulary, and every layer in the chain resolves to a published, versioned block. That means the meaning of a value is not a matter of local convention or a comment in a spreadsheet. It is machine-resolvable back to the standard that defines it, and the lineage of the definition itself, which standard each concept came from, is explicit and inspectable. When someone asks in five years what mean_biomass_herring meant and where the observation model came from, the answer is a chain of resolvable links.

The honest answer to the question “is the advantage reuse, or is it the precision of the provenance” is that it is both, and they reinforce each other. Reuse is what lets you build a small, correct profile quickly and have it understood at many levels of sophistication at once. Precise provenance is what lets you prove, later and mechanically, exactly what every value meant and which standard defined it. The chain diagram is worth including precisely because it lets a reader see both of these at a glance: read downward and it is a story of inherited capability, read upward, and it is a story of traceable meaning.

11.4 The Examples tab

Figure 17: The Examples tab, showing the source artifact and the GeoParquet header example with its JSON, JSON-LD, and RDF/Turtle serializations.

The Examples tab holds the concrete instance data that ships with the block. For the Swedish block, there are two examples. The first is the raw source artifact. It is too large to render in the browser, so the viewer offers a download button rather than an inline preview. Its presence matters because it anchors every other view to a real, unmodified input.

The second example is the GeoParquet header, and this is where the viewer shows its most useful trick. The same example is offered through three buttons labeled JSON, JSON-LD, and RDF/TURTLE. These are not three different files. They are three serializations of the same bytes. The JSON button shows the ordinary document a developer would write, with fileName, encoding, a source object, a parquetSchema list, a geo metadata block, and conversionNotes. The JSON-LD button shows that same document with its semantic annotations made explicit, and the RDF/TURTLE button shows it as triples. The ability to flip a plain JSON example into linked data and back, with no change to the payload, is the entire promise of the building block approach expressed in one control. A reader who clicks through the three buttons sees, in a few seconds, why attaching a context to plain JSON is worth the effort.

11.5 The Data Structure tab

Figure 18: The Data Structure tab, with the schema rendered as an expandable tree and inline concept links.

The Data Structure tab renders the schema as an expandable tree, which is the easiest way to read the shape of the data without reading YAML. A small banner notes that this view is a beta feature and may contain errors, so it is a reading aid rather than the normative source. The normative source is the JSON Schema tab described next.

At the root, the tree shows a one of node, which tells you immediately that an instance is allowed to be one of two alternatives. The first alternative is the SensorThings Observation. Its required members are shown at the top: @iot.id, phenomenonTime, resultTime, result, Datastream, FeatureOfInterest, and parameters. Each member carries a small badge for its type and whether it is required, and each carries a resolvable link to the concept it maps to. So phenomenonTime shows a link to the SOSA term of the same name, resultTime links to the SOSA result-time term, and result links to the SOSA has-result term. Expanding result (not extended in the figure above) reveals the model outputs themselves, with repetitions, ticks, current_year, and current_month marked required, and the biomass and catch fields such as mean_biomass_herring and mean_biomas_sprat listed as numbers . The tree preserves the source spelling of mean_biomas_sprat exactly, because the block’s job is to describe the real data, not to silently correct it. A toggle labeled Semantics on switches the inline concept links on and off, and the Expand-all and Collapse-all controls let you move between the overview and the detail. Reading this tab is how you answer the question “what fields are in here, and what does each one mean” without leaving the page.

11.6 The JSON Schema tab

Figure 19: The JSON Schema tab, offering the schema as SOURCE, FULL (YAML), and FULL (JSON).

The JSON Schema tab exposes the normative machine contract. It offers the schema in three forms through the SOURCE, FULL (YAML), and FULL (JSON) buttons, and it gives resolvable URLs to the published schema.yaml and schema.json so a tool can fetch them directly. SOURCE is the schema as the author wrote it, with its references to other blocks left as links. FULL is the same schema with every reference resolved and inlined, so a validator can consume it as a single self-contained document.

The content itself confirms what the tree showed. The schema declares the 2020-12 draft of JSON Schema, then a oneOf between a SensorThingsObservation definition and a GeoParquetHeader definition, then the $defs that spell those out. The SensorThingsObservation lists the same required members seen in the tree. The GeoParquetHeader is a single reference to the separate GeoParquet header block, which is the schema-level expression of a dependency. This tab is the one you cite when you need to say precisely and unambiguously what a conforming instance must look like, because everything else in the viewer is derived from it.

11.7 The Semantic Uplift tab

Figure 20: The Semantic Uplift tab, showing the JSON-LD context and the “Where this context comes from” attribution graph.

The Semantic Uplift tab is where structure turns into meaning. It publishes the block’s JSON-LD context at a resolvable context.jsonld URL and shows the context document on the left of the panel. The context is the mapping that gives every field a globally defined identity. At the top, it sets a default vocabulary, the SeaDOTs swedish-dt# namespace, so that any field without a more specific mapping still resolves to a stable SeaDOTs term. Then it maps the meaningful fields onto community vocabularies: phenomenonTime and resultTime map to the SOSA terms, with resultTime additionally typed as a date-time, result maps to SOSA has-result, Datastream maps to the SensorThings datastream, and parameters maps to schema.org variable-measured. Nested contexts do the same job for the source, parquetSchema, and geo structures. This is exactly the mapping that lets the Examples tab present the same payload as JSON, JSON-LD, or Turtle.

The right side of the panel, labeled “Where this context comes from,” is easy to overlook and worth pausing on. It is a small graph that attributes each part of the context to the block that contributed it, color-coded by register. It makes visible that this block did not invent most of its vocabulary. It assembled it from OGC and marine profiles it depends on and added only the Swedish-specific terms on top. The tab also offers a copy-to-clipboard button and a link that opens the context in the JSON-LD Playground, so a reader can test the uplift on live data. The lesson of this tab is that meaning in the building block world is not written in prose and hoped for. It is declared, resolvable, and reused.

11.8 The Validation tab

Figure 21: The Validation tab, listing the SHACL shape sets used to check the block.

The Validation tab shows how conformance is enforced. It repeats the green pass banner and adds a button to open the full validation report; then it lists the sets of SHACL shapes used to validate the block. For the Swedish block, there are four: the OIM Observations rules from the ILIAD register, the Observation Properties shapes from the OGC API SOSA work, and the Feature and Feature Collection shapes from the OGC geo features register. Each entry gives the machine identifier of the shape set and a resolvable link to the .shacl file.

The important thing to notice is where these shapes come from. The block itself did not author them. They are inherited from the blocks it builds on. Validation therefore checks a Swedish simulation row not only against the block’s own schema but against the community-agreed rules of every layer beneath it.

12 Current status and next steps

This framework is a living document. It describes a design that is partly implemented and partly still being built, and it is honest about which is which. The table below summarizes where the main pieces stand and what remains to close.

Capability Current state Next step
Upload of data to EDITO Working, run manually Automate the pipeline from check-in to publication
EDITO STAC catalog Records created manually Generate records automatically from check-in
EOSC discoverability Not yet tested Add DCAT 3.0 serialization and register an EOSC provider
OGC API serving Partly implemented Deploy a standard OGC API server for features and records
Vocabulary alignment Handled case by case Move to a systematic mapping index
Building-block definitions In progress Continue generating them through the check-in workflow
Indicator registry Vocabulary service deployed on EDITO Publish resolvable identifiers through the OGC register
RDF graph serving Vocabulary service deployed on EDITO Expose a public query endpoint
Validation Validation rules exist for example blocks Add a validation step to continuous integration
Format conversion Partial Complete the transformation library

Beyond closing these implementation gaps, there is one larger direction worth stating plainly. As explained in Section 1.3, this version of the framework demonstrates building blocks for the inputs to a digital twin, its outputs and their lineage. It does not yet demonstrate building blocks for the twin as a living system, meaning the coupled, running models and their feedback with the real ocean. This is the natural next area of work. A future version is expected to apply the same building block approach to the running twin itself, for example by describing model coupling interfaces, runtime state exchange and the contracts between components of a live twin. Readers should therefore treat the present document as a solid foundation for the data and provenance layer, and as an explicit invitation to help extend the approach upward to the twin itself.

13 Getting started

The approach in this document is not specific to SeaDOTs, and it is not reserved for projects that have to merge many sources at once. It is just as useful when you have a single dataset and simply want to publish it well, because the same description that makes data interoperable also makes it self-explaining. The fastest way to learn the approach is to take one of your own datasets through the steps the framework prescribes, and this section sets out that path in general terms so that a reader from any domain can begin with their own data. Any project that has to combine data from different sources can adopt it, and the fastest way to learn it is to take one of your own datasets through the same steps the framework prescribes. This section sets out that path in general terms, so that a reader from any domain can begin with their own data.

Begin by getting familiar with the OGC building block tooling and with at least one existing register. The OGC Building Blocks documentation describes the standard layout of a block and the post-processing tool that validates and builds it, and public registers give you dozens of real blocks to read. Chapter 11 walked through one such block in the viewer, and reading a few others the same way is the quickest way to build intuition before you write anything of your own.

Then take a single dataset of your own and profile it. Determine its format, its fields, its spatial and temporal coverage, and whatever metadata it already carries. This is the intake and profiling step of the check-in workflow in Section 9, and doing it by hand once teaches you what the workflow later automates.

Next, look for an existing building block that already fits your data, and prefer reuse over invention. If your dataset measures the same real-world quantity as one that is already described, you can map to the same shared concept and take the existing block as a template. If it merely shares a file structure, you can reuse the structural block and describe your own content separately. These are the two situations set out in the worked examples of Section 5, and they are the best mental model to keep in mind while you work.

Where no existing block fits, author a new one following the standard layout: a definition file, a JSON Schema that fixes the structure, a JSON-LD context that maps each field to a precise term, human-readable documentation, a set of examples, and positive and negative test cases. Reference other blocks by identifier rather than copying them, so that shared envelopes and abstract concepts are reused. This is the convention described in Section 10.2.

As you describe the data, map its meaningful fields to authoritative vocabularies first, and record anything that no existing term covers as a candidate definition rather than an opaque local label. Then run the standard post-processing command, which validates the package and builds its published form, and check that your examples pass. This mirrors the validation contract of the check-in workflow.

Finally, publish and register the result so that others can find it, trust it, and reuse it. That is the point at which your own data joins the same web of resolvable definitions and provenance that the rest of this document describes. From here the approach is the same whatever your domain: reuse what already exists, add only what is genuinely new, and make every addition testable and resolvable.

One motivation deserves to be stated on its own, because it matters even when there is only one dataset and no integration in sight. Describing data with building blocks is also what makes it ready for automated and AI-driven use. A dataset that carries a JSON-LD context, resolvable vocabulary terms, and a full provenance record does not ask a consumer to guess what a column means, which units it is in, where it came from, or how it was produced. All of that is written down explicitly and resolves to a definition. When the consumer is a language model or an autonomous agent, this is exactly what removes the need to infer the missing pieces, and inference over missing meaning and missing lineage is where hallucination enters. By supplying the meaning and the provenance up front, the framework minimizes, and for the facts it records can eliminate, the guesswork a model would otherwise fill in on its own. Producing AI-ready data is therefore not a separate exercise. It is the same building-block description, seen from the point of view of a machine that has to trust and reason over what it reads.

For a concrete, working reference you can copy from, the building blocks and tooling that SeaDOTs uses and maintains are public, and the demonstrator mapping in Section 10.3 shows how each dataset was handled in practice.

14 Conclusions

This framework defines how SeaDOTs achieves interoperability without forcing its environmental, biodiversity, socio-economic and digital-twin data into a single common schema. Specialized data profiles, catalog profiles, controlled vocabularies and reusable OGC building blocks are connected through shared catalog and provenance mechanisms, so that each class of asset keeps the structure its own community actually uses while still being discoverable, validated and published through one governed check-in workflow. The canonical formats, together with the OGC API Records, STAC and PROV-O catalog model, give the three demonstrators a common way to register workflows, executions, inputs and outputs on EDITO, with a path to EOSC discoverability through DCAT and GeoDCAT-AP.

The same description pays off well beyond interoperability. Because every dataset carries its structure, its meaning, and its provenance explicitly and resolvably, it is also self-explaining to software. A machine consumer, including a language model or an autonomous agent, can read what each field means, in which units, and where it came from, rather than inferring it. Supplying that meaning and lineage up front is what makes the data AI-ready, and it removes a principal source of hallucination, namely the guesswork a model resorts to when the meaning or the provenance is missing. This holds for a single well-described dataset just as much as for a fully integrated collection.

The framework is not yet closed, and it is not meant to be. The status section records the implementation work that remains, and the scope section records the larger conceptual step of extending building blocks to the running twin. Closing the first turns the framework from an architectural reference into an operational pipeline. Taking the second will turn it from a description of a twin’s data into a description of the twin itself. Both are within reach of the approach set out here, which is why this document is offered as much as a starting point for others as a record of what the project has built.

References

  1. EMODnet Data Guide. https://emodnet.ec.europa.eu/
  2. EMODnet Human Activities. https://emodnet.ec.europa.eu/en/human-activities
  3. OGC API — Features: OGC 17-069r4. https://docs.ogc.org/is/17-069r4/17-069r4.html
  4. OGC API — Records: OGC 20-004. https://docs.ogc.org/is/20-004/20-004.html
  5. OGC API — Coverages: OGC 19-087. https://docs.ogc.org/DRAFTS/19-087.html
  6. OGC API — Environmental Data Retrieval: OGC 19-086r6. https://docs.ogc.org/is/19-086r6/19-086r6.html
  7. OGC SensorThings API: OGC 18-088. https://docs.ogc.org/is/18-088/18-088.html
  8. GeoSPARQL 1.1: OGC 22-047r1. https://docs.ogc.org/is/22-047r1/22-047r1.html
  9. OGC Building Blocks documentation. https://ogcincubator.github.io/bblocks-docs/
  10. SOSA/SSN: W3C Recommendation 2017. https://www.w3.org/TR/vocab-ssn/
  11. PROV-O: W3C Recommendation 2013. https://www.w3.org/TR/prov-o/
  12. DCAT 3.0: W3C Recommendation 2024. https://www.w3.org/TR/vocab-dcat-3/
  13. SKOS: W3C Recommendation 2009. https://www.w3.org/TR/skos-reference/
  14. JSON-LD 1.1: W3C Recommendation 2020. https://www.w3.org/TR/json-ld11/
  15. SHACL: W3C Recommendation 2017. https://www.w3.org/TR/shacl/
  16. GeoDCAT-AP 3.0. https://semiceu.github.io/GeoDCAT-AP/releases/3.0.0/
  17. INSPIRE Directive 2007/2/EC. https://inspire.ec.europa.eu/
  18. SDMX 2.1. https://sdmx.org/
  19. NUTS classification, Eurostat. https://ec.europa.eu/eurostat/web/nuts
  20. NERC Vocabulary Server. https://vocab.nerc.ac.uk/
  21. Darwin Core terms, TDWG. https://rs.tdwg.org/dwc/terms/
  22. QUDT quantity kinds. http://qudt.org/vocab/quantitykind/
  23. ICES code lists. https://vocab.ices.dk/
  24. SeaDOTs project. http://seadots-project.eu
  25. ILIAD project. https://iliad-oceandtp.eu
  26. Ocean Information Model. https://github.com/ILIAD-ocean-twin/OIM
  27. ILIAD building blocks. https://github.com/ogcincubator/iliad-apis-features
  28. SeaDOTs building blocks. https://github.com/ogcincubator/bblocks-seadots
  29. OGC Open Science building blocks. https://github.com/ogcincubator/bblocks-openscience
  30. OGC register documentation. https://ogcincubator.github.io/rainbow-docs/

Appendix A. Standards stack

The table below shows which standard governs each layer of the framework. All the standards are open, and where an OGC standard applies the canonical specification is cited.

Layer Standard Version Authority Notes
Feature geometry GeoJSON RFC 7946 IETF All vector features
Feature geometry, extended JSON-FG OGC 21-045r1 OGC Adds 3D, CRS, instants
Coverage encoding CoverageJSON OGC 21-069r2 OGC Community Standard Grid, trajectory and point series
Multidimensional storage NetCDF-CF CF 1.11 CF Conventions Variables, units, CRS
Cloud-native storage GeoZarr OGC 23-064 OGC draft Preferred cloud format
Columnar tabular GeoParquet 1.1 OGC Replaces CSV for tabular
Statistical exchange SDMX 2.1 2.1 SDMX Eurostat and ICES data
Spatial metadata catalog STAC 1.1 1.1 Radiant Earth Primary catalog format
OGC API metadata OGC API — Records OGC 20-004 OGC Aligned with STAC
Open-science provenance PROV-O with STAC W3C REC W3C Workflow and experiment links
Open-science catalog bblocks-openscience OGC Incubator STAC and PROV-O profile
EOSC discovery DCAT 3.0, GeoDCAT-AP 3.0 W3C, EC EOSC harvesting
Observations SOSA/SSN W3C REC W3C Sensor and observed property
Ocean observations profile OIM ILIAD Profiles SOSA for marine data
Indicator vocabulary SKOS W3C REC W3C Concept schemes
Indicator provenance PROV-O W3C REC W3C Composite indicator derivation
Governance ontology Ostrom SES model Ostrom (1990) Governance framework
Quantities and units QUDT 2.1 2.1 QUDT.org Relationship weights and units
Linked data JSON-LD 1.1 W3C REC W3C Context files in building blocks
RDF serialization Turtle, RDF-star W3C W3C Provenance-annotated triples
RDF validation SHACL W3C REC W3C Validation shapes
Building-block infrastructure OGC Building Blocks OGC Incubator Schema, context and validation
Application packages APKG ILIAD CWL-aligned DT packages

Appendix B. Building-block index

This appendix lists the building blocks that SeaDOTs reuses and the ones it defines. Reused blocks come from the ILIAD features repository. Project-specific blocks are defined in the SeaDOTs repository.

B.1 Reused from ILIAD

Building block Type Purpose in SeaDOTs
OIM Observations (oim-obs) schema Base observation schema, GeoJSON feature with SOSA properties. Used by all demonstrators
OIM Observations Collection (oim-obs-cs) schema Feature collection of OIM observations
OIM STA Observations (oim-sta-obs) schema SensorThings-compatible observation variant
OIM full context library (oim) model Cross-domain context mapping many ontologies to OIM schemas
OIM Biological / TDWG (oim-bio-tdwg) schema Extends OIM with Darwin Core and Humboldt Core for biological records
ILIAD STAC multidim data (stac_multidim_data) schema STAC item profile for multidimensional data, mapping the datacube extension to OIM and EDR semantics
CoverageJSON (coverageJSON) schema CoverageJSON schema and context for grid, point-series and trajectory data
CoverageJSON Fisheries (coverageJSONFisheries) schema CoverageJSON profile for fisheries grid data
Coverage Information Model (coverage_information_model) schema OGC coverage abstract model schema and context
Zarr Array Metadata (zarr_array_metadata) schema Array metadata schema used when registering GeoZarr in STAC
Zarr Attributes Metadata (zarr_attrs_metadata) schema Attribute metadata schema mapping CF attributes to STAC and OIM
Zarr Attributes SeaDataNet (zarr_attrs_sdn) schema Attribute profile for SeaDataNet vocabulary terms
GFW Fishing Events (gfw-fishing-events) schema Global Fishing Watch fishing-event schema
Macroobservation (macroobservation) schema Schema for HELCOM macrospecies point records in the Swedish zone
Marine Protected Area (marine-protected-area-emodnet) schema Schema for EMODnet marine-protected-area features
NERC CF Standard Names (nerc-cf-standard-names) model Profile for referencing CF standard names in datasets
APKG Metadata Application Profile (apkg) model Application-package metadata, modeling DT applications as software assets
SeaDOTs Inventory (tst-inventory-seadots) schema Transform schema converting an Excel inventory into catalog-ready JSON
EMODnet Biology Occurrences (emodnet-biology-occurrences) schema Darwin Core biodiversity occurrence records as GeoParquet
Indicator Quality Requirement (indicator-quality-requirement) schema Declares the data-quality and geospatial requirements an indicator must satisfy
NINA SEAPOP catalog, target and source schema Seabird colony dataset for Norway, with discovery, target and source profiles

B.2 Defined by SeaDOTs

Building block Type Purpose
Property Relationship Ontology model Ontology defining the property-relationship pattern and its terms
Property Relationship schema Record for a directed, weighted cross-impact link between two properties, with model attribution and experiment provenance
Equation Property Relationship schema Extends the property relationship with equation, role, operator and index information
OIM Variables and Indicators model Register of variable and indicator concepts and their relationship edges
OIM Variable Observation schema Observation profile for a variable or indicator value
GeoParquet Header schema The reusable GeoParquet metadata envelope referenced by dataset-specific profiles
Swedish DT Simulations Output schema Swedish fishery simulation output, as an observation view and a GeoParquet representation
Harvest time series scenario, source and GeoParquet schema Source-faithful and GeoParquet representations of the German harvest time series
Marine Area of Interest schema Simple polygon profile for a marine area of interest used by experiments and assessments
Catalog Data schema The generic catalog record profile for a data artifact, independent of role or type
Catalog Data Multidimensional schema Catalog profile for gridded and array products
Catalog Data Tabular schema Catalog profile for tabular products, adding table and GeoParquet metadata
Catalog Data Tabular Survey schema Catalog profile for survey data, carrying social-science thesaurus and variable metadata
Catalog Application Package schema Profile for an executable package attached to a workflow record
Catalog Workflow schema Profile for a discoverable, reusable workflow, model or DT application
Catalog Execution schema Profile for one concrete execution, linking workflow, input and output records
EMODnet-compliant windfarm schema Feature profile aligned to the EMODnet human-activities windfarm model
Reef-effect family schema The cluster of blocks for the Utsira reef-effect calculation, including the colonization time factor, the submerged infrastructure geometry, the MAREANO and IMR benthic biomass observations, the reef aggregation index, the reef-effect process and output, the concrete Utsira execution, and the ODD protocol record

Appendix C. Vocabulary registry

SeaDOTs data is described with terms from the vocabularies below. Each entry gives the prefix used in the repositories, the authority and the coverage.

Prefix Authority Coverage
sosa W3C Observations, sensors, features of interest
ssn W3C System and deployment metadata
skos W3C Concept schemes and indicator vocabulary
prov W3C Provenance of indicators and experiments
dcat W3C Dataset and catalog metadata
dct Dublin Core General metadata terms
geo OGC Spatial RDF features
sf OGC Simple-features geometry types
qudt QUDT.org Quantities, units and numeric values
owl, rdfs, xsd W3C Ontology declarations, labels and datatypes
cf NERC / CF CF standard names
dwc, hc TDWG Darwin Core and Humboldt Core biological terms
iliad ILIAD ILIAD observable-property register
ind, indo, indp SeaDOTs Indicator scheme, observed properties and model parameters
ostrom SeaDOTs Ostrom governance classes
prop-rel SeaDOTs / OGC Property-relationship building block
stato OBO Statistical methods ontology

External vocabulary services used include the NERC Vocabulary Server for CF standard names and SeaDataNet parameters, the ICES code lists for fishery codes and stock identifiers, the SDMX and Eurostat code lists for statistical classifications, the SeaDOTs definition server for the hosted indicator vocabulary, and the OGC register for indicator registration.

Appendix D. Namespace registry

The table lists the RDF namespaces used across SeaDOTs files and where each is defined.

Prefix Defined in
ostrom Ostrom model file
ind, indo, indp Indicator examples in the OIM variables source
prop-rel Property-relationship ontology in the SeaDOTs repository
iliad ILIAD OIM context

Appendix E. Tools inventory

The project maintains a set of tools that support the check-in workflow. The existing project tools are summarized below.

Tool Purpose Status
Metadata augmentation notebook Turns an Excel inventory into catalog metadata, optionally with language-model assistance Working, manual
push2edito Uploads files and directories to EDITO storage Working
metadata extraction Extracts a STAC item from NetCDF Partial
STAC tools Reads an OGC features endpoint into a feature collection Working
json2rdf Converts a JSON indicator into RDF Turtle Minimal
Fuseki deployment Provides a query endpoint for RDF graph data Deployed locally

A wider set of reusable tools is available in the project repositories, including a check-in wizard that registers a data source, matches it to a building block, maps its vocabulary and stages the block, a building-block index and a vocabulary index that support matching, a transformer runner that runs and validates transformations, and a set of individual transformers for common conversions.

Appendix F. Assisted tooling inventory

The project also maintains a collection of assisted tools that help apply the framework, currently comprising a set of skills, a set of agents and a set of commands. The authoritative and always-current list is generated automatically in the tooling repository, so the summary here is deliberately brief.

The skills cover building-block cataloging, relevance ranking, validation, vocabulary alignment, context completeness checking, and a range of format conversions such as CSV to GeoParquet, GeoJSON to GeoParquet, NetCDF to STAC, and CSV to metadata, together with helpers for serving and testing building blocks through the reference OGC API implementation. The agents cover metadata generation, format detection and routing, building-block generation, dataset usability assessment, marine-data discovery and profiling, workflow orchestration, and validation. The commands provide entry points for the most common operations, including registering a new data source, assessing usability, generating a building block, generating metadata, converting formats, testing rendering and validating a block.

Appendix G. Data flow architecture

The figure below shows the detailed data flow, from a data source through the check-in wizard, validation and catalog generation, to publication on EDITO and serving through the OGC APIs and the query endpoint.

Figure 11. The detailed data-flow architecture, from source acquisition through validation and catalog generation to publication and serving.

Appendix H. Definitions, acronyms and abbreviations

Term Meaning
OGC Open Geospatial Consortium
W3C World Wide Web Consortium
STAC SpatioTemporal Asset Catalog
DCAT Data Catalog Vocabulary
GeoDCAT-AP Geospatial extension of the DCAT application profile for European data portals
EDITO European Digital Twin Ocean
EOSC European Open Science Cloud
ILIAD The ILIAD digital twin of the ocean project
OIM Ocean Information Model
RDF Resource Description Framework
JSON-LD JSON for linking data
SHACL Shapes Constraint Language
SKOS Simple Knowledge Organization System
SPARQL SPARQL Protocol and RDF Query Language
PROV-O The W3C Provenance Ontology
SOSA / SSN Sensor, Observation, Sample and Actuator, and Semantic Sensor Network ontologies
CF Climate and Forecast metadata conventions
SDMX Statistical Data and Metadata eXchange
NUTS Nomenclature of Territorial Units for Statistics
INSPIRE Infrastructure for Spatial Information in the European Community
ICES International Council for the Exploration of the Sea
NERC / NVS Natural Environment Research Council, and its Vocabulary Server
HELCOM Baltic Marine Environment Protection Commission
MAREANO Marine area database for Norwegian waters
IMR Institute of Marine Research, Norway
OBIS Ocean Biodiversity Information System
QUDT Quantities, Units, Dimensions and Types vocabulary
TDWG Biodiversity Information Standards
GFW Global Fishing Watch
CRS Coordinate Reference System
PROJJSON The JSON encoding of a coordinate reference system definition
CSVW CSV on the Web
EDR Environmental Data Retrieval, an OGC API
WCS, WFS, WMS, WMTS OGC web coverage, feature, map and map-tile services
ERDDAP Environmental Research Division’s Data Access Program
THREDDS Thematic Real-time Environmental Distributed Data Services
OWF Offshore wind farm
DT Digital twin
SES Social-ecological system
ODD Overview, Design concepts, Details protocol for describing models
APKG Application package, the OGC and ILIAD executable-package profile
CWL Common Workflow Language
FAIR Findable, Accessible, Interoperable, Reusable
AOI Area of interest
EEZ Exclusive economic zone
MPA Marine protected area

  1. SeaDOTs received funding from the European Union's Horizon Europe Research and Innovation Programme, part of the EU Missions, under Grant No. 101156488↩︎

  2. Ocean Biodiversity Information System. It is a global, open-access database that collects, integrates, quality-checks, and publishes information about where marine species have been observed around the world. It is operated under UNESCO's Intergovernmental Oceanographic Commission (IOC).↩︎

  3. An interoperability contract is an explicit, shared specification of how data or a service must be structured, described, and validated, so that independent systems can exchange and correctly use it without prior coordination.↩︎

  4. Benthic biomass is the total mass of living organisms that live on, in, or attached to the seabed↩︎

  5. An ICES area annotation is a geographic label that indicates which standardized marine management area a data point belongs to. ICES stands for the International Council for the Exploration of the Sea, which divides the North Atlantic, North Sea, Norwegian Sea, Barents Sea, and adjacent waters into a hierarchy of statistical areas and subareas used by fisheries scientists and marine researchers. These areas provide a common spatial reference for reporting observations, stock assessments, and environmental data.↩︎

  6. WoRMS (World Register of Marine Species)↩︎

  7. OBIS (Ocean Biodiversity Information System)↩︎