What Is a Satellite in Data Vault?
A satellite holds everything that describes a hub: names, statuses, amounts, and every change to them, appended and never updated. It is a slowly changing dimension type 2 without the UPDATE, with a full audit trail built in.
Does this sound familiar?
- A customer's address changed, and the old one is gone because the load overwrote it.
- A balance that changes daily drags the customer's legal name into a new row every night.
- A record disappeared in the source, and the warehouse deleted its history with it.
Hubs say what exists. Links say how things belong together. Neither stores a name, a balance or a status. That is the job of the satellite.
In classical modeling terms: the attributes of an entity, its context, become a satellite.
Where the Customer table carries name, address and segment next to customer_number,
the Data Vault keeps the number in the customer hub and the attributes in a satellite on it.
What a satellite is
A satellite holds the descriptive attributes of a hub and every change to them over time. When a customer moves, the satellite does not overwrite the address: it appends a new row, and the old one stays. Each row records when it entered the vault and which system delivered it, so every version is traceable.
If you know dimensional modeling, a satellite is a slowly changing dimension type 2, without ever running an UPDATE. Historization can also be switched off on request, which gives a non-historized satellite.
The anatomy of a satellite
| Part | Columns | Purpose |
|---|---|---|
| Parent | Hash key of the hub | Which entry of the parent hub this row describes |
| Time and audit | Load date, record source | When the row arrived, and from where |
| Payload | Descriptive attributes | Street, email, amount, status |
| Change detection (optional) | Hash diff | One hash over the whole payload |
The primary key is the parent hash key plus the load date.
The hash diff is a choice, not a rule
Textbook Data Vault says to compute a hash diff for every satellite row to detect changes. In practice it depends on the load:
- Delta loads: the stored rows already carry their hash, so hashing only the incoming rows is cheap. The hash diff wins.
- Full loads: hashing millions of rows on every run costs CPU. A direct column comparison
(
IS DISTINCT FROM) is often faster on modern engines. - Wide tables: the common claim is that wide tables benefit most from hashing. The opposite is closer to the truth: wider rows mean longer string concatenations to hash.
Datavault Builder supports the hash diff as an option, not a mandate. The hash diff gets its own article in this series.
Transactions get their own satellite
Quantities, prices and amounts of an order line describe the transaction itself. With a grain hub for the order line, as described in the article on links, they go into an ordinary satellite on that grain hub, where they are historized like any other attribute.
How to split satellites
- By source system. CRM and ERP data never share a raw satellite. Each source keeps its own lineage, unchanged.
- By rate of change. A daily balance and a legal name in one satellite copy the name into a new row every day. Fast and slow attributes belong apart.
- By sensitivity. Personal data such as name, email and date of birth goes into its own satellite, apart from attributes that identify nobody. GDPR reporting then has one clear place to look, and access rules or a deletion request apply to that satellite alone.
- Never delete. When a record vanishes in the source, its history stays. A record tracking satellite marks that the key is no longer delivered. The exception is a legal obligation to remove information, such as a GDPR erasure request, and that is where splitting by sensitivity pays off.
What changes with Datavault Builder
You map source columns to a satellite in the model. Change detection, the append-only load and the audit columns are generated, and run natively on Snowflake, Databricks, BigQuery, SQL Server, Fabric, Oracle or PostgreSQL. Two situations that classic load patterns do not handle are covered out of the box:
- Bi-temporal loads. A financial institution with end-of-day processing has two timelines: when something was true for the business, and when the warehouse loaded it. Datavault Builder processes the end-of-day history and the load history together and creates bi-temporal satellites. See bi-temporal data processing for how it works; bi-temporality also gets its own article in this series.
- More than one change per load. A data lake often delivers several versions of the same record in one batch. Classic patterns load one change per key and run, so you would have to loop over the data, and every version would be dated to the moment of the load. Datavault Builder loads all changes in one batch and places each record on the timeline where the data lake recorded it.
See It Running on One of Your Sources
Book a free demo and bring the connector that costs you the most, in money or in time.
Three Steps to a Satellite
-
Attach it to exactly one hub
Each satellite describes exactly one hub, keyed on that hub’s hash key plus the load date.
-
Split it sensibly if necessary
Where it helps: by source system, by rate of change, or personal data apart from the rest.
-
Let the loads be generated
Datavault Builder generates the change detection and the append-only load for every satellite.
How Datavault Builder Takes the Friction Out of Ingestion
-
Ingestion is built in
Batch, delta and CDC loads from databases, files, REST APIs, NoSQL and Python sources, with streams such as Kafka arriving as micro-batches. Same platform that generates the warehouse, no second invoice.
-
Your schema, not the vendor's
Source tables are mapped to a Data Vault 2.0 model you designed. A new column or a renamed table changes a mapping, not a chain of post-load scripts.
-
Only deltas move
Hubs, links and satellites load what changed. Full reloads stay in staging instead of being reprocessed downstream every night.
-
History is kept by design
Every change is retained as it arrives, so as-was reporting works even where the source overwrites its own rows.
-
Code you never hand-write
Loading, historization and lineage are generated from the model in real time and run natively on Snowflake, Databricks, BigQuery, SQL Server, Fabric, Oracle or PostgreSQL.
-
One platform, up to nine tools fewer
Modeling, ETL, CI/CD, documentation and lineage in one place. That is what makes 14.7 minutes from requirement to production possible.
Meet Our Expert
Twenty minutes with our Sales Director, and an honest answer on whether this fits your stack.
Matt Collett
Sales Director
Great, pick a time that works for you:
Other Problems This Series Covers
-
Technical Keys vs. Business Keys in Data Vault
Source systems keep the business key in the master record, but every relationship runs on technical keys. Four ways to deal with that, what each one costs, and the pattern we use: a PSA hub for the technical keys, mapped to the business key hub.
-
What Is a Business Key in Data Vault?
A business key is the identifier your employees and customers actually use: the customer number, the invoice number, the contract ID. It is stable, it is shared across systems, and it is what every hub in a Data Vault is built on.
-
What Is a Link in Data Vault?
A link records that business keys belong together: this order belongs to this customer. Datavault Builder covers the classic Data Vault link, and adds a transaction link anchored on a grain hub, which fixes the grain of a transaction and lets transactions relate to each other.
-
What Is a Hub in Data Vault?
A hub represents one core business concept, such as customer, product or account, as the list of its keys: every one the warehouse has ever seen, each exactly once. It stores identity, not state, and that restraint makes it the point where source system silos collapse.
Questions and Answers
- Dan Linstedt. He developed the method in the 1990s and published it around 2000. Data Vault 2.0, which adds hash keys, a methodology and an architecture around the model, followed in 2013.
- The standard reference is Building a Scalable Data Warehouse with Data Vault 2.0 by Dan Linstedt and Michael Olschimke (Morgan Kaufmann, 2015).
- No. Bill Inmon sees Data Vault as an evolution of his third normal form view of the enterprise data warehouse, not as a competing approach.
- No. A dimensional output is often part of a Data Vault implementation: the vault keeps the integrated history, and star schemas are built on top of it for reporting. The same vault can also deliver flat tables or a Unified Star Schema.
- With automation, a third normal form view on top of a Data Vault can be generated completely deterministically. The vault holds the data once, and the 3NF layer is derived from the model. How that works.
- Trying it without automation. Data Vault is built on a small set of strict, repeating patterns, which is exactly what makes it tedious and error prone to write by hand and straightforward to generate. That is why you should use Datavault Builder.
- Yes. Splitting keys, relationships and history into hubs, links and satellites means more tables than a normalized or dimensional model. That is why the physical layer should be abstracted by a model-driven approach like Datavault Builder, where you work on the business model and the tables are generated.
- Book a demo and see a Data Vault built on your own sources, or order a training environment and try it yourself.
- Here. Datavault Builder is a Data Vault automation solution: see pricing or book a demo.
- Datavault Builder is licensed annually, as a subscription to use the software. Perpetual licenses are available on request. The editions and what they include are on the pricing page.