What Is a Satellite in Data Vault?

A satellite holds everything that describes a hub: names, statuses, amounts, and every change to them, appended and never updated. It is a slowly changing dimension type 2 without the UPDATE, with a full audit trail built in.

What Is a Satellite in Data Vault?

The data warehouse automation solution trusted by data teams across industries

Does this sound familiar?

  • A customer's address changed, and the old one is gone because the load overwrote it.
  • A balance that changes daily drags the customer's legal name into a new row every night.
  • A record disappeared in the source, and the warehouse deleted its history with it.

Hubs say what exists. Links say how things belong together. Neither stores a name, a balance or a status. That is the job of the satellite.

In classical modeling terms: the attributes of an entity, its context, become a satellite. Where the Customer table carries name, address and segment next to customer_number, the Data Vault keeps the number in the customer hub and the attributes in a satellite on it.

What a satellite is

A satellite holds the descriptive attributes of a hub and every change to them over time. When a customer moves, the satellite does not overwrite the address: it appends a new row, and the old one stays. Each row records when it entered the vault and which system delivered it, so every version is traceable.

If you know dimensional modeling, a satellite is a slowly changing dimension type 2, without ever running an UPDATE. Historization can also be switched off on request, which gives a non-historized satellite.

The anatomy of a satellite

Part Columns Purpose
Parent Hash key of the hub Which entry of the parent hub this row describes
Time and audit Load date, record source When the row arrived, and from where
Payload Descriptive attributes Street, email, amount, status
Change detection (optional) Hash diff One hash over the whole payload

The primary key is the parent hash key plus the load date.

The hash diff is a choice, not a rule

Textbook Data Vault says to compute a hash diff for every satellite row to detect changes. In practice it depends on the load:

  • Delta loads: the stored rows already carry their hash, so hashing only the incoming rows is cheap. The hash diff wins.
  • Full loads: hashing millions of rows on every run costs CPU. A direct column comparison (IS DISTINCT FROM) is often faster on modern engines.
  • Wide tables: the common claim is that wide tables benefit most from hashing. The opposite is closer to the truth: wider rows mean longer string concatenations to hash.

Datavault Builder supports the hash diff as an option, not a mandate. The hash diff gets its own article in this series.

Transactions get their own satellite

Quantities, prices and amounts of an order line describe the transaction itself. With a grain hub for the order line, as described in the article on links, they go into an ordinary satellite on that grain hub, where they are historized like any other attribute.

How to split satellites

  • By source system. CRM and ERP data never share a raw satellite. Each source keeps its own lineage, unchanged.
  • By rate of change. A daily balance and a legal name in one satellite copy the name into a new row every day. Fast and slow attributes belong apart.
  • By sensitivity. Personal data such as name, email and date of birth goes into its own satellite, apart from attributes that identify nobody. GDPR reporting then has one clear place to look, and access rules or a deletion request apply to that satellite alone.
  • Never delete. When a record vanishes in the source, its history stays. A record tracking satellite marks that the key is no longer delivered. The exception is a legal obligation to remove information, such as a GDPR erasure request, and that is where splitting by sensitivity pays off.

What changes with Datavault Builder

You map source columns to a satellite in the model. Change detection, the append-only load and the audit columns are generated, and run natively on Snowflake, Databricks, BigQuery, SQL Server, Fabric, Oracle or PostgreSQL. Two situations that classic load patterns do not handle are covered out of the box:

  • Bi-temporal loads. A financial institution with end-of-day processing has two timelines: when something was true for the business, and when the warehouse loaded it. Datavault Builder processes the end-of-day history and the load history together and creates bi-temporal satellites. See bi-temporal data processing for how it works; bi-temporality also gets its own article in this series.
  • More than one change per load. A data lake often delivers several versions of the same record in one batch. Classic patterns load one change per key and run, so you would have to loop over the data, and every version would be dated to the moment of the load. Datavault Builder loads all changes in one batch and places each record on the timeline where the data lake recorded it.

See It Running on One of Your Sources

Book a free demo and bring the connector that costs you the most, in money or in time.

Three Steps to a Satellite

  1. Attach it to exactly one hub

    Each satellite describes exactly one hub, keyed on that hub’s hash key plus the load date.

  2. Split it sensibly if necessary

    Where it helps: by source system, by rate of change, or personal data apart from the rest.

  3. Let the loads be generated

    Datavault Builder generates the change detection and the append-only load for every satellite.

How Datavault Builder Takes the Friction Out of Ingestion

  • Ingestion is built in

    Batch, delta and CDC loads from databases, files, REST APIs, NoSQL and Python sources, with streams such as Kafka arriving as micro-batches. Same platform that generates the warehouse, no second invoice.

  • Your schema, not the vendor's

    Source tables are mapped to a Data Vault 2.0 model you designed. A new column or a renamed table changes a mapping, not a chain of post-load scripts.

  • Only deltas move

    Hubs, links and satellites load what changed. Full reloads stay in staging instead of being reprocessed downstream every night.

  • History is kept by design

    Every change is retained as it arrives, so as-was reporting works even where the source overwrites its own rows.

  • Code you never hand-write

    Loading, historization and lineage are generated from the model in real time and run natively on Snowflake, Databricks, BigQuery, SQL Server, Fabric, Oracle or PostgreSQL.

  • One platform, up to nine tools fewer

    Modeling, ETL, CI/CD, documentation and lineage in one place. That is what makes 14.7 minutes from requirement to production possible.

Recognized by BARC in The Data Fabric Survey 26

Meet Our Expert

Twenty minutes with our Sales Director, and an honest answer on whether this fits your stack.

Matt Collett

Matt Collett

Sales Director

What are you looking for?

By submitting you agree to our Privacy Policy.

Other Problems This Series Covers

  • Technical vs. Business Keys

    Technical Keys vs. Business Keys in Data Vault

    Source systems keep the business key in the master record, but every relationship runs on technical keys. Four ways to deal with that, what each one costs, and the pattern we use: a PSA hub for the technical keys, mapped to the business key hub.

  • Business Key

    What Is a Business Key in Data Vault?

    A business key is the identifier your employees and customers actually use: the customer number, the invoice number, the contract ID. It is stable, it is shared across systems, and it is what every hub in a Data Vault is built on.

  • Link

    What Is a Link in Data Vault?

    A link records that business keys belong together: this order belongs to this customer. Datavault Builder covers the classic Data Vault link, and adds a transaction link anchored on a grain hub, which fixes the grain of a transaction and lets transactions relate to each other.

  • Hub

    What Is a Hub in Data Vault?

    A hub represents one core business concept, such as customer, product or account, as the list of its keys: every one the warehouse has ever seen, each exactly once. It stores identity, not state, and that restraint makes it the point where source system silos collapse.

Questions and Answers