Blog — Data Vault & Data Warehouse Automation
Articles on Data Vault 2.0, data warehouse automation, Snowflake, Databricks, and practical data engineering.
-
Datavault Builder and the DAMA-DMBOK Framework — How We Support the Eleven Knowledge Areas
How Datavault Builder supports the 11 DAMA-DMBOK knowledge areas — simplifying data warehousing, integration, modeling, and automated data governance.
Read more → -
Data Sovereignty & AI: Enterprise DWH
How do you keep your data under control when AI models need access to everything? This webinar shows how to combine data governance with modern AI capabilities.
Read more → -
Data Vault & Databricks Medallion Stack
Data Vault and Lakehouse are often seen as conflicting approaches. This webinar shows why the opposite is true.
Read more → -
How I Think About Business Value in Data
Most data projects don’t fail because of technology. They fail because value arrives too late, in the wrong form, or at too high a cost.
Read more → -
ISO 27001:2022 Certification & SOC 2 Type 2
In today’s rapidly evolving digital landscape, data security is not just a priority—it's a necessity.
Read more → -
Import Flow.BI AI generated Data Models
Watch how to migrate Flow.BI AI-generated metadata into Datavault Builder using the Migration Vault approach — mapping hubs, links, and automated ETL pipelines.
Read more → -
Integrating and Unioning Data
In Data Vault Modeling, we use hubs to integrate data. This is one of the main reasons we choose Data Vault modeling.
Read more → -
Near Real-Time DWH Analytics
In today's data-driven world, credit cards, networks, IoT sensors, and numerous data sources provide real-time data. How can this be processed effectively?
Read more → -
SSO on Snowflake Snowpark Container Services
Using Datavault Builder with Snowflake's Snowpark container services lets data teams set up and manage data solutions without relying on infrastructure teams.
Read more → -
Bi-Temporal Data Processing
Whether you're dealing with customer accounts, financial transactions, or insurance claims, this webinar equips you to handle bi-temporal data.
Read more → -
DVB on Snowflake Snowpark Container Services
In this video we do demonstrate how simple it is to run Datavault Builder on Snowflake Container services.
Read more → -
Direct Model GIT Check-In and Check-Out
Welcome to our short presentation on Datavault Builder 7.1 and its new GIT check-in and check-out feature for agile development.
Read more → -
The Migration Vault Concept
A Migration Vault is a specialized metadata store designed as a Data Vault. It facilitates migrating data from an old format to a new one, ensuring compatibility.
Read more → -
Talend Open Studio Alternative Compared
Datavault Builder offers a large variety of features and modules for the same cost as previously free ETL tools.
Read more → -
Unified Star Schema Automation
In this video, viewers delve into the world of automated data warehousing with the Unified Star Schema, a concept by Francesco Puppini and Bill Inmon.
Read more → -
Yale University: Revolutionizing Data Management for 75% Savings and Faster Insights
How Yale University used Datavault Builder to cut 75% of billed consulting hours, expand their internal data team, and accelerate time to insight.
Read more → -
BI-SPEKTRUM Case Study: C&A on Snowflake
BI-SPEKTRUM published an article in issue 2023/3 how one of our clients is using the Datavault Builder to integrate its SAP Data.
Read more → -
DDVUG Willibald Use Case
Watch the DDVUG Willibald use case: building a full Data Warehouse with 2 data sources in under 3 hours using Datavault Builder.
Read more → -
Is Data Modeling dead
Do we still need data models? What is the value of models? Why did data modeling fail in the past? How can we create value by modeling?
Read more → -
Data Vault Bi-temporal: Inscription Time
How to load bi-temporal data into the Data Vault using Inscription Time — patterns, pitfalls, and practical examples.
Read more → -
Data Vault vs. Data Mesh?
Should I still do Data Vault if there is Data Mesh? In the past few weeks and months, I got these very interesting questions which brought up a several times.
Read more → -
CI/CD with Datavault Builder on Snowflake
How to use Snowflake's Zero Copy Cloning with Datavault Builder to build a powerful CI/CD pipeline for your data warehouse.
Read more → -
Do Equi-Joins always matter?
It happens that from time to time I comeacross some statements about how databases work and how they shall be queried. And I like to read those recommendations.
Read more → -
3NF and Data Vault: Nothing to Fear
From time to time we receive an interesting question: does Datavault Builder support 3NF? The answer is yes — and here is how it works.
Read more → -
DWH Temporality Pt.4: SCD Type 2 Dimensions
Kimball Style dimensions - SCD Type 2 Output If you haven't read them I recommend reading the first 3 parts first.
Read more → -
DWH Temporality Part 3: Outputting Timelines
Although in many cases it is not necessary to output the timelines in the reports, there are some cases where the output of timelines is important.
Read more → -
DWH Temporality Part 2: Reducing Complexity
How to reduce temporal complexity in the Data Vault. Datavault Builder users frequently ask how to map changes over time correctly.
Read more → -
DWH Temporality Part 1: The Challenge
In the past years, I was confronted with the demand to create a reporting with SCD type 2 dimensions.
Read more → -
On Multi-Active Satellites in Data Vault
Petr Beles on implementing Multi-Active Satellites as Document Satellites in Data Vault: patterns, trade-offs, and practical guidance.
Read more → -
On Links
Petr Beles on Data Vault links representing transactions: patterns, pitfalls, and design decisions when modeling transaction links.
Read more → -
Qlik
Do Your Qlik Apps and Power BI Reports Show Different Numbers?
Neither Qlik nor Power BI is wrong. Each has its own load logic, its own definitions and its own lineage, so the same metric is calculated twice and nobody can reconcile them. The rules belong in one governed warehouse model both tools read.
Read more → -
Tableau
Does Your Tableau Server Have Too Many Versions of the Same Number?
Every published .tdsx was reasonable on the day it was made. Together they are hundreds of private definitions of the same metric, and no way to tell which one is right.
Read more → -
Power BI
Do Your Power BI Reports Show Different Numbers for the Same Thing?
Finance, Sales and Operations each built their own semantic model, and each is internally consistent. The definitions were never wrong in one place. They were never agreed in any place.
Read more → -
Qlik
Do You Clean the Same Data Again in Every Qlik Load Script?
The same mapping loads, string fixes and deduplication are written again in every app and QVD layer, and the copies drift apart. The cleansing is repeated per script because no integrated warehouse layer does it once.
Read more → -
Qlik
Are Your Qlik Set Analysis Expressions Too Long and Too Slow?
Set analysis is precise for genuine comparisons. Most of the long expressions in your app are there because the model never delivered history, flags or a single grain, so the chart rebuilds them on every selection.
Read more → -
Tableau
Are Your Tableau LOD Expressions Too Complex to Touch?
FIXED, INCLUDE and EXCLUDE are precise tools for genuine multi grain questions. Most of the ones in your workbook are there because the warehouse never resolved the grain or kept the history.
Read more → -
Power BI
Is Your Power BI DirectQuery Report Slow on Every Click?
DirectQuery and Direct Lake promise live data. What you get is a thirty second visual and a compute bill nobody wants to explain. The mode is not the problem. The schema underneath it is.
Read more → -
Qlik
Do Your Qlik Apps Keep Creating Synthetic Keys?
Synthetic keys and circular references are Qlik associating exactly what it was given. They appear because the data arrives without conformed dimensions or real keys, so every app has to invent them.
Read more → -
Qlik
Do Your Qlik Reloads Keep Failing as the Data Grows?
The in-memory engine is fast because everything sits in RAM. Reloads fail and apps slow down when that RAM is filled with row level detail and transformations that no warehouse did beforehand.
Read more → -
Tableau
Do Your Tableau Extract Refreshes Keep Failing or Running Late?
The backgrounder times out, the .hyper file keeps growing, and the dashboard shows yesterday. The extract is large because it is carrying raw rows that were never aggregated upstream.
Read more → -
Power BI
Is Your Power BI DAX Getting Too Long to Maintain?
Two hundred lines of CALCULATE and FILTER is not a sign of advanced DAX. It is usually a sign that the warehouse never gave you the keys, the history or the grain you needed.
Read more → -
Tableau
Is Your Tableau Dashboard Slow Every Time You Change a Filter?
Twenty seconds of "Executing Query" on every filter click. Tableau is not rendering slowly. It is waiting for a database that was handed a question it cannot answer quickly.
Read more → -
Power BI
Does Your Power BI Refresh Keep Failing Overnight?
Scheduled refresh times out, Power Query runs out of memory, and you find out when someone opens the dashboard. The cause is almost never Power BI. It is what Power BI is being asked to do.
Read more → -
dbt
The dbt Trade-Off: Code Sprawl, Hidden TCO, and the Case for Automated Modeling
dbt brought software engineering to SQL, and teams are right to want that. The trade is a transformation layer that grows in code, in cost and in compute, on top of an ingestion tool that is still a separate product. A generated Data Vault takes that whole layer off the code base. It does not need dbt, but it can be combined with it.
Read more → -
Azure Data Factory
Azure Data Factory Pain Points: The Hidden Cost of UI Pipelines and Spark Transformations
Azure Data Factory is a good transport layer inside Azure and a poor place to keep a data warehouse. Used as the modeling and transformation suite it brings a canvas nobody can read, releases that fail on ARM templates and Spark clusters for loads that fit in one SQL statement. And every one of those pipelines exists only in Azure.
Read more → -
dbt
Does dbt Still Leave You to Land the Data Yourself?
dbt is a transformation tool by design. It assumes the raw data is already in the warehouse, so the landing step is a second product with a second bill, even now that Fivetran and dbt Labs are one company. A warehouse platform that ingests and models in one place closes the gap without a second contract.
Read more → -
SSIS
SSIS Pain Points: Why Lifting Packages to the Cloud Moves the Debt, Not the Problem
SSIS estates have three problems that a cloud VM does not fix: a package per table nobody wants to open, releases that DevOps cannot reach, and a design that stops at SQL Server. Azure-SSIS Integration Runtime carries all three to Azure intact. A model that generates the warehouse retires them instead.
Read more → -
dbt
Is Your dbt Run Quietly Driving Up Warehouse Compute?
dbt executes everything in the warehouse, so a full refresh where an incremental would do, or a table materialization that rebuilds every night, shows up as credits, not as an error. Incremental logic is optional in dbt. In a generated Data Vault it is the only way loads are written.
Read more → -
Azure Data Factory
Are Azure Data Factory Mapping Data Flows Costing More Than the Data They Move?
Mapping Data Flows run on a managed Spark cluster that takes minutes to start and bills by the vCore hour. For a large nightly transformation that is reasonable. For a few hundred thousand rows it is a cluster spun up to do what one SQL statement would do inside the warehouse.
Read more → -
dbt
Does Your dbt Project Have More Models Than Anyone Can Explain?
ref() makes a new model a one line decision, and a project of hundreds of loosely governed SQL files is the result. Changing a business key upstream then means finding every model that inherited it. The fix is not more tests. It is a model that generates the structure instead of accumulating it.
Read more → -
SSIS
Does SSIS Stop Where Your Cloud Warehouse Starts?
SSIS was built to move data between on premises SQL Server instances. When the warehouse moves to Fabric, Snowflake, Databricks or BigQuery, the choices are lifting the packages onto an Azure-SSIS Integration Runtime, buying third party connectors, or rewriting. A model that generates natively for the new platform is the fourth choice, and the only one that does not carry the packages along.
Read more → -
dbt
Is Your dbt Bill Growing in Seats, or in Engineers?
dbt Core is free to run and expensive to operate. dbt Cloud is easy to operate and priced per developer plus usage. Either way the transformation layer has a cost that grows with the team, and neither path is the one you chose for that reason.
Read more → -
Azure Data Factory
Does Every Azure Data Factory Release Turn Into an ARM Template Fight?
Under the visual editor, an Azure Data Factory is JSON: pipelines, datasets, linked services and the ARM template that deploys them. Promoting a change from Dev to Prod means parameter files, global parameters and a template that fails on one type mismatch. Releases should be generated from a model, with the rollback included.
Read more → -
Fivetran
Fivetran Cost and Schema Headaches? How Automated Modeling Fixes Ingestion Friction
Plug and play ELT is fast to start and slow to control. Two things erode over time: what the pipeline costs, and who decides the shape of the data. An automated Data Vault layer gives both back without giving up the connectors that work.
Read more → -
SSIS
Is SSIS the One Part of Your Stack That Still Cannot Do CI/CD?
A .dtsx file is XML that diffs badly and merges worse, so two engineers on one package end in a rebuild. Environments live in SSISDB variable mappings maintained by hand. Modern DevOps stops at the SSIS project. Releases should be generated from a model, per environment, with the rollback included.
Read more → -
Fivetran
Is Your Fivetran Schema Forcing You to Rebuild the Model After Every Load?
Fivetran lands each source in its own standardized schema. Business keys, custom fields and legacy structures then have to be reshaped in SQL that breaks when the connector changes. The fix is not a better script. It is a model that owns the shape.
Read more → -
Azure Data Factory
Has Your Azure Data Factory Canvas Outgrown the People Who Built It?
A drag and drop pipeline is quick to build and slow to change. Past a few dozen activities the canvas turns into the documentation, the wiring takes over the logic, and every new source is another copy activity nobody wants to touch. The fix is not a tidier canvas. It is a model that generates the pipelines.
Read more → -
Fivetran
Does Your Fivetran Bill Jump Every Time a Source Table Gets Busy?
Fivetran charges by Monthly Active Rows. A bulk update or a schema migration upstream touches every row again, and the invoice follows. The rows were never the problem. Paying per row for data you then reprocess yourself is.
Read more → -
SSIS
Do Your SSIS Packages Outnumber the People Who Understand Them?
Hundreds of .dtsx packages, one per table, each built in Visual Studio by whoever had the ticket, with control flows and data flows that only open one at a time. Adding a column means opening the packages one by one. The fix is not a package template. It is a model that generates the loads, natively on SQL Server, Azure SQL or Fabric.
Read more → -
Building Blocks
Technical Keys vs. Business Keys in Data Vault
Source systems keep the business key in the master record, but every relationship runs on technical keys. Four ways to deal with that, what each one costs, and the pattern we use: a PSA hub for the technical keys, mapped to the business key hub.
Read more → -
Building Blocks
What Is a Business Key in Data Vault?
A business key is the identifier your employees and customers actually use: the customer number, the invoice number, the contract ID. It is stable, it is shared across systems, and it is what every hub in a Data Vault is built on.
Read more → -
Building Blocks
What Is a Satellite in Data Vault?
A satellite holds everything that describes a hub: names, statuses, amounts, and every change to them, appended and never updated. It is a slowly changing dimension type 2 without the UPDATE, with a full audit trail built in.
Read more → -
Building Blocks
What Is a Link in Data Vault?
A link records that business keys belong together: this order belongs to this customer. Datavault Builder covers the classic Data Vault link, and adds a transaction link anchored on a grain hub, which fixes the grain of a transaction and lets transactions relate to each other.
Read more → -
Building Blocks
What Is a Hub in Data Vault?
A hub represents one core business concept, such as customer, product or account, as the list of its keys: every one the warehouse has ever seen, each exactly once. It stores identity, not state, and that restraint makes it the point where source system silos collapse.
Read more →
Nothing here yet.