hello@datascale.de+49 89 921 35 623tracked cookie-free · /openDEEN

Search services, integrations and blog posts.

HomeServicesData Platform & Governance

Service 04 · Data Platform & Governance

Your composable CDP and marketing data warehouse, plus governance

We build your Marketing Data Lakehouse on your own cloud (BigQuery/Snowflake), plus the governance layer on top: binding definitions, automated quality tests, and an access model so the numbers stop drifting apart.

First call within 48 h · 2 weeks · from €4,500 net

You're in the right place if:

Shop, ad systems, and CRM show three different revenue numbers.

Meetings start by debating which numbers are the right ones.

LTV, CAC, or attribution can't be calculated cleanly internally.

One spreadsheet ties every source together by hand, once a week.

BigQuerySnowflakedbtfunnel.io

What we build

Five building blocks. One platform.

From architecture to governance, bookable individually or as a chain.

01

Data architecture & lake

BigQuery EU · Snowflake · CMEK

BigQuery or Snowflake as the central, EU-hosted data source. We decide tool and schema for the concrete context, not from a standard template.

Current-state analysis of all sources, target architecture document, BigQuery vs. Snowflake decision with reasoning, cost control.

02

ELT pipelines & modelling

funnel.io · dbt · Airflow

Every relevant source runs into the lake automatically and becomes tested models in dbt. Staging, marts, business logic.

funnel.io plus custom connectors for shop, CRM, finance. dbt transformations, daily orchestration with monitoring.

03

Marketing activation & reverse-ETL

Hightouch · Census · Klaviyo · HubSpot

The tested data goes back where it acts: into ads, CRM, and marketing automation. No export spreadsheet in between.

Attribution (data-driven instead of last-click), LTV on your own transaction data, audience segments via reverse-ETL.

04

Definitions & business glossary

Notion / Confluence · dbt docs · DataHub

If "lead" means three different things across marketing, CRM, and finance, no dashboard fixes it. We define the central metrics once, bindingly.

Stakeholder mapping, binding glossary for lead, MQL, SQL, conversion, revenue, mapping onto the fields in GA4, CRM, and ERP.

05

Data quality & access model

dbt tests · Great Expectations · BigQuery IAM

Definitions without tests drift again. Automated checks warn before silent data loss, and an RBAC model protects production pipelines from accidents.

dbt tests and Great Expectations for critical pipelines, Slack alerts with diagnostic context, RBAC and schema-change management with review gates.

How it works

From source to activation, in four steps.

One central lake, tested and documented, instead of point-to-point exports.

01funnel.io · dbt

Ingest

funnel.io and custom connectors load shop, ads, CRM, and finance into the lake automatically.

02dbt · marts

Model

dbt normalises, enriches, and sets the business logic for attribution and LTV.

03dbt tests · Slack

Test

Automated checks for completeness, freshness, and plausibility. Slack alert on drift, before a dashboard lies.

04Hightouch · Census

Activate

Reverse-ETL pushes the tested audiences back into ads, CRM, and Klaviyo.

Process

Four phases, fixed order.

always starts with phase 1 · no blind build

Phase 1 · 2 weeks

Audit Sprint

Five layers audited, findings ranked, effort estimated. The result is a report you could act on without us.

Phase 2 · 1–2 weeks

Architecture

Data contract, event design, target architecture. We fix where each number is produced and who guarantees it.

Phase 3 · 6–10 weeks

Build Sprint

Delivery in sprints, every module signed off on its own. Your team stays involved, not locked out.

Phase 4 · ongoing

Managed Evolution

Monitoring, release support, platform updates. Optional; plenty of clients run the setup themselves.

What you get

A report, not a workshop afterglow.

The Audit Sprint ends in a document: findings per layer, severity, effort, sequence. Not a slide deck full of recommendations in the subjunctive.

→ Findings with severity and reproduction path

→ Effort estimate per finding, in person-days

→ A draft data contract for the core events

→ An implementation plan another agency could execute

Request an anonymised sample →

Deliverables · Data Platform & Governance

4 groups · 14 items
01Data architecture3 items
02Pipelines & modelling3 items
03Governance & data quality4 items
04Activation & handover4 items

Structure taken from this page's scope of delivery. The concrete scope comes out of the audit.

Scopes

Three ways in, one starting point.

Recommended start

Audit Sprint

from €4,500 net

2 weeks

We audit what is wrong. Prioritised report + action plan.

Request an Audit Sprint →

Build Sprint

Fixed price

6–10 weeks

Fresh build or restructure, built to spec.

Discuss a Build Sprint →

Managed Evolution

monthly

3-month minimum

Ongoing partnership. Analytics as a product.

Request Managed Evolution →

From the integrations catalog

Tools this service works with.

Category: CDP & Event PipelinesCategory: CRMCategory: Data WarehouseCategory: Data Integration & ETL

Asked often

Cleared up front.

Different question? Write to us directly, reply within 48 h.

A classic data warehouse (Redshift, on-premise systems) is highly structured and optimised for SQL queries. A modern data lake or lakehouse (BigQuery, Snowflake) combines the flexibility of a lake with the query performance of a warehouse. For marketing use cases BigQuery is the better choice in most setups today.

funnel.io alone is sufficient for most marketing dashboards (→ Revenue Intelligence). A data lake becomes necessary when shop transaction data, CRM data, and marketing data need to be unified, when LTV or attribution models should be calculated, or when ML applications are planned.

Yes, BigQuery offers EU regions (europe-west3 Frankfurt, europe-west4 Netherlands). Datascale configures all projects in EU regions by default. All data stays in the EU region.

BigQuery is configured exclusively in an EU region (europe-west3 Frankfurt or europe-west4 Netherlands). Data residency is contractually assured by Google. Residual risk: the CLOUD Act targets US parent companies. For highly sensitive data we combine BigQuery encryption with CMEK (Customer-Managed Encryption Keys), optionally with an External Key Manager from EU vendors like Fortanix or Thales. The decryption key then sits outside CLOUD Act reach. For maximum sovereignty: Snowflake on AWS Frankfurt with the same External-Key setup, or an open-source lake on StackIT or IONOS.

An enterprise CDP (Segment, mParticle, Tealium) typically runs €80,000 to €250,000 per year, depending on MTU volume. Plus implementation. A composable setup on BigQuery EU sits in a different order of magnitude: BigQuery storage and compute together stay under €1,500 per month for most DACH mid-market companies, dbt Cloud Team plan from €100 per month, Hightouch Starter from €350 per month. Implementation runs as a Build Sprint at a fixed price set after the audit. Year-one total cost is typically €35,000 to €60,000. The biggest difference is not the price. It is data control: storage belongs to you, not to the CDP vendor.

Data quality measures individual values (completeness, format, range). Data reliability guarantees that a definition means the same thing across systems, and that automated tests keep it that way when someone changes a schema or connects a new source. Data quality is a snapshot; data reliability is a process.

No. We start with the sources that already exist. GA4, Plausible, CRM exports, ad APIs. If the audit shows a central layer is missing, that becomes a recommendation, not an upfront investment. Most reliability problems can be addressed without a finished lakehouse, as long as definitions and tests are properly documented once.

Someone internal, from the data or analytics team. We build the framework, document the processes, and onboard the first steward, but governance only works if it's owned internally. In our experience, externally-staffed steward roles don't last 12 months.

After a 2-week Audit Sprint the 3–5 most critical pipelines have alerts in place. The full build with data catalogue, RBAC model, and governance processes typically runs over 6 to 10 weeks, together with Saloid for the technical implementation.

Next step

Data platform and governance: architecture conversation.

Strategy call about lakehouse architecture, reverse-ETL, and governance. Full-cycle implementation together with Saloid.

Juri Saloid

Your contact

Juri Saloid

Founder & Managing Director

hello@datascale.de