sviat.infocase studieslakehouse

Case 02case.data

39 sources, one source of truth.

Embedded at an international logistics company’s factory, I designed the ETL flows with the operations staff and delivered the Lakehouse that became the source of truth for 5 analytics and ML teams — then cut end-to-end pipeline runtime by 37%.

Client
International logistics company
Role
Senior Data Engineer · Sigma Software Group
Where
On-site at the client’s factory · Münster, Germany
When
Jul 2023 – Jan 2025

case 02 · factory_ops → source_of_truth

  • 39 → 1 sources
  • 5 teams served
  • −37% runtime

demopausedstatic frame · reduced motionsimulated

Animated diagram of the Lakehouse case: records from 39 operational sources flow through ETL and incremental PySpark jobs into one source-of-truth Lakehouse that serves dashboards and 5 analytics and ML teams. Simulated demo.

  1. 39 operational sourcesMapped on-site at the client’s factory with operations staff
    • source 01
    • source 02
    • source 03
    • …
    • source 39
  2. ETL flowsETL flows designed on-site with the operations staff
  3. incremental PySparkIncremental PySpark on Databricks · −37% end-to-end runtime
  4. one source of truth39 operational sources consolidated into one Lakehouse
  5. dashboards · 5 teamsServing 5 analytics & ML teams, plus dashboards for factory & logistics stakeholders
    • dashboards
    • analytics teams
    • ML teams
  • raw record
  • clean record
  • failed check

Hover or focus a stage for proofTap a stage for proof

The data pipeline from the homepage, rewired for this case. Record rates are simulated — the numbers above are the real results.

The problem

The company’s operational data was spread across 39 sources, with no single source of truth for its analytics and ML teams.

The job: one platform the analytics and ML teams could all build on, designed with the people who run the operations.

What I did

  1. 01 embed

    Work where the data is made

    I worked on-site at the client’s factory and designed the ETL flows together with the operations staff.

    ↑ stage 01 · Sources

  2. 02 map

    Turn shop-floor workflows into data models

    Mapped shop-floor workflows directly into data models and reporting, and designed analytical dashboards for factory and logistics stakeholders on-site.

    ↑ stage 05 · Serve

  3. 03 ship

    Consolidate 39 sources into one Lakehouse

    Brought 39 operational sources into one Lakehouse — the source of truth for 5 analytics and ML teams.

    ↑ stage 04 · Lakehouse

  4. 04 prove

    Make the pipelines incremental

    Built incremental PySpark pipelines on Databricks, cutting end-to-end pipeline runtime by 37%.

    ↑ stage 03 · Process

Results

  • 39 → 1operational sources consolidated
  • 5analytics & ML teams served
  • −37%end-to-end pipeline runtime

Stack & methods

  • Databricks
  • PySpark
  • Lakehouse
  • ETL
  • Dashboards

↑↓ navigate↵ selectesc close> terminal

↵ run↑ historytab complete⌫ back

Type > for the terminal