sviat.infocase studieslakehouse
Case 02case.data
39 sources, one source of truth.
Embedded at an international logistics company’s factory, I designed the ETL flows with the operations staff and delivered the Lakehouse that became the source of truth for 5 analytics and ML teams — then cut end-to-end pipeline runtime by 37%.
- Client
- International logistics company
- Role
- Senior Data Engineer · Sigma Software Group
- Where
- On-site at the client’s factory · Münster, Germany
- When
- Jul 2023 – Jan 2025
Animated diagram of the Lakehouse case: records from 39 operational sources flow through ETL and incremental PySpark jobs into one source-of-truth Lakehouse that serves dashboards and 5 analytics and ML teams. Simulated demo.
- 39 operational sourcesMapped on-site at the client’s factory with operations staff
- source 01
- source 02
- source 03
- …
- source 39
- ETL flowsETL flows designed on-site with the operations staff
- incremental PySparkIncremental PySpark on Databricks · −37% end-to-end runtime
- one source of truth39 operational sources consolidated into one Lakehouse
- dashboards · 5 teamsServing 5 analytics & ML teams, plus dashboards for factory & logistics stakeholders
- dashboards
- analytics teams
- ML teams
- raw record
- clean record
- failed check
Hover or focus a stage for proofTap a stage for proof
The data pipeline from the homepage, rewired for this case. Record rates are simulated — the numbers above are the real results.
The problem
The company’s operational data was spread across 39 sources, with no single source of truth for its analytics and ML teams.
The job: one platform the analytics and ML teams could all build on, designed with the people who run the operations.
What I did
01 embed
Work where the data is made
I worked on-site at the client’s factory and designed the ETL flows together with the operations staff.
↑ stage 01 · Sources
02 map
Turn shop-floor workflows into data models
Mapped shop-floor workflows directly into data models and reporting, and designed analytical dashboards for factory and logistics stakeholders on-site.
↑ stage 05 · Serve
03 ship
Consolidate 39 sources into one Lakehouse
Brought 39 operational sources into one Lakehouse — the source of truth for 5 analytics and ML teams.
↑ stage 04 · Lakehouse
04 prove
Make the pipelines incremental
Built incremental PySpark pipelines on Databricks, cutting end-to-end pipeline runtime by 37%.
↑ stage 03 · Process
Results
- 39 → 1operational sources consolidated
- 5analytics & ML teams served
- −37%end-to-end pipeline runtime
Stack & methods
- Databricks
- PySpark
- Lakehouse
- ETL
- Dashboards