Work/HEMA/Turnover
Data Engineering

What was sold, what was paid, and where the two do not meet

HEMA finance is building one source of truth across tills, orders, payments and gift cards. Before that can exist, every source has to publish clean data. We are building those pipelines, starting with the hardest one.

Customer HEMA
Client HEMA
Department Finance
Service Data Engineering, Cloud Native Engineering
Platform AWS
Scope Data mesh datasets, turnover foundation
The challenge

A SALE LIVES IN ONE SYSTEM, THE TRANSACTION IN ANOTHER

Orders arrive from the tills and from order management. Payments arrive from the payment providers. Gift cards arrive from a third party. It is hard to put the three next to each other and see the gaps.

Finance wants two questions answered. Was everything we sold actually paid? Was everything that was ordered actually delivered? And when an auditor asks what happened on a given day, HEMA wants to answer that day.

This project delivers the system that answers all three. HEMA is moving from a file sharing platform and point-to-point APIs to a data mesh, where every source publishes to a shared catalog.

So the first work is not the system. It is the sources. And the hardest of them produces several thousand events a minute, every minute the shops are open.

HEMA shop floor
HEMA shop floor. Every transaction at every till is an event the pipeline has to accept and keep.
What we did

Build the sources first, in the order they are needed

Turnover is a finance system with a queue in front of it. We work at the front of that queue: the datasets, the ingestion, and the model turnover will read.

01

Sit with the other teams

The biggest problem in this project is not technical. Turnover needs data that belongs to other teams, so we joined those teams and built the datasets with them rather than filing requests.

02

Accept everything, validate later

A till cannot wait. API Gateway takes the request and does no validation, so the intake never builds a backlog that takes the system down with it. Validation happens further along.

03

Draw one write boundary

The till is only told the event was accepted once Kinesis confirms it. Past that line the event is ours, it cannot be lost, and it can be replayed.

04

Buffer the fast against the slow

Tills fire continuously, analytics reads in batches. Kinesis sits between the two so each runs at its own pace, and leaves room to hang a real-time consumer off the same stream later.

05

Publish tables, not files

Firehose lands the raw events in S3. Lambdas run Athena queries over them and produce Glue tables. Those tables are the POS dataset, and any team in the mesh can subscribe to it.

06

Never scan twice

Athena bills for data scanned, so the layout was designed to pinpoint only the newest data. A run touches around 15 MB instead of the whole history.

The POS chain, simplified
POS terminal HTTP event
API Gateway Accept, no validation
Kinesis Write boundary, replay
Firehose → S3 Raw events, retained
Lambda → Athena Newest partitions only
Glue tables The POS dataset
The result

The hard part is done, and it costs pocket change to run

10 mln
Records pushed through in a stress test, two working weeks of POS data
5 min
Time the pipeline needed to process all of it, end to end
15 MB
Data scanned per run, which is what keeps the query bill at nothing worth measuring
4 of 4
Source datasets migrated to the data mesh: tills, order management, gift cards and payments

HEMA usually sees a few thousand POS events a minute. The pipeline handled ten million in one go, so ten times the current number of stores would not trouble it. The alternative route would have added several hundred euro a month, thousands a year, and left turnover reading data over an API instead of from the catalog.

The sources

Four satellites, one catalog

Each of these used to hand its data over as files or through a bespoke API. Each now becomes a published dataset that turnover, and anyone else, can read.

Migrated

POS

Sales at the till. By far the most complex of the four: several thousand events a minute, arriving whether or not anything downstream is ready for them.

Migrated

OMS

Orders from the online channels. The other half of the sold side, and the source that tells finance whether an order was actually delivered.

Migrated

Gift cards

Issued and redeemed value from the gift card provider. Money that moves on a different timeline from the sale it belongs to.

Migrated

Payments

The payment provider, and the card schemes behind it. The paid side of the ledger, and where the anomalies turnover is meant to catch will show up.

The turnover pipeline feeding the published datasets in the catalog.
What we delivered

The pipelines turnover will be built on

A POS ingestion system that does not flinch

Serverless intake sized well past what the estate produces today, with a persistence boundary the tills can trust and replay when something downstream needs rebuilding.

Published datasets in the mesh

POS, order management and gift cards as structured Glue tables in the shared catalog, available to any team that subscribes, not only to turnover.

A transformation layer built for the bill

Lambdas that run Athena over the newest partitions only. Cost was a design input from the first day, not a surprise on the first invoice.

Room for real time, unused

Analytics reads in batches, so the pipeline is near real time by choice. A real-time consumer can be attached to the same stream without touching what is there.

The turnover data model, in progress

Modelling and analysis of what finance needs to reconcile sold against paid.

Carsten Klomp, Brighting

“A single source of truth is only as good as the sources. So we started there, with the one that produces a few thousand events a minute.”

Carsten Klomp, Brighting
What we brought

Engineers who work across the teams, not around them

A finance system that reads from four other departments is an organisational problem before it is a technical one. We build in each of those teams, which is why the datasets exist.

Stress test before you promise

Ten million records through the pipeline is not the load HEMA has. It is the load we wanted proof of before saying the system scales, and it cleared in five minutes.

Run cost as an architecture decision

Athena charges by data scanned, so the partitioning was designed around that. The difference between a considered layout and a naive one is thousands of euro a year.

Say what is not finished

Turnover is not done. The foundations under it have, and the business value lands when finance starts reconciling on them.

AWS API Gateway Kinesis Data Firehose S3 Lambda Athena Glue SageMaker Unified Studio Data mesh

More of our work

All cases →
ASICS

Full readiness assessment for AI scalability across 42 capabilities

42 capabilities scored · 2-week assessment
ASICS

Eight weeks from signing to the first agent live

8 weeks to first agent live
HEMA

Real-time pricing on every shelf, in every store

800+ stores live · 4-year partnership

Can your finance team answer what an auditor asks?

Tell us which systems hold your sales, your payments and your deliveries today. We will tell you what it takes to read them as one.

+31 20 210 13 90 info@brighting.nl Grasweg 183, 1031 HX Amsterdam