Skip to main content

Dev Station Technology

Data engineering

The plumbing that has to work before any of the clever stuff

Ingestion, storage, movement and migration. We have run readings from device fleets into time series storage, moved years of sensitive records with nothing lost, and put reporting in front of people who actually open it.

Book a 30-minute scoping call See the data work

Most stalled analytics projects are stuck here, not at the model.

Pipeline health, one working day.
Readings ingested MQTT, CoAP and HTTPStreaming
Rows reconciled Source and destination counts matchMatched
Query on a year of history Time series storageFast
Failed batch Retried, then escalated1 alert
Report refreshed Before the morning meeting06:00
Boring when it works, which is the point.

Data work we have delivered

  • IoT WorkzTelemetry into TimescaleDB and Kafka
  • Norwegian EdTechYears of records migrated, none lost
  • Power BI reportingOn a platform used by schools
  • Self-hostedThe client owns the data outright

The honest question first

You may not have a data problem yet

Data engineering is expensive and invisible. It is worth doing when the alternative has started costing more.

Leave it alone when

A spreadsheet is still fine

  • One person exports one report once a week
  • Volumes fit comfortably in a spreadsheet
  • Nobody disputes which number is correct
  • No regulator asks where a figure came from

Building a warehouse for this is a hobby, not an investment.

Fix the plumbing when

The numbers have started disagreeing

  • Two departments quote different figures for the same thing
  • Data arrives faster than anyone can process it by hand
  • You are moving off a system that is past end of life
  • An analytics or AI project has stalled waiting for clean inputs
  • You cannot answer where a number came from

This is the work we do, and it is usually what an AI project actually needs first.

Plain definitions

Two questions that come up before anyone agrees a budget

The second one in particular gets answered by architecture diagrams when it should be answered by a count of your sources.

What is data engineering?

It is the work of getting data from where it is produced to where it is used, in a state somebody can rely on. Ingestion, storage, movement, reconciliation and the monitoring that tells you when a feed has stopped.

Data science gets the attention and data engineering gets the calls at six in the morning. When an analytics or AI project stalls, this layer is almost always the reason.

Warehouse, lake or lakehouse?

A warehouse holds structured data modelled in advance for reporting. A lake holds raw files of any shape and defers the modelling. A lakehouse is the attempt to have both, with warehouse-style querying over lake storage.

For most companies asking us, the honest answer is that none of the three is the first problem. One reliable feed and one agreed definition of the number people argue about will change more than any of these labels. We would rather build that and let the architecture question wait until the volume forces it.

What we build

Six pieces, in the order they usually matter

Ingestion
From devices over MQTT, CoAP and HTTP, from systems over APIs, and from files nobody wants to admit still exist.
Storage that suits the shape
PostgreSQL for relational work, TimescaleDB where readings arrive continuously, object storage for everything raw.
Movement and streaming
Kafka where volume or decoupling demands it, scheduled batch where it does not. We do not reach for streaming by reflex.
Migration with proof
Row counts, checksums and spot checks on both sides. On the Norwegian platform a separate team did nothing else, and nothing was lost.
Reporting
Dashboards and exports in what your team already uses, including Power BI and Grafana.
Monitoring the pipeline itself
Prometheus and alerting, so a broken feed is found by the system rather than by an executive noticing a flat chart.

What people ask us for

Six data jobs that arrive most often

Almost every enquiry is one of these. Two of them are usually happening at once.

Streaming ingestion from equipment
Readings arriving continuously from devices or instruments, sized for the fleet rather than for a demo. This is the IoT Workz ingestion path over MQTT, CoAP and HTTP.
System to system integration
Where an ERP, a CMMS and two things bought separately hold overlapping records, and people retype between them every week.
Migration off a system past end of life
Its own workstream with reconciliation on both sides. The Norwegian platform moved years of sensitive school records with none lost.
One agreed source for a disputed number
The build that starts with two departments quoting different figures. Most of the work is agreeing the definition, not moving the data.
Reporting people will actually open
Dashboards and exports in the tool your team already uses. We have handed this over in Power BI and in Grafana.
Monitoring retrofitted to pipelines that already exist
Where the feeds work but nobody finds out when one stops. Prometheus and alerting, so a flat chart is not the detection method.

Tell us where the numbers break

Three fields. An engineer reads it, not a sales rep. You get an answer within one working day.

  • We map sources, volumes and who disagrees with whom
  • You get a written view on what to fix first, with a cost band
  • No obligation, and no phone number needed to start

Our work

Two jobs where the data was the project

Case study, IoT platform

Readings that never stop arriving, stored so they stay useful

IoT Workz needed measurement data from device fleets without renting it back from a SaaS vendor. We built the ingestion path over MQTT, CoAP and HTTP, put readings into TimescaleDB, used Kafka where decoupling mattered, object storage for raw payloads, and Prometheus with Grafana so the pipeline reports on itself.

  • TimescaleDB
  • PostgreSQL
  • Kafka
  • MQTT
  • MinIO
  • Grafana
  • Prometheus
Platform and data ownership
100%
Critical alarms in live operation
Zero
History kept queryable
Years

Case study, Norway

Moving years of school records while schools kept working

The client had to leave a system that could no longer handle its volumes or meet security standards, without interrupting the schools using it. Migration ran as its own workstream with a dedicated team, reconciled on both sides, alongside a phased rollout and 60 days of warranty support.

  • .NET
  • Angular
  • AWS
  • Power BI
Records lost
Zero
Warranty support after go-live
60 days
Compliance with national data law
Met

How the work runs

One feed end to end before we widen

  1. Week 1

    Map the sources

    Where data comes from, how much, how often, and which number people argue about.

  2. Week 2 to 3

    One feed, end to end

    A single source ingested, stored, monitored and reported, so the shape is proven before scope grows.

  3. Week 4 to 12

    Widen and migrate

    More sources, then any migration, rehearsed before the real run.

  4. Handover

    You run it

    Pipelines, monitoring and documentation are yours, deployed in your own tenant.

Technology

What we build on

Everything here is running somewhere in production. We have left off the tools we have only read about.

Ingestion
MQTT, CoAP and HTTP from devices, APIs from systems, and scheduled file drops from the things nobody admits still exist.
Storage
PostgreSQL for relational work, TimescaleDB where readings arrive continuously, MinIO and object storage for raw payloads.
Movement
Kafka where volume or decoupling demands it, scheduled batch where it does not. We do not reach for streaming by reflex.
Migration
Row counts and checksums on both sides, spot checks agreed with your people, and a rehearsal before the real run.
Reporting
Power BI and Grafana, chosen by what your team already opens rather than by what is fashionable.
Monitoring
Prometheus with alerting on the pipeline itself, so a broken feed raises an alarm rather than a question in a meeting.

Who does the work

Twenty engineers, eight of them senior

Pipelines get handed over and then live for years without us. That only works if the person who built yours wrote documentation somebody else can follow, which is easier to insist on across twenty people than across two hundred.

Engineers in Ho Chi Minh City
20
Senior engineers
8
Average experience
5+ yrs
Working with US, UK and EU teams
10+ yrs

You interview whoever we put forward. Ask what their alerting looks like when a feed fails at two in the morning, because that answer separates the careful from the quick.

How we work

Five engagement models, and you pick how much you keep

The models differ in one thing: how much of the management you hand over. Everything else, including who owns the code, is the same in all five.

Changing model later is normal. Augmentation into a dedicated team is the common direction, and the people stay.

Engineers join your team and work inside your process, your board and your code review.

Team control
You manage the day to day
Pricing
Monthly per person, by role and seniority
Minimum
1 month

A team that works only on your product, with a lead on our side running delivery.

Team control
Shared. Our lead runs delivery, you set priorities
Pricing
Monthly per role, lead included
Minimum
3 months

A long-term engineering site under your standards, where we carry recruitment, HR, payroll and equipment.

Team control
You direct the work, we run operations
Pricing
Monthly per role plus site costs
Minimum
6 months

One price for a scope that is genuinely settled, with the overrun carried by us.

Team control
We deliver, you accept against written criteria
Pricing
Fixed bid, paid against milestones
Minimum
Project based, usually 10 to 14 weeks

Everything in the centre model, with the whole team moving into your own Vietnamese entity on an agreed date.

Team control
Transfers from us to you
Pricing
Monthly per role, transfer price agreed up front
Minimum
2 to 6 years

FAQs

Questions we get on the first call

Do we need a warehouse or a lakehouse?

Usually neither, at first. Most of the value comes from reliable ingestion and one agreed source for the numbers people argue about. We would rather fix that than sell you an architecture diagram.

Can you migrate us off a system nobody supports any more?

Yes. We treat migration as its own workstream with reconciliation on both sides, and we rehearse it before the real run.

Will our reporting tools still work?

Yes. We build to what your team already uses, including Power BI and Grafana, rather than making everyone learn something new.

Where does the data live?

Your own tenant, in the region your policy requires. On the IoT platform the client hosts everything themselves.

How do we know a feed has broken?

The pipeline monitors itself and alerts. Finding out because a chart looks flat is not a monitoring strategy.

Is this the first step towards AI?

Almost always. Models fail on inputs far more often than on algorithms. If an AI project has stalled, this is usually why.

What does data engineering cost, and how long does it take?

One feed ingested, stored, monitored and reported end to end is usually three weeks. A working platform across several sources, with migration behind it, takes three to four months. Migration off a legacy system across regions runs longer again. Which engagement model you pick decides how it is billed, and the models are set out above.

Do you work with Snowflake, Databricks or BigQuery?

We have not shipped on them, and we will say so rather than learn on your budget. Our production work is PostgreSQL, TimescaleDB, Kafka and object storage, self-hosted or in your own cloud tenant. If your organisation has already standardised on one of the big platforms, hire a team that lives in it.

ETL or ELT, and does the difference matter?

ETL transforms the data before it lands, ELT loads it raw and transforms it afterwards. ELT is the more forgiving default, because raw data you kept can be reprocessed and raw data you transformed away cannot. It matters less than the two things people skip, which are reconciliation and monitoring.

Tell us which number people argue about

That argument usually points straight at the broken feed. Thirty minutes is enough to find it and say what fixing it involves.

Book a 30-minute scoping call Read when to leave it alone

Let's Talk