Hiring a data engineer
Seventy-five applications listing Airflow, dbt, Snowflake and Spark in roughly that order, with nothing anywhere about how much data any of it was moving.
Scale is the signal, and it is the one thing nobody states. A pipeline moving a few thousand rows nightly and one moving several billion events are the same words and completely different engineering. Everything that makes the second hard — backfills, late-arriving data, idempotency, what happens when a job fails at 4am and the downstream report is due at 8 — is invisible on a CV that lists the orchestration tool. Worse, the tooling has become so standardised that the stack list genuinely does not differentiate: almost every candidate has used almost the same things. The classic false positive is a candidate who built a tidy pipeline once, in a course, and the false negative is someone who spent three years keeping an ugly legacy system correct and does not think that counts.
How much data, and what happens when it breaks
Selene asks for volumes and failure behaviour, which are the two things that make pipeline experience comparable.
Volume, freshness, and who is downstream
Three facts describe a data platform role: how much data, how fresh it has to be, and who notices when it is wrong. Freshness requirements in particular change the engineering completely, and are almost never in the job ad.
One sentence changes the guide
It describes the failure mode this discipline is prone to. Data platforms accrete: each new question becomes a new pipeline, nobody deletes anything, and after three years there are four tables that nearly agree. An engineer who has done the unglamorous work of collapsing that is worth several who have added to it.
Exceptional, not just good —
“Someone who has deleted a pipeline instead of building another one.”
→ Removing or consolidating infrastructure became a must-have; pipelines built stopped being counted.
What the guide ends up measuring
Six criteria, one naming a tool. Correctness leads because finance is downstream — on a role feeding a product analytics dashboard, latency and flexibility would take those top two slots instead.
No gates, and the stack list is the false one
Nothing here is a yes-or-no fact. Teams commonly gate on named tools instead — must have dbt, must have Airflow — which on a discipline this standardised filters out almost nobody it should and a few people it should not, since the concepts transfer in days and the judgment does not.
Scored on scale and correctness
Selene never receives the candidate’s name or the raw CV. She scores a blind, structured profile, with any stated age, sex, nationality and religion stripped out before scoring runs.
She reads for stated volumes and described failures — the two things that make one pipeline CV comparable to another, and the two things applicants most often leave out.
How the blind scoring works →What stays with you
Selene's half
- Getting volume, freshness and who is downstream on the record before writing criteria
- Reading all 75 against the same guide, blind, with the evidence attached
- A ranked shortlist that is not ordered by which tools appear on the CV
- An interview brief that asks for the volumes each candidate left out
Yours
- Getting data engineers to apply — Selene doesn’t post to job boards or source candidates
- Any SQL or modelling exercise; she reads what people wrote about their work
- The interview, with her brief in hand
- The offer, and every call that matters
Give her the volume
Rows per month and how fresh it must be — two numbers, and the guide stops being a stack list.