Data Pipelines Should Be Safe to Run Twice
Key takeaway: A pipeline is ready for recurring use when a retry, a late record or an interrupted load has a defined outcome. A successful first run is only the beginning.
Specify what one row means
Suppose a daily pipeline combines order exports with payment records for an operations dashboard. Before selecting an orchestrator, define the grain of the output: one row per order, order line or payment? Each choice changes which joins and aggregations are valid.
Record the business key, timestamp convention, currency, nullable fields and treatment of cancellations. If two systems use different meanings for “completed”, preserve those distinctions until the business rule for combining them is explicit. Otherwise, a technically successful join can still produce a misleading number.
Design retries as normal behaviour
Apache Airflow recommends treating tasks like database transactions and producing the same outcome when a task is rerun. For the order pipeline, appending every downloaded file directly to the final table would make a retry capable of doubling yesterday’s sales. Apache Airflow: best practices
A practical design is to load into a staging area, validate the batch and then merge by a stable key or replace the intended partition. The right approach depends on the source and database. Keep the input batch identifier so that an operator can trace which records a run was meant to process.
Check the data, not just the exit code
An empty export can be valid on a quiet day and a serious problem on a busy one. Combine structural checks, such as key uniqueness and expected columns, with business checks such as freshness, plausible volumes and reconciliation against source totals. Set tolerances with the people who understand the process.
Decide which failures should block publication and which should quarantine records for review. Preserve the last known usable dataset when appropriate, but show its age clearly. A dashboard should not silently present stale values as if a refresh had succeeded.
Rehearse recovery before it is urgent
Use a small test batch to simulate four situations: the same file arriving twice, a late update to an existing order, an interruption after staging, and a renamed source column. Check both the final rows and the run log. Recovery behaviour should be observable, not something inferred from a green status badge.
The handover should name an owner, explain how to replay a bounded period and identify which downstream reports may be affected. A modest pipeline with clear recovery instructions can be easier to operate than an elaborate one whose failures require its original author to investigate every time.
