Idempotent pipelines: the one habit that saves your weekends
Every pipeline fails eventually. An API times out, a disk fills up, a deploy lands halfway through a run. The question is not whether a job will need to be re-run, but what happens when it is.
If running the same job twice produces duplicate rows, double-counted revenue or a half-updated table, someone ends up fixing data by hand on a Saturday. If the job is idempotent — the result is the same no matter how many times it runs — you simply press retry.
1. Write by partition, not by append
Instead of appending new rows, each run owns a slice of the table — usually a day — and replaces it completely.
BEGIN;
DELETE FROM orders_daily WHERE day = :run_date;
INSERT INTO orders_daily SELECT ... WHERE day = :run_date;
COMMIT;
Re-running yesterday's job rewrites yesterday. Nothing else is touched.
2. Upsert on a natural key
When partitions don't fit, use a key that identifies the record in the source system and let the database resolve conflicts with INSERT ... ON CONFLICT DO UPDATE.
3. Make time an input, not a side effect
A job that reads now() produces a different result every time. Pass the logical run date as a parameter, and backfills become a loop instead of a project.
The payoff
Once every job is safe to re-run, alerts become boring: something failed, retry it. That is exactly what you want at 3 a.m.