A bad week
A bug in our ingestion pipeline had been corrupting records for about a week before anyone noticed. The data looked plausible, which is the worst way for data to be wrong — nothing alerted, nothing crashed, and downstream systems had been happily consuming it the entire time.
Fixing the bug took an afternoon. Undoing it took much longer, because re-ingesting a week of records meant doing it without disturbing the traffic still flowing in.
The process we had for this was not really a process. An engineer triggered batch jobs by hand. Progress was invisible unless you went and looked. Anything that failed had to be found and retried manually, and verifying the result was a separate manual exercise that someone did afterwards with a spreadsheet. Start to finish, days.
The shape of the fix
Cloud Storage → Cloud Functions → Pub/Sub → Dataflow → BigQuery
Serverless throughout, because re-ingestion is bursty by nature. It runs hard for an hour and then nothing happens for three weeks, and I didn't want to pay for idle capacity between incidents.
Cloud Storage holds the raw files — JSON, CSV, Parquet — and an upload is what starts everything. There's no separate trigger to remember.
Cloud Functions handle orchestration: validating and preprocessing files, splitting the large ones into chunks that can be worked in parallel, publishing to Pub/Sub, and coordinating verification once processing finishes.
Pub/Sub is the part that made this survivable. At-least-once delivery, retry with exponential backoff, dead letter queues for messages that never succeed, and replay when you need to run something again. During an incident, the ability to replay is worth more than almost anything else in the stack.
Dataflow running Apache Beam does the actual work, across as many workers as the volume justifies.
Verifying it, which is the hard half
Moving data is straightforward. Convincing yourself it arrived correctly is the part that takes design.
During ingestion, each record goes through schema validation, business rule checks, referential integrity, and duplicate detection. Failures route to a dead letter queue with an error message attached, so the thing that failed and the reason it failed stay together.
After ingestion, a reconciliation job compares record counts between source and destination, runs checksum validation on a sample, executes the business-critical queries whose answers we already know, and writes a report.
That last check is the one that earns its place. Counts matching tells you that you moved the right number of rows. Running a query whose correct answer you know independently tells you that you moved the right rows.
Validation accuracy landed at 99.9%, and it caught problems that would otherwise have reached production a second time.
Scale, and what it actually bought
Dataflow scales from one worker to a hundred and more as the load demands. Pub/Sub handles the message volume without tuning. Cloud Functions scale on their own. BigQuery takes the writes.
Concretely: a 5GB dataset that used to take six hours now finishes in under thirty minutes.
Manual intervention dropped roughly 95%. Operations that ran for days run in hours. No data has been lost during a re-ingestion since. And the team stopped spending its week babysitting batch jobs.
What I'd carry forward
Assume failure. Retry logic, circuit breakers, and graceful degradation at every layer, because in a distributed system the question isn't whether a component fails mid-run but which one and how far in.
Idempotency is the feature that makes the rest possible. Every operation produces the same result when re-run. This sounds like a purity concern and is actually a practical one: it means that when something breaks halfway, the recovery procedure is "run it again" rather than a careful reconstruction of what did and didn't complete.
Instrument before you need it. Logging and monitoring went in from the first version, not after the first incident. Every hour spent on that came back several times over the first time something went wrong at 2am.
The system nobody thinks about is the one that's working. This one gets remembered about twice a year, which is the correct number.
