Where this started
Support agents kept filing tickets that were, underneath, the same request: "can someone run a query for me?" The data existed. They just had no way to reach it. So the request went into an engineering queue, waited its turn behind actual feature work, and came back a day or two later — by which point the customer conversation it was meant to inform had usually moved on.
Twenty-four to forty-eight hours for a data pull. Nobody was happy with that, least of all the engineers writing the same SELECT for the fifth time that month.
The obvious fix is "give support SQL access." The obvious reason nobody had done it is that handing a production database to fifty people who've never written a JOIN ends badly in a number of ways. So the actual problem was narrower: how do you give people the shape of SQL without giving them a SQL prompt?
What it's built on
React on the front, with React DnD driving the canvas. Node and Express behind it. PostgreSQL read replicas as the data source, never the primaries. A query builder in the middle that turns canvas state into SQL. The whole thing runs in Docker on an internal Kubernetes cluster.
The builder
Users drag tables onto a canvas, draw joins between them, tick the columns they want, and fill in filters through form inputs. Before anything executes, they see the SQL that their canvas produced.
That preview turned out to matter more than I expected. It was meant as a transparency feature — show people what you're about to run on their behalf. What it became was the main way people learned SQL. Agents would build something visually, read the generated query, and after a few weeks start predicting what the SQL would look like before they saw it. A couple of them eventually stopped using the canvas entirely and asked for a raw query box.
Keeping it safe
Every connection is read-only, so the worst case is a slow query rather than a mangled table. Only approved table patterns are reachable — an allowlist, not a denylist, because a denylist is a promise you can't keep. Each user gets rate limited so one person's cartesian product doesn't starve everyone else. Queries hard-timeout at thirty seconds. Every execution is logged with the user attached.
The thirty-second timeout caused the most friction early on, and it was also the setting I was least willing to move. A query that takes longer than thirty seconds in this tool is nearly always a mistake — a missing join condition, usually — and failing fast teaches that faster than any documentation.
Making it fast enough to trust
Results cache in Redis for thirty minutes. Large result sets stream back in chunks rather than materializing in one go. Users can pull up an explain plan when something feels slow. Common aggregates sit behind materialized views.
The caching was the single biggest perceived-speed win, and mostly for a boring reason: support queries repeat. Several agents investigating the same incident run near-identical queries within the same hour. The first one pays the cost and the rest come back instantly.
Where it landed
Data request tickets to engineering dropped by about 90%. The waiting went from a day or two to right now. Around fifty people across three support teams use it daily, running roughly 500 queries between them.
The number I actually watch is the ticket count, because it's the one that says engineers got their afternoons back.
What I'd tell myself at the start
Scoping the first version to read-only queries was the decision everything else depended on. It made the security model small enough to reason about in one sitting, and it removed the entire category of "what if someone drops a table" from the design.
Advanced features — raw SQL, explain plans — stay hidden until someone goes looking. Putting them on the main surface would have made the tool look like a database client, which is precisely the thing that scares off the people it was built for.
Onboarding deserved more of my time than I gave it at first. What eventually worked was sample queries that answered real questions people already had, so learning the tool and doing the job were the same activity.
And the query logs were worth more than the monitoring dashboards. Watching which queries people actually ran showed me which parts of the schema were confusing, and where someone was quietly struggling and about to give up.
