Public filing data is public, technically. Usable is a different word entirely.
This is what I ran into building Insidr, which tracks SEC insider trades (Form 4 filings) and congressional stock trades (STOCK Act disclosures), and what actually worked to turn it into something a product could use.
The data isn't one format, it's several pretending to be one
SEC EDGAR gives you Form 4 filings in a structured XML format — that part's genuinely fine, it's machine-readable and consistent once you know the schema. Congressional trade disclosures are a different world. They're filed as PDFs, sometimes scanned, with formatting that varies office to office and hasn't been standardized. You're not parsing one data source, you're parsing one clean one and one genuinely messy one, and treating them the same will break your pipeline the first week.
Ask AI to find the schema, not guess it
The move that actually worked: feed it real sample filings and ask it to map out every field, every variant of how a field can appear, and every edge case it can find in the samples — not "write me a parser for SEC Form 4," which gets you something that works on the one example you showed it and breaks on the next filer.
Assume every filer formats it slightly differently, because they do
A ticker symbol shows up as "AAPL," "Apple Inc," and a CIK number depending on who filed it. A trade date shows up in three date formats across different congressional offices. None of this is a bug in your code — it's the actual state of the data. Build the normalization step assuming inconsistency exists, don't discover it in production.
Cross-reference before you trust it
If a name, ticker, or amount doesn't reconcile against a second source (I cross-check tickers against a market data source, names against known filer lists), don't silently drop it — flag it for review. Silent failures in financial data are worse than loud ones. A missing trade is bad. A wrong trade shown as real is worse.
The lesson underneath all of it
"Public data" and "usable data" are not the same claim. Anyone building a product on top of public filings should budget real time for exactly this — not model logic, not UI, just getting the raw mess into something you'd trust enough to show a paying customer.
This is the filing pipeline behind Insidr — see it live.
Try Insidr →Brands: partner with me
Keep reading
- Case StudyHow I Built Insidr Using AI as My Only Dev TeamThe Insidr build from the first idea test to Stripe, security bugs, and the first production incident.
- BuildingHow I Use Claude/ChatGPT to Build a Real SaaS SoloHow I go from a blank page to a live product with AI doing most of the heavy lifting.
- PromptsThe Exact Prompts I Use to Go From Idea to Live Product in a WeekendThe prompts I use to go from a rough idea to something working without spending days going in circles.