Case study · Development of mission critical software systems
Kabuken: a bilingual data platform built on Japanese regulatory filings
The problem
Japanese listed companies publish their financial statements through EDINET, the Financial Services Agency's disclosure system, as XBRL filings. The data is public, but using it is hard: the filings are numerous, the taxonomies are in Japanese, and the same figure can change from one report to the next.
What we built
Kabuken is our own platform for fundamental analysis of Japanese equities. It extracts financial data from EDINET XBRL filings, stores it in structured form, and makes it available in Japanese and English — through a web application, and through an MCP server that lets Claude or any MCP-compatible AI assistant query company filings directly.
We designed, built and operate Kabuken ourselves. Below are two decisions in depth and one short closing note — the kind of engineering we bring to client systems.
Decision 1 — Measuring before spending
Constraint. Our database platform bills every index entry as a write. Each index on a table multiplies the billed cost of every row inserted — and processing financial filings is write-heavy by nature.
What we did. Indexes had accumulated as the schema evolved, which is normal and rarely revisited. Rather than assume they were earning their cost, we audited them: two independent traces through the codebase, plus query-plan analysis against the live schema, established that every reachable query resolved by document identifier, and that only a few indexes were ever actually used. The rest were removed.
Outcome. Billed write operations per document fell sharply, and storage use fell with them.
Why it matters. Nothing was rewritten and no functionality was lost. The cost had simply never been checked against how the data was really queried. That check cost days; the saving is permanent and recurring.
Decision 2 — Filings are history, not a cache
Constraint. Japanese listed companies restate figures between filings. The same financial concept reported in a quarterly report and again in the annual report may legitimately differ.
What we did. Facts belong to the specific filing that reported them. Reprocessing a filing may replace only that filing's own facts. Values are never deduplicated or merged across documents, and a concept is never identified by name alone — always together with the reporting context it was stated in.
Why it matters. The obvious optimisation here is wrong. Collapsing repeated values across filings would shrink the dataset and delete the information users care about most, because the change between filings is itself the signal — it is how a restatement becomes visible at all. Storage is cheap; a destroyed audit trail cannot be recovered.
Closing note — A spend ceiling in front of our own jobs
Every tool that writes to production checks cumulative billed writes against a hard ceiling before it starts, and refuses to run if it cannot verify that figure. A backfill that cannot confirm its own cost does not begin. Bulk reprocessing against a metered database is exactly where unbounded cloud spend comes from.
Working with us
The same engineers work on client systems: architecture consulting and development of mission critical software.