Why I’m Building LedgerBase
Financial data is easy to find.
If I want Apple’s revenue, net income, assets, or EPS, I can get the number from dozens of websites or APIs within seconds.
The problem starts when I ask a different question:
Where did that number come from?
Which SEC filing?
Which period?
Which XBRL fact?
Was the value reported by the company, calculated somewhere in the pipeline, or normalized by the data provider?
And if the number looks wrong, can I trace it all the way back to the original filing?
That is the problem I started running into while building LedgerBase.
The data already exists
Public companies in the US already publish a huge amount of financial data through the SEC.
Most modern filings include XBRL, so at first the problem looks straightforward:
- Download the filing.
- Extract the XBRL facts.
- Store them.
- Expose an API.
I originally expected a large part of LedgerBase to work like this.
It did not take long to find the problems.
A company can report several facts that look like they represent the same metric.
The concept used for a metric can change between filings.
Fiscal periods overlap.
Dimensions can change what a fact actually represents.
Two unrelated rows can contain exactly the same number.
Older SEC filings can have completely different document structures.
Different companies can also report the same financial idea using different XBRL concepts.
So extracting numbers is not really the hard part.
The hard part is deciding which number is the one you should trust.
The API response is only half of the answer
Suppose an API gives me this:
{
"revenue": 391035000000
}
For many applications, that is enough.
Until it isn't.
If the value looks suspicious, I need to know what happened before it reached the API.
I want to know which filing it came from, which reporting period was selected, which XBRL fact was used, and where that value exists in the source document.
Without that information, debugging financial data becomes surprisingly painful.
This gets worse when the data is consumed by another system.
A dashboard may use the number.
Then an analytics job uses it.
Then an AI agent generates a summary based on it.
By the time somebody notices the result is wrong, the original mistake may be several layers away.
That led to one of the main rules I use while building LedgerBase:
A financial metric should keep a connection to the evidence behind it.
Keeping the path back to the filing
LedgerBase stores normalized financial metrics, but I do not want normalization to remove the connection to the original SEC data.
The rough flow looks like this:
SEC filing
↓
Raw document
↓
XBRL facts
↓
Normalized metric
↓
Source evidence
The normalized metric is what an application usually wants.
The source evidence is what I need when I want to verify it.
Both matter.
For example, an application should be able to retrieve a company's financial history and, when necessary, inspect why LedgerBase returned a specific value.
Provenance is not something I want to keep hidden inside the ingestion pipeline.
I want it to be part of the product.
Sometimes the correct result is no result
One thing I did not expect when starting this project was how important refusal would become.
There are cases where the data is ambiguous.
A filing may use a layout the extractor cannot handle reliably.
Several facts may look like valid candidates for the same metric.
A value may appear several times in the document, and there may not be enough evidence to prove which occurrence is the correct one.
In those situations, it is tempting to choose the most likely answer.
For financial data, that can be dangerous.
A plausible number is still wrong if the system cannot justify why it selected it.
So in several parts of LedgerBase, I prefer a fail-closed approach.
If there is not enough evidence, the system should say so rather than quietly guess.
That means there will be edge cases where LedgerBase returns less data than another provider might.
I am fine with that tradeoff.
I would rather know that a value is unsupported than receive a value that only looks correct.
AI makes provenance more important
This also matters for the way financial data is starting to be used with AI.
LLMs are good at working with financial information.
They can summarize filings, compare companies, calculate ratios, explain trends, and turn raw financial statements into something much easier to understand.
But there is an important separation here.
An LLM can explain a number.
It cannot make an unreliable source number reliable.
For example:
Revenue was X.
is useful.
But this is much better:
Revenue = X
→ selected from this XBRL fact
→ for this fiscal period
→ from this SEC filing
→ located here in the source document
The AI can handle interpretation.
The data layer still needs to handle evidence.
I think that separation will become more important as more software starts using AI to analyze financial information.
What LedgerBase is becoming
LedgerBase is a financial fundamentals API built around SEC filings.
The goal is not just to make financial metrics easy to retrieve.
I want developers to be able to retrieve normalized data and still have a path back to the original evidence when they need it.
The project currently covers areas such as:
- normalized financial statements
- historical metrics
- SEC filing metadata
- metric provenance
- source explanations
- links back to the original document evidence
There are still plenty of difficult parts.
Historical filings are inconsistent.
XBRL normalization has more edge cases than I expected.
And mapping a normalized metric back to the exact row and value in an SEC document turned into a much larger problem than simply parsing the filing.
Those problems are also the reason I decided to start writing about the project.
Writing about the parts that went wrong too
I do not want this blog to be a list of product announcements.
I want to document the engineering problems behind LedgerBase.
That includes things like:
- how I normalize SEC XBRL data
- why one metric can map to multiple facts
- how financial statements are reconstructed
- how source lineage works
- why highlighting the exact source value is difficult
- how I test whether a selected value is actually correct
- where deterministic logic works well
- where AI is useful, and where I do not want it making decisions
Some of these posts will be about solutions that worked.
Some will probably be about approaches that looked reasonable and failed.
Both are useful.
LedgerBase is still being built, and there are still areas I am not satisfied with.
But the idea behind it has become much clearer while working on the system:
Financial data should not just be easy to query. You should be able to prove where it came from.
Comments
No comments yet — be the first to comment.