Introducing

Kepler Builds Provably Correct AI for Financial Analysts with Massive

Aug 20, 2026

About Kepler.ai

Kepler is a provably correct AI platform, and its first users are financial analysts: buy-side analysts at hedge funds and PE firms, sell-side and independent equity research analysts, and bankers building comps. For that audience, one wrong number can carry multi-million-dollar consequences, so a single hallucinated figure forces an audit of the whole dataset and erases any time the tool saved. Kepler's promise is AI that proves it's right. Its answers are deterministic and every figure traces back to source, which puts a strict requirement on the data underneath. Kepler standardized on Massive for prices, splits and dividends, earnings, and news, because the series is clean, replayable, and source-attributable.

Kepler.ai

Separating the Model From the Data

An analyst presents a number to a PM, and the PM asks where it came from. If the analyst can't show the trace in one click, the number doesn't go in the model. AI tools that can't survive that handoff don't get used.
- Rui Zhang, Founding Product Lead

Kepler's promise has two pieces. Every number, metric, or data point behind an answer links back to the exact source it came from: a specific 10-K line item, an analyst estimate, a price tick. And the work that produced that number is deterministic. Same question, same data, same output, every run.

The architecture makes that guarantee structural, not aspirational. Kepler separates the AI layer from the data retrieval layer by design. If you let a model retrieve and compute directly, it will sometimes invent the source, pull the wrong filing, or miscalculate. Worse, it will be confidently wrong and tell the user it checked something it cannot actually check. Each of those is a credibility-destroying error in finance.

So Kepler splits the work. The model handles intent: what is the user asking, which entity, which period, which metric. Once intent is locked, Kepler runs the retrieval and the math against their semantic layer, and the answer carries the trace. There is no path in the system where a model invents a number.

The outcome that matters most to Kepler's team isn't any single endpoint. It's that an analyst can show a PM where every number came from in one click.

Key takeaways

  • A Provable Answer Needs a Provable Data Layer: Kepler's answers are deterministic and traceable to source, so the market data underneath has to be deterministic and traceable too. Massive supplies a clean, replay-able price series Kepler can point a customer directly back to.
  • Bad Data Doesn't Bias One Number, It Propagates: If a price is off by a cent or a corporate action isn't applied, the EV calculation is wrong, the P/E is wrong, the comps table is wrong, and every conclusion above it loses credibility. Kepler accepts data only on a deterministic, source-attributable basis, and Massive cleared that bar across thousands of automated and manual tests.
  • One API Pattern Across a Growing Surface: Prices, splits and dividends, analyst estimates and actuals, and news all follow the same Massive request and response conventions, which let a small team integrate fast and expand coverage without rebuilding the foundation.

Why a Provable Answer Needs a Provable Source

SEC filings, earnings transcripts, and regulatory documents are the core of Kepler's document corpus. A 10-K tells you a company's history. A transcript tells you the qualitative story. Neither tells you what the stock did on the day of the print, what the EV/EBITDA multiple is right now, or what the historical multiple band looks like.

That is where Massive comes in. Kepler needs a clean price series and accurate quotes to build comps, EV, P/E, and dividend yield, and to anchor any valuation conversation in market reality. The transcript and the 8-K provides the raw context, while the pricing data provides the market's reaction to that context.

Because every figure Kepler returns has to be reproducible, the requirement that flows down to the data layer is strict. If a price is wrong by a cent or a corporate action isn't applied, the error propagates through every calculation above it. So Kepler accepts market data on three conditions: it has to be a deterministic series the team can replay, they have to be able to point a customer at exactly where it came from, and they have to have full confidence in its reliability.

Before settling on Massive, Kepler evaluated providers against three filters: a clean API a small team could integrate quickly, single-vendor breadth across the products they needed (prices, corporate actions, earnings, news), and a partnership posture that matched startup pace rather than enterprise procurement.

Massive fit. Finance has a high-trust bar, so anything sitting under Kepler's outputs has to clear Kepler's bar. The proof was in the testing. The early days were a barrage of checks: thousands of automated and manual tests against the data, splits and corporate actions, edge cases on tickers with messy histories. Kepler kept building on Massive because everything passed.

We have a high bar on the output of the data. The two things that matter to us are that the data passes our test suite, and that the engineering process on the other end is at the same bar as ours, so we have no concerns with reliability or consistency.
- Susannah Meyer, Founding Engineer, Kepler

Benzinga Estimates in Kepler.ai, via Massive

Building Kepler's Data Layer on Massive

Market data is fetched at query time, when an answer depends on a current price or a price series. Three categories of Massive data flow into both Kepler's analytical workflows and its Company Profile UI, and every value is attributed back to source, consistent with how Kepler traces everything else.

  • Prices, Splits, and Dividends: Raw market data powers the price series behind comps tables, multiple bands, and every chart Kepler renders. Splits and dividends carry the adjustment factors Kepler needs to normalize a historical series so the math stays correct across corporate actions.
  • Earnings via Benzinga: Massive's Earnings API returns reported figures alongside analyst estimates and surprise metrics, which powers Kepler's actual-versus-estimate analysis after an earnings print.
  • News and Earnings Coverage: This powers the news feed on the Company Profile and the article previews surfaced in chat citations.
  • Source Attribution as a First-Class Output: Retrieved values render into the answer with a source tag. The user clicks through to the underlying series, the timestamp, and the vendor of record. That attribution is the whole point. It's how an analyst defends the number.
  • One HTTP Pattern Across the Surface: Prices, corporate actions, earnings, and news share request and response conventions, so one client and one integration layer cover every product. Adding coverage is configuration, not a new vendor.

Retrieved values are rendered into the answer with a source tag. The user can click through to see the underlying series, the timestamp, and the vendor of record.
- Susannah Meyer, Founding Engineer, Kepler

Benzinga Analyst Ratings data in Kepler.ai, via Massive

Results: What the Data Layer Guarantees

What Kepler can stand behind:

  • Every Number Traces to Source: Retrieved values render with a source tag, and the analyst can click through to the underlying series, timestamp, and vendor of record. The number survives the handoff to a PM.
  • A Deterministic, Replayable Series: Kepler can replay the same series and get the same result, so models and backtests line up with the filings and transcripts in the rest of the corpus. The analyst compares like to like.
  • Corporate Actions Applied Correctly: Splits and dividends carry the adjustment factors, so historical math holds across corporate actions and tickers with messy histories.
  • Massive Cleared Kepler's Test Bar: Thousands of automated and manual data-quality tests passed, spanning splits, corporate actions, and edge cases on difficult tickers.
  • One Vendor, Room to Grow: US equities today, with options, futures, and crypto reachable through the same schema-consistent layer instead of a new vendor per market.

The result isn't a throughput number. It's that Kepler never has to ask an analyst to trust a figure it can't show them the source for.

Benzinga's Bulls Say Bears Say in Kepler.ai, via Massive

What's Next

Kepler's edge isn't a load number. It's that no answer rests on a figure the analyst can't trace. For Kepler's architecture, historical depth matters more than latency. Users build models and run backtests, studying trends over varying periods and market conditions, so a clean, deep series that replays deterministically and lines up with the filings and transcripts is what the product depends on.

Today Kepler's Massive usage is focused on US equities. As customer demand pushes Kepler into new markets, the team can extend coverage through the same data layer rather than onboarding a new vendor each time. Options, futures, and crypto are anticipated down the line, and Kepler made the single-vendor, schema-consistent bet with that growth path in mind. The ceiling is set by what the team can build, not by what the data layer can afford to serve.

Look for someone who can match your bar. The data has to pass your test, every series, every edge case, every corporate action applied correctly. And the people on the other end have to pass it too. They have to know the data, respond quickly, and care when something is off. Most vendors fail one of those two. Pick the one that does both. 
- Rui Zhang, Founding Product Lead, Kepler

If your team is building an AI finance product, a research platform, or a trading tool that depends on accurate market data you can trace back to its source, reach out to sales@massive.com. We would be happy to learn about your project and help you build on the same foundation powering Kepler. You can also get started at massive.com and explore the API at massive.com/docs.

Keep building something massive.

From the blog

See what's happening at Massive