Back to the Lab
/ ESSAY·8 MIN READ·LONG-FORM
/ LONG-FORM

What Building a Board Deck KPI Extractor Taught Me About Portfolio Health

Extraction isn't the hard part with board decks — trusting the number after is. What broke, what worked, building this for real portfolios.

What Building a Board Deck KPI Extractor Taught Me About Portfolio Health
/ TL;DR

Extraction isn't the hard part with board decks — trusting the number after is. What broke, what worked, building this for real portfolios.

IExtraction was never the hard part

Here is the thing nobody tells you before you build one: pulling a revenue number off a board slide is the easy 20% of the job. Trusting that number — and making it mean the same thing across every company in a portfolio — is the other 80%. That is the whole lesson.

I learned it building our Board Meeting Scanner, the tool that catches inbound board decks in the inbox, runs them through a structured extraction, and posts a Monday-morning partner digest with the KPI trends. The extraction worked early. The trust part took a rebuild. This is what broke, and what actually held up.

IIWhy board decks are the hardest documents in a fund's inbox

A board deck is the worst-case document for any parser. It mixes structured tables with narrative commentary, charts, and screenshots — and there is no standard layout, because every portfolio company builds its deck differently. Portfolio-monitoring vendor Standard Metrics, which does this at scale, calls board decks among the hardest files to parse for exactly that reason: about half of decks embed charts and screenshots, and a single deck can spread revenue, headcount, and runway across dozens of slides with context stretched over 50+ pages.

The chart problem is the sneaky one. As V7 Labs puts it, financial figures are "almost always presented as charts... not as editable text," and a tool that reads only text returns nothing for a revenue trend that lives inside a slide graphic. Half the numbers a partner actually cares about — the growth curve, the burn line, the runway cliff — are drawn, not typed. If your extractor cannot see, it is blind to the most important slide in the deck.

IIIThe manual alternative, by the numbers

The reason anyone builds this is that the manual version is brutal. Here is the time it eats:

And it is not glamorous work. An analyst at a roughly $908M-AUM firm described the status quo to PortfolioIQ: "we're... looking at two different screens, one of the PDF of the board deck, and then we're plugging in the numbers manually... it is cumbersome, and it takes a lot of time." That is a highly paid person retyping numbers off a PDF. Multiply it across a portfolio and you see why the demand exists.

aFailure mode #1: confidently wrong charts

The first thing that broke was trust in the charts. Multimodal models read a chart and hand you back a clean table with the right structure — and several values that are just wrong. A chart-extraction benchmark study caught exactly this testing GPT-4o: the table looked right, but a quick visual comparison turned up values that were clearly off, and the researcher's conclusion stuck with me — "if even I can spot these errors at a glance, how can I trust the rest?"

That is the same failure I wrote about putting AI invoice OCR into production: the model is not wrong often, it is wrong occasionally and confidently — which is worse, because it kills your ability to skim. A board-deck growth chart read 12% too high does not announce itself. It just quietly poisons the digest.

bFailure mode #2: a flag with no memory

The second break was architectural. For a multi-stage European fund, our first version attached a single enriched "performance flag" to each portfolio company — one summary signal pulled from term sheets, investment records, and follow-on notes. It was fast to ship and it demoed beautifully.

Then a partner asked a follow-up question, and it fell apart. The flag had no memory of which round, which document, or which conversation it came from. "Is this the number from the Q2 deck or the one they restated in the update email?" — the system could not say. A number you cannot trace is a number you cannot defend in front of an LP, and a flag that cannot answer the second question is just decoration.

cFailure mode #3: definitional drift is the real KPI problem

Even when the numbers extracted cleanly and traceably, they still did not line up — because portfolio companies do not define their metrics the same way. Burkland calls this definitional drift, and it is the stealthiest way a KPI stack loses its value: "if your CFO, your head of sales, and your pitch deck are all calculating ARR differently, you don't just have a KPI problem, you have a trust problem." The same drift shows up in CAC (does it include brand spend?), gross margin (fully burdened support costs or not?), and runway (gross or net burn?).

This is why generic, off-the-shelf parsing underdelivers here. Standard Metrics is blunt about it: generic models parse data fine, but "they break down when it comes to mapping those parsed values to a firm's source of truth" without a memory and a custom "profile" layer — and without a human in the loop verifying, accuracy stays "far from meeting institutional requirements." Extraction without normalization just produces wrong answers faster.

IVWhat actually worked: the redesign

The rebuild for that European fund threw out the single flag and replaced it with four principles that have held up since:

  • Per-round structured extraction. Every KPI and its surrounding commentary get pulled per round, not flattened into one summary — so history is preserved instead of overwritten.
  • Source-linked numbers. No figure is ever shown without its originating quote or document one click away. If a partner asks where a number came from, the answer is right there.
  • Peer-group and historical baselines. No number is read in isolation — it is always shown against the same company's own history and against same-stage, same-industry peers. A 15% growth month means nothing until you know whether it is up, down, or middle-of-the-pack.
  • A human review queue. Ambiguous or conflicting figures do not get silently resolved. They get routed to a person. Silent auto-resolution is how a wrong number becomes a "fact" nobody questions.

None of these are about better extraction. They are about what you wrap around the extraction so the output is trustworthy.

aWhy the ranking stays behind a curtain

One design decision surprised me, and it was not technical. Once you can rank a whole portfolio on a metric, you can build a stack-ranked league table — genuinely useful for the people making capital-allocation calls. So we built it. But we deliberately tiered who can see it: the full ranking is visible only to fund leadership. The rest of the platform team sees the same underlying metrics and notes, without the ranking layer.

The reasoning was simple, and I would apply it to any internal tool: a ranking that leaks past the people who requested it stops being a management tool and starts being a morale problem. Visibility is a design decision, not an afterthought.

VWhat this means if you're the one reading the decks

If you are a GP, a VP of Operations, or a platform lead weighing build versus buy on this, the failure modes above are your evaluation checklist. Whatever you buy or build, demand four things:

  • Numbers you can trace back to the exact slide or sentence they came from.
  • Per-company definitions reconciled to your fund's own source of truth — not just scraped.
  • A review queue for anything ambiguous, instead of confident silence.
  • Access tiers that match who is actually supposed to see a ranking.

That is the shape our Board Meeting Scanner ended up in — inbound decks caught in Gmail or Outlook, extraction tuned per company because no two decks match, outliers flagged, and the whole thing landing as a partner digest in Slack on Monday morning. It is built for funds with 20+ portfolio companies, where the board-reporting cycle genuinely eats partner time.

And it does pay off. A separate build — a portfolio platform for an anonymized $200M+ climate-tech fund — cut a fund's reporting time by roughly 90%, turning a four-hour spreadsheet exercise into a single Slack question. Different tool, same lesson: the value is not in the extraction, it is in the trustworthy layer you build on top. If you want the honest math on when that is worth a custom build versus off-the-shelf software, I wrote separately about what a build-and-operate retainer actually replaces.

VIFAQ

What KPIs does the Board Meeting Scanner track?

The recurring set is revenue, growth, burn, runway, and headcount — the numbers a partner scans first — with the trends and any outliers surfaced in the Monday digest.

What happens when it isn't sure about a number?

It flags rather than guesses. Outliers and anything ambiguous surface in the digest instead of being silently filled in. The principle across every version I have built is the same: a number the system cannot stand behind should reach a person, not the LP report.

Does it work across different deck formats?

That is the point of tuning extraction per company. Every portfolio company's deck looks slightly different, so the extraction is profiled per company rather than assuming one universal layout. It watches Gmail or Outlook for inbound decks and posts to Slack.

What's needed to set it up?

It is a custom install, not a self-serve signup — built and operated for you on a stack of Gmail or Outlook, Claude, Slack, and a Retool or Airtable dashboard. The setup work is mostly profiling each company's deck and wiring the inbox and Slack digest to how your partners actually review.

Michael Rouveure

/ WORKING WITH BLACK MATTER VC

If this was useful,
you should book a call.

$10k / month. Whatever your fund needs, shipped that month. 30-min intro, no deck — I’ll tell you which three systems I’d ship first.

Or follow along on LinkedIn / X.