Round operations

Why AI-generated investor lists are usually wrong

A generated investor list is a candidate pile until it is validated against your actual round.

Jun 25, 20269 min readRound operations

The list looked right for about four minutes

You ask the model for "50 seed investors for an AI infrastructure startup." It returns a clean table in seconds. Names, firms, focus areas, check sizes, a tidy "why they fit" column. It reads like research. You paste it into a spreadsheet and feel ahead of schedule.

Then you start checking rows. The first fund stopped writing seed checks two years ago and only does Series B now. The second partner left to start their own fund. The third "AI infra" match is a firm that did one ML deal in 2021 and writes about it constantly. The fourth has a portfolio company that competes directly with you, which means a conflict, not a lead. By row ten you are doing the actual work the model skipped: figuring out whether each name is real for your specific round, this quarter, given who you are.

The list was never wrong because the model is dumb. It was wrong because you asked it to do the one part of list-building that requires information it does not have: your company, your constraints, and your judgment about fit.

What the model is really doing under the hood

A language model generating an investor list is doing pattern completion over public text. It has seen thousands of "top seed investors" posts, fund websites, and portfolio pages. When you ask for investors in a category, it returns the names most associated with that category in its training data.

That produces three predictable failure modes.

Stale fit. The model's picture of a fund is an average of everything it ever read about them. Funds change strategy, stage, and check size faster than that picture updates. A firm that was the canonical seed AI fund in 2022 may be doing growth now. The model will still hand them to you for a seed round.

Surface fit. "AI" or "fintech" in a focus line is a keyword, not a thesis. The model matches keywords. It cannot tell the difference between a fund whose entire conviction is your exact market and a generalist who lists twelve sectors on their homepage. Both come back looking equally relevant.

Invented fit. The most dangerous one. Asked to justify each name, the model writes a plausible "why they fit" sentence whether or not it is true. It will confidently tell you a partner "focuses on developer tools and AI infrastructure" because that sentence is statistically reasonable for a VC, not because it checked. This is where founders get burned: the reasoning looks like diligence and is fluent guessing.

None of this means generation is useless. It means generation is step one of four, and most founders ship step one as if it were the finished list.

The reframe: generate candidates, then earn the target list

Stop thinking of the AI as a list builder. Think of it as a candidate generator. A candidate is an unverified name that might belong on your list. A target is a name you have checked against your round and decided to pursue, in a specific order, for a specific reason.

The work that turns candidates into targets is not generation. It is three things the model cannot do alone:

  1. Constraints come from you. Stage, check size, geography, lead vs. follow, sector conviction, conflicts, and any "never pitch" names. The model has none of this unless you give it, and most prompts give it almost nothing.
  2. Enrichment comes from current sources. What this fund did in the last 12 months, who the live partner is, whether they are deploying right now or paused. This is fresh data, not training data.
  3. Taste comes from you again. Given two plausible fits, which one matches how you want to build, who you would want on your cap table, and where a warm path exists.

Generation is fast and cheap. The other three are where the round is won. Inverting that ratio, spending minutes generating and seconds verifying, gets it backwards.

Required inputs before you generate anything

Most bad lists trace back to a thin prompt. "Seed investors for an AI startup" gives the model nothing to constrain on, so it falls back to fame. Before you generate, write down your round's constraints. This is the spec the candidates have to satisfy.

Template
ROUND CONSTRAINTS (fill before generating)

Stage:                e.g. seed, raising now
Round size:           e.g. $2.5M
Check size wanted:    lead $1M+ / follow $100–500k
Lead or fill:         need 1 lead, then fills
Geography:            where partners must be able to invest
Sector thesis:        the specific belief, not the keyword
                      e.g. "infra for AI eval, not generic ML tooling"
Stage rule:           must write FIRST seed checks (not seed-prepared Series A)
Conflicts:            funds backing [competitor], do not pitch
Hard nos:             [names you will not approach and why]
Proof you have:       traction/team facts that make a fund a real fit
What you want from
the investor besides
money:                intros, recruiting, design-partner access

That block is the difference between "name famous AI funds" and "name funds that write first seed checks of $1M+, in my geography, with live conviction in my exact thesis, no competitor conflict." The second prompt produces far fewer candidates and a far higher hit rate.

A prompt structure that produces checkable candidates

The goal of the prompt is not a finished list. It is candidates formatted so you can verify them fast, with the model forced to admit what it does not know.

Template
You are helping me build a CANDIDATE investor list. These are not final picks.

My round constraints:
[paste the ROUND CONSTRAINTS block above]

Generate up to 30 candidate investors that plausibly match.

For each candidate return columns:
- Name (person) and firm
- Why plausibly a fit (one specific sentence)
- Stage they're associated with
- Confidence: HIGH / MEDIUM / LOW that this is current
- What I must verify before trusting this row
- Source type you're inferring from (portfolio page / news / general knowledge)

Rules:
- Do NOT invent check sizes, fund dates, or partner names you are unsure of.
- If you are guessing, mark Confidence LOW and say so.
- Prefer leaving a field blank over filling it with a plausible guess.

This does two useful things. It makes the model rank its own reliability, which surfaces the rows worth your time. And it converts the "why they fit" column from confident prose into an explicit to-verify list, so you stop reading guesses as facts.

False-fit examples: what a wrong row looks like

These are the patterns that pass a skim and fail a check. Learn to spot them in your own generated list.

Candidate as generatedWhy it looks rightWhy it's a false fit
"Partner X, focuses on AI infrastructure"Keyword match to your sectorPartner X left the fund 8 months ago. The model is averaging old bios.
"Fund Y, seed-stage AI investor"Reads like your stageFund Y raised a $400M fund and now leads Series B. "Seed" is residual reputation.
"Fund Z, invests in developer tools"Plausible adjacencyZ's only dev-tools deal was 2020. Their live thesis is consumer. Surface keyword, no conviction.
"Partner W, backs technical founders"Flattering and genericTrue of every VC alive. Says nothing about fit for your round.
"Fund Q, AI portfolio includes [competitor]"Strong sector signalThis is a conflict, not a lead. They will pass and may share notes.
"Angel V, writes early checks in AI"Right stage, right spaceV stopped angel investing after going GP at a fund with a no-angel policy.

The shared tell: every false fit is built from a true-sounding fragment that is stale, generic, or pointed the wrong way. The model is not lying. It is completing a pattern, and a pattern can be confidently out of date.

The validation workflow

Run every candidate through the same five checks before it earns a place on the real list. This is the part that protects your outreach.

Template
VALIDATION CHECKLIST (per candidate)

[ ] 1. PERSON IS LIVE
       Partner still at the firm? Check the firm's team page or their
       own profile, not the model. Departed partner = delete row.

[ ] 2. STAGE IS CURRENT
       Did they lead/write a check at MY stage in the last ~12 months?
       One recent stage-matched deal beats a reputation. No match = demote.

[ ] 3. THESIS IS REAL, NOT KEYWORD
       Find one concrete signal they care about my exact space:
       a relevant portfolio company, a written thesis, a public take.
       Keyword-only = LOW priority.

[ ] 4. NO CONFLICT
       Do they back a direct competitor? If yes, it's a conflict, not
       a lead. Move to a separate "do not pitch" list.

[ ] 5. WARM PATH OR HONEST COLD
       Is there a real intro path? Score it. No path is fine, but then
       the row needs a strong cold reason to exist.

SCORE EACH: keep / demote / cut
Only "keep" rows with a stage + thesis match become the target list.

A candidate that clears all five is a target. A candidate that fails stage or has a conflict is not a "maybe later," it is a cut. The discipline is treating the checklist as a gate, not a suggestion.

Score the output so the list ranks itself

Once candidates are validated, score them so your outreach order is obvious instead of arbitrary. A simple additive scorecard works.

SignalPoints
Wrote a stage-matched check in last 12 months+3
Specific, current thesis in your exact space+3
Warm path exists (and how warm)+1 to +3
Right check size / can lead if you need a lead+2
Public signal they're actively deploying now+1
Conflict with a portfolio competitordisqualify
Partner departed / fund changed stagedisqualify

Sort descending. The top of that list is where your first, best, freshest warm intros go. The bottom is where cold outreach goes if you have capacity. The model gave you none of this ordering, because ordering depends entirely on facts about your round that live in your sources, not its training data.

Where this connects to RoundOS

The reason raw generation fails is that the three steps that matter, constraints, enrichment, and ranking, all depend on context the model never sees: your stage, your conflicts, the partner who replied last week, the intro path sitting in a cofounder's inbox.

RoundOS works the other way around. You upload the sources where your round already lives: your investor spreadsheet, email and calendar, meeting notes, LinkedIn exports. It enriches each fund and partner against current context, maps the warm paths you have, and ranks next moves by fit instead of by fame. Generation becomes one input, not the answer. You generate candidates, then enrich and score them against your round so the names that reach your outreach queue are the ones that survived a check, in the order that respects your warm intros.

The model names people. Your context decides who is real. Keep those two jobs separate and the list stops embarrassing you in row ten.

Generate candidates, then earn the target list.

Validate the top rows against your round before sending a single intro request or cold email.