Skip to main content

Startups I Want to See #1

The AI-Native Hedge Fund


20269 min readRead on Substack ↗

Disclosure: I was the youngest employee at Balyasny Asset Management, and I scout for venture funds. I want someone to build this. If you are that person, message me on LinkedIn. Nothing here is investment advice and nothing here is specifically directed at the mentioned funds, including any past firms I worked at.

Every fund now claims to use AI. Most of what that means in practice is ML-assisted feature engineering inside a human research loop: a researcher has an idea, an embedding model helps parse transcripts, gradient boosting ranks the signal, and a human still sits at every decision point. The iteration cost of that loop, from hypothesis to reviewed backtest, is measured in researcher-weeks.

The fund I want to see inverts this. Agents own the research loop. Humans own capital, compliance, reward specification, and the off switch. And because "AI-native multi-strat" is not a fundable pitch for a startup with no track record, I will start with the wedge.

Part 1: A replaceability index for public SaaS

The market agrees AI compresses software margins and cannot agree on which tickers. That disagreement is the opportunity, and the instrument to price it does not exist yet. So build it:

Take the universe of public SaaS companies. For each one, decompose the product into its feature surface, then hand a coding agent a fixed compute budget, say 48 agent-hours, and let it attempt a functional clone. You are not shipping the clone. The clone is a measurement. A judge model scores coverage against the real product, and the residual is where it gets interesting, because the residual is a decomposition of the moat: SOC 2 and FedRAMP certifications, integration depth, data gravity, switching costs embedded in workflows. The stuff a coding agent cannot replicate in 48 hours is, definitionally, what the company is worth.

Publish a Replaceability Score per ticker, re-run quarterly. The level is somewhat informative. The first derivative is the product. When a name goes from 40 to 75 in two quarters because coding models improved, the market is watching model benchmarks and you are watching that company's moat evaporate in a controlled experiment.

Two monetization paths, in order. First, trade the signal market-neutral, short high-and-rising replaceability against low-and-stable, sector and beta hedged, which also gives you the live track record. Second, once the index has public credibility, wrap it in an inverse ETF. Retail demand for "short legacy SaaS" is real and currently has no clean expression. Single-name shorting has borrowing costs and unbounded downside that make it wrong for retail; a rules-based product with a published methodology is the honest version of the trade.

The wedge logic: this competes with nobody on day one, the artifact markets itself, and one correct out-of-sample call, flagging a moat collapse two quarters before the miss, is worth more than any pitch deck.

Part 2: Closing the research loop

The longer-term thesis rests on one identity that every quant PM has internalized: the fundamental law of active management. Your information ratio is roughly your skill per bet times the square root of how many independent bets you make. IR ≈ IC × √breadth. My best Citadel PM mentors are generally only right 53% of the time. Renaissance did not win by having a higher IC than everyone else on any single trade. They won on breadth, thousands of small, weakly correlated, high-turnover positions where a 51/49 edge compounds into the Medallion track record.

Breadth is bounded by research throughput, and research throughput is bounded by headcount, and headcount is expensive and slow. That constraint just broke. A hypothesis-backtest-review cycle that costs a pod a week costs an agent system minutes and single-digit dollars of inference. Four orders of magnitude on iteration cost does not give you the same research faster. It changes what regions of hypothesis space are economical to search at all. Hypotheses with a 2% prior of working were never worth a researcher-week. At $3 per loop, you test all of them.

The naive version of this dies immediately, and it dies of multiple hypothesis testing. Run a million backtests and thousands will clear any Sharpe threshold by chance; this is the false discovery problem that Bailey and López de Prado spent a decade formalizing (deflated Sharpe ratios, combinatorial purged cross-validation, the whole apparatus exists because human quants were already fooling themselves at a few hundred backtests a year). An agent system that generates hypotheses at scale without an equally scaled falsification apparatus is a machine for manufacturing overfit.

So the architecture that matters is generator-verifier, and the asymmetry runs in your favor: verification is easier than generation. The system needs a hypothesis agent searching over public data (filings, transcripts, hiring flows, app telemetry, card panels), a backtest agent that enforces point-in-time data hygiene, and then the component that is the entire company: an adversarial reviewer rewarded purely for killing ideas. It hunts for lookahead leakage, regime dependence, capacity limits that make the signal untradeable after costs, and crowding. Above all of it, a meta layer treats research allocation as a bandit problem, shifting compute toward the regions of hypothesis space where survivors cluster. That reallocation loop is the recursive part, and it is also where Goodhart lives: agents optimizing backtest statistics will find degenerate solutions unless the reward specification is treated as the core engineering problem rather than a config file. Anyone who has trained RL systems knows this failure mode arrives on day one, not eventually.

I will be honest about the objection I cannot fully answer. If every fund's agents run on the same three foundation models, the equilibrium looks bad: shared priors, correlated discovery, alpha decaying at the speed of an API rollout. My partial answer is that edge migrates to the scaffolding, the proprietary falsification data you accumulate, and the meta-learned search policy, none of which ship with the base model. But I hold this at maybe 70% confidence, and if you have a sharper answer, that answer is your seed round.

The part you cannot do

One extension of this idea keeps coming up when I describe it to people, so let me kill it in writing: agents generating organic-looking hype on Reddit and X to pump your own ETF to retail. Undisclosed promotion of a product you profit from is the core thing the SEC prosecutes, and laundering it through an agent changes nothing about liability except adding an intent trail in your logs. The compliant version works better anyway: publish the methodology, publish the scores, and let a correct call do the marketing.

Shape of the firm

Humans: capital, legal, counterparties, reward specification, kill authority. Maybe twelve people. The regulatory surface is the underrated constraint here, because "the agents decided" is not an acceptable answer to an SEC examiner, which makes interpretability a licensing requirement rather than a research luxury. Position-level attribution needs to trace back to a stated, testable hypothesis. This is annoying and it is also, quietly, a feature: a fund that can explain every position is running better risk than most humans-only funds I have seen.

One more objection I like rather than fear: if the replaceability index works and gets famous, SaaS companies will manage to it, hardening the parts of the product the builder agent can clone. Good. An index that changes corporate behavior is an index the market prices. Ask MSCI.

If you are building this

Sequence: index first, published free, six months of public out-of-sample calls. Then trade it. Then raise for the engine. If you are anywhere on this path, message me on LinkedIn. I scout for venture funds and I would like a reason to make an introduction.

Building this?

If you are anywhere on this path, reach out. I scout for venture funds and would like a reason to make an introduction.