· Venture Capital Tracker Research · venture-capital
Preference Model Raises $16M Seed for Frontier-AI Training Environments
Preference Model raised a $16 million seed led by Andreessen Horowitz to build reinforcement-learning environments and difficult technical data for frontier AI labs.
Preference Model has emerged from stealth with a $16 million seed round led by Andreessen Horowitz. SignalFire, South Park Commons and Scale Angels participated, alongside angel investors including Fei-Fei Li, Ian Goodfellow and Julian Schrittwieser.
The company builds reinforcement-learning environments and evaluation infrastructure for frontier AI laboratories. It also open-sourced Karotte, a framework for constructing and grading environments designed to resist reward hacking.
The financing at a glance
| Item | Detail |
|---|---|
| Company | Preference Model |
| Financing | $16 million seed |
| Lead investor | Andreessen Horowitz |
| Institutional participants | SignalFire, South Park Commons and Scale Angels |
| Selected angels | Fei-Fei Li, Ian Goodfellow and Julian Schrittwieser |
| Valuation | Not disclosed |
| Founders | Jennifer Zhou and Ning Cao |
| Product focus | RL environments for AI research and ML engineering |
The bottleneck is moving from labels to environments
The first generation of AI-data businesses organized large volumes of labeled text, images and code. Frontier laboratories now need something harder: expert tasks, executable environments and grading systems that can distinguish a genuinely capable model from one that found a shortcut.
Preference Model focuses on machine-learning engineering and AI research tasks. These can involve debugging training runs, optimizing kernels, curating data or designing experiments. They are long-horizon problems with many possible paths and many opportunities for an agent to exploit a flawed grader.
That changes the supplier's role. The job is no longer simply to recruit experts and collect answers. It is to build reliable software environments, define verifiable outcomes and attack the evaluation system before a model does.
Karotte and the reward-hacking problem
Karotte is the framework Preference Model says it uses in production to build its environments. Andreessen Horowitz says it has been hardened through more than one million evaluation runs and controlled red-teaming.
The framework includes defenses such as killing stray processes before grading and rejecting files designed to crash the grader. These details may sound operational, but they address a fundamental reinforcement-learning problem: a model can maximize reward without completing the intended task if the environment exposes a loophole.
Open-sourcing Karotte serves two purposes. It gives researchers a tool they can inspect and use, and it creates a distribution channel for Preference Model's paid environment-building work. The tension is that open source also lowers the barrier for labs and rivals to reproduce some of the infrastructure.
Founder-market fit
Co-founder Jennifer Zhou was an early Anthropic employee who worked on pretraining data infrastructure, tokenizers and Claude datasets. Co-founder Ning Cao was an early employee at DatologyAI.
That experience is relevant because the company's customers are technically sophisticated and often capable of building internally. Preference Model must provide more than outsourced labor. It needs research judgment, engineering speed and a growing library of adversarially tested environments that would be expensive for each lab to recreate.
A competitive expert-data market
Preference Model overlaps with Scale AI, Surge AI, Turing and a growing group of specialist data providers. It also competes with the frontier labs' own data and evaluation teams.
The market is attractive but concentrated. A small number of model developers account for much of the demand, giving customers negotiating leverage and creating material revenue-concentration risk.
Potential defensibility can come from:
- proprietary environments and difficult task libraries;
- trusted relationships with frontier labs;
- expert contributor networks;
- grading reliability and red-team data;
- tools that identify model weaknesses and generate targeted tasks; and
- operational knowledge accumulated across millions of evaluations.
The strongest companies in this category will likely look less like labeling marketplaces and more like research-infrastructure partners.
What competitors covered
Search results for “Preference Model funding” currently surface company databases, the Wilson Sonsini financing notice, Andreessen Horowitz's investment thesis, Dealroom and smaller funding trackers. Most focus on the investor list and the Karotte launch.
The missing analytical angle is procurement power. Frontier labs urgently need high-quality environments, but they are also the buyers most able to internalize successful methods. Preference Model's long-term value will depend on staying ahead of its customers technically while serving them operationally.
What the seed round must prove
The $16 million financing gives Preference Model capital to hire researchers and engineers, expand its environment portfolio and support more evaluation runs. The most important evidence of progress will be:
- repeat and expanding work with frontier labs;
- environments that remain useful as model capability improves;
- independent adoption of Karotte;
- measurable resistance to reward hacking; and
- lower dependence on any single customer.
Bottom line
Preference Model's financing is a confirmed $16 million seed equity round. Andreessen Horowitz led, with SignalFire, South Park Commons, Scale Angels and prominent AI researchers participating. No valuation was disclosed.
The round is smaller than the headline financings flowing to model developers, but it targets an enabling layer those labs increasingly need: difficult, verifiable training grounds for models that are already good at finding easy answers.
Editorial note: AI tools assisted with research, structure, or drafting. Venture Capital Tracker retains human editorial responsibility for factual accuracy, relevance, and source quality before publication.