< Date >
< Customer >
A short pre-read for our call

Ship AI agents you can trust

TL;DR  On Harvey LAB, we optimized a DeepSeek V4 Flash verifier:
as accurate as the default per-criterion grading, about 430x cheaper.
Ahmad Beirami · Co-founder / CEO
© 2026 Fidian, Inc · Confidential · Do not distribute

We have deep expertise in
frontier model research & enterprise AI

Fidian is a research-driven product company with decades of combined experience building/shipping AI.

Google DeepMind
Gemini model alignment via controlled decoding and InfAlign; 2x inference speedups via speculative decoding
Meta
Led cross-org prototype of first tool-using LLM agent (2020); responsible AI research, adversarial evaluation, red teaming
Apple
A decade-plus of AI/ML engineering building production-scale ML systems
Electronic Arts
Delivered first RL-driven play testing (bug discovery, balance tuning) plus production ML engineering
Palo Alto Networks
Architected AIOps for NGFW (now Strata Cloud Management); 30+ years backend systems, SRE leadership
Cisco Tetration Analytics
First enterprise-grade data pipelines, ML-driven dependency mapping, regression frameworks
Educational background (PhD) Georgia Tech MIT University of Michigan USC Harvard University Duke University
© 2026 Fidian, Inc · Confidential · Do not distribute 2

The future of development

Fidian builds a self-evolving agent platform by solving reliable verification first

1 · Verifiers

Frontier-level accuracy at a fraction of the cost.

Today

2 · Maintainable evals

Your evaluations regenerate from your product specs, so they stop rotting as your product evolves.

3 · Harness learning

Your harness improves automatically against those evaluations without overfitting. The optimized harness is distilled into a new model, and the loop continues.

Fidian’s self-evolving agents platform

You maintain the spec; the platform evolves evals and the harness.

© 2026 Fidian, Inc · Confidential · Do not distribute

Why verification first

A flawed verifier
lets a worse agent outscore a perfect one

Interactive · dragHarvey LAB · 1,251 tasks · all-pass scoring (all criteria must pass) · simulated error sweep

verifier accuracy = 100%
Perfect agent (true 100%) Bad agent (true 20%) true all-pass rate

Optimizing against flawed verification leads to reward hacking.

© 2026 Fidian, Inc · Confidential · Do not distribute

Measured results

We reached near-ceiling accuracy
at a fraction of the cost

Verifier quality: share of real failures caught vs share of correct work accepted

9596979899100 80859095100 Real failures caught (%) Correct work accepted (%) ideal ↗ Per-criterion · Sonnet Batch · DeepSeek V4 Flash Per-criterion · DeepSeek V4 Flash Optimized · DeepSeek V4 Flash

Verifier cost per 1,000 criteria (USD)

$0$20$40 Batch · DeepSeek Optimized · DeepSeek Per-criterion · DeepSeek Per-criterion · Sonnet $0.08 $0.10 $1.63 $43
As accurate as the default per-criterion grading at about 430x lower cost.

Verifiers measured against silver labels (frontier consensus) on Harvey LAB.

© 2026 Fidian, Inc · Confidential · Do not distribute
Ship in your VPC, with privacy and security Close support from our AI Research Engineering team Shape the product design based on Harvey AI’s needs

Collaboration proposal

01

We improve Harvey’s evaluation lifecycle: faster, cheaper, more accurate.

02

This reliable signal enables model post-training and learning better harnesses.

03

Harvey’s product itself becomes faster, cheaper, and more accurate.

© 2026 Fidian, Inc · Confidential · Do not distribute