Lattice

llattice MODEL BENCHMARKING PLATFORM
THE QUESTION STATIC LEADERBOARDS CANNOT ANSWER

What is the best model for this workload,
under these constraints, right now?

Lattice turns model evaluation from an occasional research project into continuous infrastructure — designed around real workloads, not synthetic leaderboards, in a world where models, providers, and prompts change every week.

DESIGNMEASURECOMPAREDECIDEREPEAT
// not the center of the universe

Not the leaderboard the world uses.
The benchmark your workload needs.

Lattice is a private benchmarking platform. Every org — and every product inside every org — runs its own suites, on its own workloads, with its own graders and constraints. It's the cheapest, fastest, and most accurate way to know which models actually work for your systems.

PUBLIC LEADERBOARD
#1
one number.
for the whole world.
  • Someone else's test set
  • Someone else's constraints
  • Weeks-old at best
  • Can't be verified against your traffic
YOUR LATTICE — private, per product
acme / checkout-agent
winner · flash-2 @ 0.4¢ / task
4,988 cases · updated 12 min ago
acme / support-bot
winner · core-8 @ p95 1.4 s
2,102 cases · updated 3 h ago
widgets-co / research-agent
winner · max-large @ quality 0.94
812 cases · updated 41 min ago
widgets-co / ocr-intake
winner · self-hosted q4 @ 0.09¢
multimodal · 1,240 cases
Every product picks the model that actually wins its job.
CheapestAdaptive testing skips obvious losers — you don't pay to re-prove them.
FastestA model released this morning can be evaluated on your workload before lunch.
Most accurateYour traces. Your graders. Your constraints. No stand-ins.
// benchmark designer

Benchmarks built around real work — not a generic test suite.

Design in minutes. Run across any model or provider. Version, fork, compose, tag, and reuse test suites across teams. A benchmark can be twenty carefully chosen examples or millions of automated trials.

TEST CASES IT COMBINES
  • 01Prompt & response cases
  • 02Multi-turn conversations
  • 03Tool-using agent tasks
  • 04Structured-output requirements
  • 05Coding & execution tasks
  • 06Retrieval & research workflows
  • 07Image & multimodal
  • 08Long-context tests
  • 09Adversarial & edge cases
  • 10Production traces (anonymized)
  • 11Synthetic test-set generation
  • 12Parameterized dynamic cases
SUITE / checkout-agent-v14
agenthappy-path checkoutmulti-turn · 4 tools · 128 cases
edgeexpired card fallbackmulti-turn · 3 tools · 48 cases
prodreplay: apr traces2,412 sanitized sessions
codestripe intent creationstructured output · 320 cases
advprompt-injection sweepadversarial · 640 cases
imgreceipt OCR extractionmultimodal · 200 cases
genparameterized · region × tiergenerates 1,240 cases
TOTAL 4,988 CASES VERSION v14 · FORKED FROM v12 TAGS checkout, agent, prod
// evaluation

Some things are numbers.
Some things are preferences.

Lattice supports both — without pretending a piece of writing deserves an arbitrary 8.3/10.

RATINGSobjective & quantitative
Accuracy0.923
Pass / fail0.860
Exact match0.712
Schema compliance0.984
Tool-call correctness0.809
Execution success0.881
Latency (p95)2.14 s
Cost / task$0.007
Refusal rate0.121
Judge score0.764
RANKINGSpairwise & preference
image · vibe
R>C
341 pairs
creative writing
R<C
208 pairs
UI generation
R>C
184 pairs
tone & style
R>C
276 pairs
summarization
R~C
tie · 162 pairs
research quality
R<C
98 pairs
AGGREGATED (Bradley-Terry)
1Model R1642 elo
2Model C1584 elo
3Model M1531 elo
4Model A1489 elo
Blind evaluation prevents model identity, provider, price, or reputation from contaminating judgment.
// pareto frontier

There isn't one best model.
There's a frontier.

Lattice measures the complete execution profile — quality, latency, tail latency, cost, throughput, reliability — and identifies which models dominate different portions of the trade space under your constraints.

tiny-3 flash-2 core-8 max-large $0.001 $0.01 $0.05 $0.25 / task 1.0 0.7 0.4
QUALITY →
COST PER TASK →
on frontier dominated frontier
// what you didn't know you wanted

Lattice benchmarks the benchmark.
And every model becomes a fingerprint.

A leaderboard reduces a model to a position. Lattice builds a multidimensional profile instead — and it watches the tests themselves for staleness, redundancy, contamination, and unstable graders.

CAPABILITY FINGERPRINT — model core-8
  • instruction adherence0.94
  • tool use0.88
  • reasoning · long ctx0.81
  • visual taste0.46
  • fast extraction0.91
  • code · agentic0.79
  • refusal calibration0.83
  • creative style0.58
Models become searchable by demonstrated behavior — not marketing descriptions.
BENCHMARK INTELLIGENCE
  • discriminates strong models
  • judge agreement 0.87
  • !3 redundant cases detected
  • !2 cases show grader instability
  • 7 cases suspect memorization
  • capability gap: multi-modal reasoning
DISAGREEMENT MINING
case #04217core-8 > max-large
case #08811judge-A judge-B
case #12904temp 0.2 temp 0.9
case #01366humans model judge
High-information cases roll forward automatically into the next benchmark generation.
DECISION ENGINE
tool-call success ≥ 95%  ·  p95 latency < 2 s  ·  cost per task ≤ $0.03
→  3 models qualify · continuously verified as the frontier moves
Design.·Measure.·Compare.·Decide.·Repeat.

Model benchmarking at the speed of innovation.

Models arrive weekly. Providers change. Prices fall. Context windows expand. Lattice continuously rebuilds the evidence you need to make good decisions — as infrastructure, not a dashboard someone remembers to check.

lattice.