Benchmark models on the code you actually ship.
Synthetic benchmarks don't predict how a model handles your dense, real production code. Benchy evaluates models on real corpora — recall, compose, simplify, docgen, repair — and shows you exactly where each one breaks.
Leaderboard scores don't
survive contact with your repo.
Public benchmarks are built on synthetic, self-contained puzzles. Production code is dense, interdependent, and full of context — and model selection based on the wrong numbers is expensive guesswork.
LeetCode-style puzzles reward pattern memorization, not comprehension of real, messy modules.
A single accuracy score hides where a model fails — which functions, which task types, at what context depth.
Confidently invented APIs and phantom functions don't show up in pass/fail scores until they reach review.
"Is the 8B good enough for docgen?" deserves a measured answer before you build a workflow on it.
Five tasks. Real corpora.
Failure modes made visible.
Benchy builds evaluation sets from real production code and runs every model through the same gauntlet — then the analysis layer shows not just accuracy, but the shape of each model's failures.
The bench engine extracts functions and context windows from production repositories — the evaluation set is your code, not a puzzle book.
Recall-verbatim, compose, simplify, docgen, and repair — each probing a different way teams actually use models on code.
Hallucination rate, recall vs position, decay curves, and task-accuracy heatmaps — the metrics that explain the score.
API models and local models run through the same tasks with the same scoring — comparable by construction.
pick the top model →
it invents your APIs →
find out in review
5 tasks × N models →
leaderboard + heatmap →
per-function drill-down →
choose with evidence
Corpus to conclusion,
in six stages.
Every run is scored against ground truth from the corpus itself. Results land in dashboards built for deciding, not admiring.
Pull functions and context from real production repos.
Generate recall, compose, simplify, docgen, and repair items.
Wire API and local models into one harness.
Every model, every task, same corpus, same scoring.
Leaderboards, heatmaps, hallucination and decay curves.
Per-function failures show what the aggregate hides.
An analysis layer,
not just a runner.
Running prompts is the easy half. Benchy's value is in what it measures and how it shows it.
Bench sets built from production code with schema-normalized dumps.
Recall-verbatim, compose, simplify, docgen, repair — one corpus, five lenses.
Model ranking per task and overall, with result bars and comparisons.
Model × task grids that expose specialization and blind spots at a glance.
Invented APIs and phantom code measured per model, per task.
Accuracy vs position and context depth — where memory actually fades.
Every failure traceable to the exact function and answer that produced it.
API and local models configured side by side, scored identically.
Real screens.
Real runs.
The dashboard and analysis views from actual benchmark runs — rankings, per-function results, and recall behavior.





One run.
A decision you can defend.
A benchmark run resolves to ranked models with the failure analysis attached — so "we chose this model" comes with receipts.
Benchy is model evaluation grounded in your code — not a public leaderboard, not a vibes check, not a demo prompt.
It is the harness for teams who pick models for engineering work and want the failure modes on the table first.