This dashboard benchmarks each model and CAD-engine combo on a fixed set of prompts, scoring prompt adherence, aesthetics, cost, speed, and code errors. The short version: medium reasoning is the value sweet spot, Opus 4.7 is our reliability benchmark, and cheap models usually cost more once you count error-repair loops and latency.
9.50%
What is the average R-squared of all the runs? This tells us linear regression alignment human vs AI. 100% means AI perfectly predicts a human vote. 0% means AI doesn't predict it at all.
17.14%
When we set the model according to all pairs from all runs what is the R-squared?
109
Total number of evaluation runs.
$721.69
The sum of costs for all evaluation runs.
$6.62
The average cost of a single evaluation run.
7169m 49s
The sum of durations for all evaluation runs.
65m 47s
The average duration of a single evaluation run.
3264
The sum of all generations for all evaluation runs.