Which LLM builds the best CAD? Twelve models run through one parametric 3D-parts benchmark, scored on fidelity, watertight solids, geometry checks and cost.
On OwlCAD’s published parametric-parts benchmark, GPT-5.6 Sol Pro ranks first, averaging 4.3 out of 5 for fidelity with 99% of its generations producing a watertight solid. Grok 4.5 and Gemini 3.6 Flash come next, and Gemini 3.6 Flash gets there at roughly a quarter of the cost per generation. The ranking is a weighted composite of fidelity, watertightness, geometry checks, cost and speed, over 12 benchmark runs.
Start with the absolute numbers. The composite score only ranks the models against each other.
How faithfully the part matches the request, graded by a blind judge: 5 means the right features, placement and proportions, ready to print. The number that matters most.
The share of generations that built a non-empty, watertight solid in the CAD kernel. A model that writes valid output but no solid gets no credit here.
The share of prompts whose automatic expectations all passed, for example that holes are really cut rather than faked, or that a standard generator was used where one exists.
Average provider cost in US dollars and average wall-clock time, both measured through OpenRouter: comparable across models, not identical to what a direct route costs.
A weighted blend of fidelity 45%, watertight 30%, checks 10%, cost 10% and speed 5%, each normalised across the models in the table. It ranks this field; it is not an absolute grade.
Published August 31, 2026 · 12 benchmark runs · 752 generations
AI models ranked on OwlCAD’s parametric CAD benchmark
Open real parts from the models Pro runs, in the editor, with no account.
Every model gets the same set of part prompts — brackets, gears, enclosures, print-in-place mechanisms, lofts and revolved shapes — through the same system prompt OwlCAD uses in production. Each answer is validated, built into a solid by the Manifold CAD kernel, checked against the prompt’s geometric expectations, and graded from 1 to 5 for fidelity by a judge model that is not told which model produced it.
It is disclosed rather than hidden. The judge, Claude Opus 4.8, never sees which model produced a tree, and the grades it gave its own generations are counted in the figures. The table says how many of its grades were self-graded, so you can discount them if you prefer.
Pro runs Gemini 3.6 Flash, Gemini 3.5 Flash, Claude Opus 4.8 and GPT-5.6 Sol on your monthly credits. Every other model in the table, including GPT-5.6 Sol Pro and Grok 4.5, runs with your own provider key and spends no credits.
So every model is measured through one transport and one price list. OwlCAD reaches some models directly in production, so real cost and time can differ from these figures; fidelity and watertightness do not depend on the route.