§.FAQ

Which LLM model should I pick?

Decision matrix: speed vs quality vs cost for each supported model.

Updated 2026-04-13 · By Jon Lasley
Use caseTop pickRunner-up
Complex reasoning, long context, highest qualityClaude Opus 5GPT-5.6 Sol or Claude Opus 4.8
Most general-purpose prompt workClaude Sonnet 5GPT-5.6 Terra or Claude Sonnet 4.6
Fast, cheap, high volumeClaude Haiku 4.5GPT-5.6 Luna or Gemini 3.1 Flash-Lite
Multi-modal (images, audio, video, PDF)Gemini 3.7 FlashGemini 2.5 Pro
Cost-optimized inference at scaleGemini 2.5 Flash-LiteGPT-5.6 Luna
Hardest reasoning problemsGPT-5.6 Sol at max effortClaude Opus 5 or Claude Opus 4.8
Reasoning at low costGPT-5.6 LunaGPT-5.4 Mini or GPT-5.4 Nano
Coding-task promptsGPT-5.3 Codex (target-only)GPT-5.6 Sol or Claude Opus 5
Latency-criticalClaude Haiku 4.5GPT-4.1 Mini (no reasoning tokens)
Picks marked target-only can be set as a prompt's Target Model but are not offered as a Workbench Model. Six entries are target-only today: Claude Fable 5, GPT-5.5 Pro, GPT-5.4 Pro, GPT-5.3 Codex, Llama 3.3 70B, and Mistral Large. Claude Fable 5 is the pick when you are authoring for multi-hour autonomous runs. One scheduling note: GPT-4.1 Nano retires on 2026-10-23, and OpenAI's stated replacement is GPT-5.6 Luna.
Experiment in the playground
Two paths in the playground: switch the model dropdown and re-run sequentially (each run recorded in test-run history with latency / tokens / cost), or click Compare models to run the same prompt against 2–5 models in parallel and see outputs, costs, and latencies side-by-side. With a judge step you can also score every model's output against a shared rubric and surface the best per criterion. See Comparing models in the playground.