What we shipped, week by week.
The public log of features, improvements, fixes, and migrations on the Prompt Assay workbench. New entries land most weeks. Subscribe via RSS or follow @PromptAssay on X.
Entries
GPT-6 Astra and Claude Fable 5.1 join the model pickers
Two new models, a reasoning-effort fix for Claude Sonnet 5 and Opus 5, and clearer pricing and context data across the model pickers.
Claude Opus 5, Sonnet 5, the GPT-5.6 family, and the Gemini 3.x line
Twelve new models across all three providers, a new Max reasoning-effort level, and a cheaper Sonnet. Plus corrected pricing and capability data for models you were already using.
Convert Prompt to Skill, Skills REST API, public prompt links, kind switcher
Convert turns any prompt into an Agent Skill bundle. The REST API gains skill-file and Skill Report endpoints, prompts get public share links, and the sidebar adds a Prompts and Skills switcher.
Skills workbench launches · Behavioral Eval · evaluations open on every tier
Author Agent Skills in Prompt Assay, score them with a six-dimension Critique, and run a Behavioral Eval across Claude, GPT, and Gemini. Evaluation suites now open on every tier.
Demo mode, cost drill-down, the GPT-5 lineage, and a much sharper AI pair
A big week of ships: free demo runs for new accounts, per-run cost receipts, GPT-5 and Gemini 2.5 Flash-Lite, shareable Critique and Compare, plus cross-model testing in the Playground.
Vote on what we ship next
A two-board feedback system where you can file bugs, request features, and vote on what we ship next. Plus a brainstorm timeout fix you may have hit on long Opus runs.
Schedule eval batches without watching the timer
The new Schedule batch action queues eval-suite runs through Anthropic's Batches API so you can walk away. Results post back when the batch returns.