Overview

Benchmarks don't tell you which model designs a better interface. Taste does. So instead of scores, I ran a practical test: the same prompt, the same constraints, judged on the things that make UI feel finished. The video shows the full head-to-head; this is the written verdict.

Both are excellent. The differences are in the details, which is exactly where design lives.

First passFable 5Looks shipped out of the gate: considered spacing, realistic data, intentional interactions.
Attention to detailFable 5Fewer AI tells (emoji in cards, logos as text, skipped graphics) across a full layout.
IterationFable 5Targeted feedback changes the thing you asked about without regressing the rest.
Bold starting pointsGPT-5.6More ambitious first drafts when you want an experimental direction.
List price (July 2026)GPT-5.6 Sol $5 / $30 vs Fable 5 $10 / $50Per million input / output tokens. Fable 5 costs more per token; Sol is the cheaper starting point.

Which model produces the better first pass?

First impressions matter because most of a project's direction is set in the opening minutes. Fable 5 tends to produce a first pass that already looks shipped: considered spacing, realistic data instead of lorem ipsum, interactions that feel intentional. GPT-5.6 comes out strong too, often with ambitious touches, but the polish can be less even from section to section.

If you judge only the opening screen, Fable 5 usually looks the more finished of the two.

GPT-5.6 vs Fable 5: The Ultimate UI Design Test - Which model produces the better first pass?

Which model holds detail across a full layout?

This is where a design test is won or lost. The tells of AI-generated UI, emoji stuffed into cards, generic gradients, real company logos rendered as plain text, complex graphics quietly skipped, are all about detail. Fable 5 is the stronger of the two at avoiding those tells and holding quality across an entire layout. GPT-5.6 can produce impressive individual moments but is more likely to drop detail as the screen gets complex.

For work where the finish is the point, that consistency is the deciding factor.

Which model takes design feedback better?

No first pass is the final answer, so how a model takes feedback matters as much as its opening move. Both handle plain-English direction well. The question is whether targeted feedback improves the specific thing you asked about without regressing everything else. In my testing Fable 5 iterated more predictably, while GPT-5.6 occasionally reworked more than requested.

The prompting approach matters a lot here regardless of model. I broke down how to get consistently great output from Fable 5 in a dedicated guide: Claude Fable 5 for UI Design: Beautiful Output Every Time

GPT-5.6 vs Fable 5: The Ultimate UI Design Test - Which model takes design feedback better?

What the two models are, and what they cost

GPT-5.6 went into general availability on July 9, 2026 across ChatGPT, the API and Codex, as a family of three: Sol (the flagship I tested), Terra for everyday work and Luna as the cheapest tier. Sol lists at $5 per million input tokens and $30 per million output. Claude Fable 5 is Anthropic's first publicly available Mythos-class model, priced at $10 per million input and $50 per million output.

On paper that makes GPT-5.6 the cheaper model per token. In practice the bill depends on how many tokens each model burns to finish, which is why I look at usage on real work rather than list price. Fable 5 is token-hungry; a single dashboard prompt moved my Claude Max usage 23 points, and I documented that in the Fable 5 guide: Claude Fable 5 for UI Design: Beautiful Output Every Time

Prices as of July 2026 from OpenAI's and Anthropic's pricing pages; both change often.

Which model I'd trust with UI work

For pure UI design work, Fable 5 is the one I reach for: more consistent polish, better detail retention, more predictable iteration. GPT-5.6 is genuinely capable and worth using, especially when you want a bolder or more experimental starting point. The good news is you don't have to marry one, and both fit into a wider stack. See where each model sits in the full guide: AI Design Tools for Designers (2026): The Complete Guide

September 2026 update: GPT-6 Astra and Fable 5.1

Both models have a successor now, and the two launched in the same week. Claude Fable 5.1 keeps Fable 5 pricing, cuts cached-input cost by 75%, and in my three-brief test made stronger first-pass calls on hierarchy, typography and density than 5 did: I Tested Claude Fable 5.1 as a UI Designer

GPT-6 Astra is the bigger shift for web design specifically, because inside Codex it can call Images 2.0 and generate the page's photography as part of the build. On an image-heavy landing page it is now my first pick; for product UI I still reach for Claude Code: I Tested GPT-6 Astra for Web Design (with Images 2.0)

The verdict below stands as a record of the July test, and the judging criteria have not changed: first pass, detail, iteration.

FAQ

For consistent, finished-looking UI, Fable 5 is my pick. GPT-5.6 is very capable and can be better for bolder, more experimental first passes.

No. As of September 2026 both have successors: GPT-6 Astra from OpenAI and Claude Fable 5.1 from Anthropic. The design criteria in this post still apply; my tests of both newer models are linked in the update section above.

Often, yes. Grounding the model in real references and realistic content improves both models more than switching between them.

Absolutely. Many designers start a concept in one and refine in the other. They fit alongside each other in a wider AI design stack.

At the July 2026 test, GPT-5.6 Sol listed at $5 per million input tokens and $30 per million output; Fable 5 at $10 and $50. Real cost depends on how many tokens each burns to finish a task.

Continue reading

Claude Design Tutorial for Beginners (Full Walkthrough)AI Design Tools for Designers (2026): The Complete Guide