R&D · Methodology · 2026-06-15
The image model defaults to the average of every tasteful brand board it has seen — airy voids, mediumless vector, an untreated wordmark. The cure isn’t more iteration. It’s naming a medium, diverging before committing, and treating the toolkit as a direction-communicating Style Tile — with a safe↔bold dial. We tested it. The same model, on the same first attempt, went from generic to expressive.
The problem
Ask for “a brand toolkit for a pizzeria” and you get the statistical mean: a cork-board of floating white cards, clean mediumless vector, a plain serif wordmark, tasteful pill swatches with caption names, and thin line-art icons. Competent — and completely generic. (This one even garbled FLAME → FLAM.)
This is not the model’s ceiling — it’s its default. Everything below is about overriding that default deliberately.
Diagnosis
Grounded in visual forensics across 16 prior runs, the impeccable skill internals, and external research on generative homogenization; then corrected by an adversarial pass.
An unconstrained ask lands in a dense “safe attractor.” The model supplies a specific medium readily when named — but never reaches for one unprompted.
The playbook generates one toolkit board, gated approve/redo — a 1-wide search. It even disabled the visual-direction probe. The first thing the user sees is one safe guess.
The gate checks only colors-present / legible / no-chrome. A perfectly generic board passes every item. The toolkit is framed as a palette lock, not a direction it communicates.
The anti-generic intelligence already exists in the impeccable skill — the AI-slop test, reflex-reject lists, “Safe = invisible,” the Restrained↔Drenched axis. The graphic-craft playbook just loads none of it at the toolkit step.
The fix
Six process changes (PC1–PC6), drafted as a proposal — the live playbook is intentionally untouched. The load-bearing move: reframe the toolkit from a swatch sheet into an in-medium Style Tile that predicts the mock.
| # | Change | Tier |
|---|---|---|
| PC1 | Divergence gate — 3 named “territory” boards first | adopt |
| PC2 | Toolkit → in-medium Style Tile | adopt |
| PC3 | Positive expressiveness rubric (auto + human) | adopt |
| PC4 | Medium+composition contract (out-of-band) | adopt |
| PC5 | Branch, don’t “make it bolder” | adopt |
| PC6 | SAFE↔BOLD dial (1–5), logged per run | adopt |
Flash scoring, exact divergence count, and verbatim-contract-in-prompt are gated behind validation — they touch unproven instruments or contradict an existing finding.
The experiment
Control = the current generic prompt. Treatment = named-medium-first + Style-Tile format + regression-to-mean patterns, at dial L4, across three distinct media. One attempt each, no human iteration. Wordmark fonts kept SAFE-tier (Work Sans / Playfair / Montserrat) so expressiveness never broke buildability.


Three media, all first-attempt



Results
Scored by a new calibrated Gemini-Flash scorer (the playbook’s flash_eyes.py can’t judge). It first passed calibration against human-labeled pairs (expressiveness 3/4, forced-choice 4/4), then scored the new images blind.
| Cell | Expr | POV | /9 | Medium named |
|---|---|---|---|---|
| A1 control | 3 | 3 | 5 | none / clean vector |
| A2 control | 3 | 3 | 5 | none / clean vector |
| A3 control | 2 | 2 | 3 | none / clean vector |
| B1 riso | 4 | 4 | 7 | risograph |
| B2 sign | 4 | 4 | 8 | enamel sign |
| B3 screen | 4 | 4 | 8 | silkscreen |
| Comparison | Stronger POV |
|---|---|
| A1 vs B1 (riso) | B · 3–0 |
| A2 vs B2 (sign) | B · 3–0 |
| A3 vs B3 (screen) | B · 3–0 |
| L2 vs L3 (dial) | bolder · 3–0 |
| L4 vs L5 (dial) | bolder · 3–0 |
3 of 4 success criteria cleanly met; the 4th (mean-expressiveness +1.5) missed only at +1.33 — the judge’s scalar is compressed. The discriminating signals (unanimous forced-choice, perfect medium-naming, +3.4 rubric items) are decisive, and the human eye agrees.
The dial
The same riso direction across four dial levels. POV climbs 2.67 → 3.33 → 4 → 4; rubric items 4 → 6 → 7 → 8. L5 breaks category codes — memorable, but it stops reading as “a pizzeria,” which is exactly why it’s opt-in only.




Key nuance: the dial governs brand boldness; the toolkit’s own craft stays high even at L1–L2 — so a deliberately safe brand still ships a custom-looking, direction-communicating tile, never a generic one.
Buildability & what’s next
The adversarial pass’s most valuable catch: the strongest expressive levers collide with the project’s hard-won build machinery. The guardrails keep an expressive direction buildable.
Treatment (knockout/offset/paint) is metric-neutral — keep it. The font must be SAFE-tier or a fixed SVG logo; AVOID-tier display bakes narrower → HTML overflows.
Flood feature bands only; exempt text-dense sections. WCAG AA contrast is a hard gate that outranks the density bar.
Keep generation prompts minimal; carry the art-direction contract as a verification artifact, never pasted into the prompt (verbose prompts degrade font fidelity).