KUMYU / CATALOG

Research archive · AI slop controls

AI Slop Lab

Compare exact prompts, raw outputs and recorded settings for AI-generated websites, motion and images. The source map turns reported fixes into testable hypotheses. These first outputs are exploratory pilots; unavailable model controls and changed factors remain explicit.

This is a research record. It does not rank models or treat a planned run as executed.

11 references · 9 cases · 19 completed outputs

Evidence first

References used in these cases

The compact list contains the original material used by the current cases. Open the annotations only when you need source-level context.

Frontend prompt instructions

OpenAI · OpenAI's frontend prompt gives concrete guardrails for context-aware, coherent interfaces.

Provider guidance · Full text

Designing delightful frontends with GPT-5.4

OpenAI · The guide links underspecified briefs to generic layouts and recommends real content, references, constraints, and restrained composition.

Provider guidance · Full text

What's New in WCAG 2.2

W3C WAI · WCAG 2.2 adds requirements around visible, unobscured focus and target sizing, grounding visual polish in usable interaction.

Provider guidance · Full text

Image generation

OpenAI · OpenAI’s image API guide documents generation and image-editing controls as product capabilities rather than a universal aesthetic recipe.

Provider guidance · Full text

Prompt Basics

Midjourney · Midjourney recommends concise, visually specific prompting and names subject, medium, environment, lighting, color, mood, and framing as useful dimensions.

Provider guidance · Full text

FLUX image best-practice core principles

Black Forest Labs · BFL’s maintained guidance proposes descriptive natural language, front-loaded priorities, and one-variable iterative refinement.

Provider guidance · Full text

Human Learning by Model Feedback

arXiv · The study examines how people revise prompts after seeing generations, including adaptation to a model’s preferences.

Study · Full text

Art-style prompt troubleshooting

r/midjourney · A practitioner thread proposes checking weighting, raw mode, style references, and regional edits when an intended style is missed.

Anecdote · Full text
All source annotations (48)

Provider guidance · OpenAI · Access: Full text

Image generation

OpenAI’s image API guide documents generation and image-editing controls as product capabilities rather than a universal aesthetic recipe.

Published: unknown · Accessed: 2026-10-04

Source context

Intervention Record model, size, quality, input images, exact prompt, and every edit pass beside the output.

Caveat Provider documentation explains interfaces, not independent proof of visual quality or authorship.

Provider guidance · OpenAI · Access: Full text

Images and vision

OpenAI documents image inputs alongside generation, making visual references a first-class input to an iterative workflow.

Published: unknown · Accessed: 2026-10-04

Source context

Intervention Give each reference a named job: subject, layout, palette, material, or unacceptable pattern.

Caveat A reference can change multiple attributes at once; a result does not identify which attribute caused the change.

Provider guidance · Midjourney · Access: Full text

Prompt Basics

Midjourney recommends concise, visually specific prompting and names subject, medium, environment, lighting, color, mood, and framing as useful dimensions.

Published: unknown · Accessed: 2026-10-04

Source context

Intervention Write the brief as a shot card with only the two most important constraints marked as non-negotiable.

Caveat Conciseness is model-specific guidance, not proof that short prompts outperform detailed prompts elsewhere.

Provider guidance · Midjourney · Access: Full text

Style Reference

Style Reference separates visual feel from depicted content and exposes a style-weight control for testing influence.

Published: unknown · Accessed: 2026-10-04

Source context

Intervention Use style reference for palette, medium, texture, or lighting; place the desired subject and action in text.

Caveat Version changes can alter reference interpretation, so a reference code alone is not reproducible metadata.

Provider guidance · Black Forest Labs · Access: Full text

FLUX official inference repository

The official repository distinguishes text-to-image, fill, structural conditioning, variation, and image-editing model variants with different licenses.

Published: unknown · Accessed: 2026-10-04

Source context

Intervention Choose an operation before prompting: generate, preserve structure, vary, or edit; record the exact model and license.

Caveat Open weights and permission for a particular commercial use are separate questions.

Provider guidance · Black Forest Labs · Access: Full text

FLUX image best-practice core principles

BFL’s maintained guidance proposes descriptive natural language, front-loaded priorities, and one-variable iterative refinement.

Published: unknown · Accessed: 2026-10-04

Source context

Intervention Put subject and action first, then context, lighting, material, and framing; alter one field per revision.

Caveat Claims about ideal length or lighting impact are provider recommendations, not general experimental findings.

Provider guidance · Black Forest Labs · Access: Full text

Prompting Guide - FLUX.2

BFL publishes model-family-specific prompting guidance, reinforcing that controls must be read against the actual selected model.

Published: unknown · Accessed: 2026-10-04

Source context

Intervention Attach a model-version field to every gallery card and prevent cross-model comparisons from masquerading as one experiment.

Caveat Documentation may change after capture; retain the URL and accessed date with the result.

Study · arXiv · Access: Full text

Design Guidelines for Prompt Engineering Text-to-Image Generative Models

This paper studies prompt structures and failure modes for text-to-image systems, treating prompting as an interaction design problem.

Published: 2021-09-28 · Accessed: 2026-10-04

Source context

Intervention Expose prompt parts in the page UI so a reader can identify which creative decision each phrase encodes.

Caveat The work predates current image models; its design framing is useful, but model behavior must be re-tested.

Study · arXiv · Access: Full text

Human Learning by Model Feedback

The study examines how people revise prompts after seeing generations, including adaptation to a model’s preferences.

Published: 2023-11-20 · Accessed: 2026-10-04

Source context

Intervention Show revision history and the author’s intent so viewers can distinguish learning from simply optimizing for a model’s default look.

Caveat It describes observed behavior; it does not declare one prompt style artistically better.

Provider guidance · W3C WAI · Access: Full text

Understanding SC 2.3.3: Animation from Interactions

W3C explains that non-essential interaction-triggered motion must be disableable at WCAG AAA, because it can distract or cause physical symptoms.

Published: 2025-09-16 · Accessed: 2026-10-04

Source context

Intervention Give every animation a declared information or feedback purpose and a reduced-motion alternative.

Caveat This criterion is Level AAA guidance; product conformance still needs an accessibility review in context.

Provider guidance · MDN · Access: Full text

prefers-reduced-motion CSS media feature

MDN documents the widely supported CSS preference for removing, reducing, or replacing non-essential motion.

Published: unknown · Accessed: 2026-10-04

Source context

Intervention Treat reduced motion as a designed mode: retain state changes while replacing travel and scale with opacity or instant transitions.

Caveat The media query signals a preference; it does not decide which motion is meaningful in your interface.

Design practice · web.dev · Access: Full text

prefers-reduced-motion: Sometimes less movement is more

web.dev gives implementation-oriented guidance for honoring motion preference as part of ordinary frontend work.

Published: unknown · Accessed: 2026-10-04

Source context

Intervention Make reduced-motion screenshots and interaction states part of every animation experiment’s published evidence.

Caveat Preference support does not validate an animation’s visual hierarchy, timing, or content relevance.

Provider guidance · Motion · Access: Full text

Create accessible animations in React

Motion documents React-oriented mechanisms for respecting reduced-motion preferences while keeping interaction feedback available.

Published: unknown · Accessed: 2026-10-04

Source context

Intervention Store a motion intent beside each component: orient, confirm, reveal, or decorate; only the first three have a default justification.

Caveat A library feature cannot substitute for deciding whether a transition helps the reader.

Provider guidance · C2PA · Access: Full text

C2PA Technical Specification

C2PA specifies signed provenance assertions for digital media so a viewer can inspect a record of origin and edits when it is present.

Published: unknown · Accessed: 2026-10-04

Source context

Intervention Publish a human-readable generation ledger even when a file carries credentials: prompt, model, settings, references, edits, and export date.

Caveat A provenance record describes declared history and integrity; it does not by itself establish factual truth or aesthetic value.

Provider guidance · Content Authenticity Initiative · Access: Full text

Content Credentials

Content Credentials presents provenance information as inspectable context for media rather than a visual style marker.

Published: unknown · Accessed: 2026-10-04

Source context

Intervention Place a compact disclosure beside every public output and link to the complete experiment record.

Caveat Credentials can be missing or stripped in distribution; the page must remain understandable without them.

Anecdote · r/midjourney · Access: Full text

Reference syntax discussion

A community answer separates character, general image, and style-reference syntax, reflecting a practitioner distinction worth testing.

Published: unknown · Accessed: 2026-10-04

Source context

Intervention Label the intent of each reference before generation instead of treating all uploaded images as interchangeable.

Caveat This is a single community explanation and product syntax can change.

Anecdote · r/midjourney · Access: Full text

Art-style prompt troubleshooting

A practitioner thread proposes checking weighting, raw mode, style references, and regional edits when an intended style is missed.

Published: unknown · Accessed: 2026-10-04

Source context

Intervention Treat each proposed fix as a separate hypothesis; never stack weights, style mode, and reference changes in one diagnostic run.

Caveat Advice is anecdotal and tied to a specific product version and user’s goal.

Anecdote · r/midjourney · Access: Full text

Reference-image limitation discussion

A user reports a detailed object not being retained from a reference; replies frame reference influence as non-deterministic.

Published: unknown · Accessed: 2026-10-04

Source context

Intervention For critical geometry or product details, use a layout or edit workflow and inspect the exact feature at 100% before publishing.

Caveat One failure report does not quantify reliability across prompts, models, or reference types.

Anecdote · r/midjourney · Access: Partial

Different from the reference

A community question about divergence from a reference is retained as a signal for a controlled reference-fidelity test.

Published: unknown · Accessed: 2026-10-04

Source context

Intervention Pre-register which reference attributes must survive before selecting an output.

Caveat Page retrieval was partial, so do not infer a general technique from its discussion.

Anecdote · r/midjourney · Access: Partial

Style-control discussion

A community reply suggests lowering stylization or using raw style to reduce default styling, a claim suitable for an A/B test.

Published: unknown · Accessed: 2026-10-04

Source context

Intervention Define the unwanted default tendency before changing a style control, then evaluate only that tendency first.

Caveat The claim is anecdotal and parameter semantics can shift between product releases.

Anecdote · r/midjourney · Access: Partial

Content versus style reference discussion

A community reply distinguishes style-reference use from using a regular image prompt for content or composition.

Published: unknown · Accessed: 2026-10-04

Source context

Intervention Use a reference board with separate columns for content, composition, and visual treatment.

Caveat This is practitioner interpretation, not a guarantee of model behavior.

Anecdote · r/midjourney · Access: Partial

Reference-strength equivalence discussion

A question about reference-strength equivalents highlights that controls cannot be assumed portable between image systems.

Published: unknown · Accessed: 2026-10-04

Source context

Intervention Publish a control-mapping table with an explicit “no equivalent” state instead of inventing cross-tool parity.

Caveat Partial retrieval; the entry is a research lead, not evidence that any controls are equivalent.

Anecdote · r/midjourney · Access: Partial

How to generate images like this?

A request for reproducing an image through references is a useful reminder to separate inspiration, controllable attributes, and prohibited copying.

Published: unknown · Accessed: 2026-10-04

Source context

Intervention Describe observable properties—layout, light direction, material, palette—rather than asking to copy a source image or artist.

Caveat Partial community page; it does not establish legal, ethical, or technical rules.

Provider guidance · OpenAI · Access: Full text

Frontend prompt instructions

OpenAI's frontend prompt gives concrete guardrails for context-aware, coherent interfaces.

Published: unknown · Accessed: 2026-10-04

Source context

Intervention Supply audience and existing-system context; ask for a CSS color scan and coherent non-overlapping layout.

Caveat Provider guidance is a recipe, not evidence that a listed style is inherently bad.

Provider guidance · OpenAI · Access: Full text

Designing delightful frontends with GPT-5.4

The guide links underspecified briefs to generic layouts and recommends real content, references, constraints, and restrained composition.

Published: 2026-03-20 · Accessed: 2026-10-04

Source context

Intervention Start with a mood board, define tokens and narrative, then use a single focal composition and test floating elements at breakpoints.

Caveat It targets a specific model family and presents practical advice rather than a controlled design study.

Provider guidance · OpenAI · Access: Full text

Prompt engineering

OpenAI recommends explicit responsibilities, concrete tool examples, a quality rubric, and validation for agentic work.

Published: unknown · Accessed: 2026-10-04

Source context

Intervention Ask the model to make an internal rubric, implement, inspect the rendered result, and report failed criteria.

Caveat A rubric can turn into another shared default unless its criteria arise from the product and audience.

Provider guidance · OpenAI · Access: Full text

Image prompting

The image guide treats UI previews as product artifacts: describe hierarchy, real elements, real copy, and concrete constraints.

Published: unknown · Accessed: 2026-10-04

Source context

Intervention Use an artifact specification with canvas, hierarchy, real text, and a prohibition on decorative clutter instead of a style-only request.

Caveat Image previews can communicate direction but are not responsive, accessible, or production implementation proof.

Provider guidance · Anthropic · Access: Full text

Prompting best practices

Anthropic advises direct, explicit instructions and warns that unguided frontend generation can converge on generic patterns.

Published: unknown · Accessed: 2026-10-04

Source context

Intervention State deliverable, audience, visual direction, exclusions, and review procedure as separate prompt sections.

Caveat Prompt specificity can improve adherence while still preserving a model's learned visual bias.

Provider guidance · Anthropic · Access: Full text

Harness design for long-running application development

Anthropic describes predictable but visually unremarkable layouts as a self-evaluation problem and makes subjective quality gradable.

Published: unknown · Accessed: 2026-10-04

Source context

Intervention Turn visual taste into named, observable criteria and feed the agent screenshots rather than relying on code inspection alone.

Caveat A gradable harness only tests the criteria it encodes; it cannot certify originality or audience fit.

Provider guidance · Anthropic / Claude Code · Access: Full text

Frontend design skill

The official skill asks an agent to commit to an intentional visual direction instead of repeating generic defaults.

Published: unknown · Accessed: 2026-10-04

Source context

Intervention Require a written design-direction decision before implementation, including purpose, audience, constraints, and a memorable element.

Caveat The skill is guidance, not an independently validated anti-slop detector.

Design practice · AkyRayy · Access: Full text

Frontend Design Skill

This community skill packages anti-patterns with typography, color, layout, motion, accessibility, and performance guidance.

Published: unknown · Accessed: 2026-10-04

Source context

Intervention Audit one surface at a time—type, palette, layout, motion, or content—rather than applying a single style blacklist.

Caveat Its pattern list is authored guidance, so test it against the product rather than treating it as universal law.

Design practice · nattergabriel · Access: Full text

unslop

Unslop separates prevention from cleanup and routes rules by surface, including prose, UI copy, naming, code, and frontend work.

Published: unknown · Accessed: 2026-10-04

Source context

Intervention Record whether the run is preventive generation or audit-and-rewrite, then evaluate only the relevant surface rules.

Caveat The repository describes its own method; its claims of coverage are not external outcome evidence.

Design practice · Studio Groei · Access: Full text

anti-slop

The project argues that a generic anti-slop skill can itself become a convergence source and proposes curation, tooling, and broad iteration.

Published: unknown · Accessed: 2026-10-04

Source context

Intervention Build a project-specific reference library and choose an aesthetic taxonomy before coding; expose parameters for iteration.

Caveat A curated library can still reproduce its curator's narrow taste or source rights constraints.

Design practice · herzigma · Access: Full text

Frontend Design Architect — System Prompt

This prompt kit combines a style selector, exclusions, and execution guidance for AI-assisted frontend work.

Published: unknown · Accessed: 2026-10-04

Source context

Intervention Select one visual direction and one major visual moment, then test whether implementation decisions support that direction.

Caveat Universal prompts risk convergence when reused unchanged across unrelated products.

Anecdote · r/SideProject · Access: Full text

Fighting the AI slop!

A maker shared one-shot prompt-to-design examples and explicitly left visible flaws as a demonstration of expected limitations.

Published: unknown · Accessed: 2026-10-04

Source context

Intervention Label one-shot output as a baseline rather than a finished result; capture a second pass separately.

Caveat One maker's self-report and examples cannot establish typical model behavior.

Anecdote · r/claudeskills · Access: Full text

I built a skill that makes landing pages look LESS AI SLOP

A practitioner proposes screenshot input and an explicit create-review loop after finding one-pass results insufficient.

Published: unknown · Accessed: 2026-10-04

Source context

Intervention Make the agent inspect a screenshot after each major change and require a stated next defect to fix.

Caveat The report does not provide a controlled comparison, so treat it as a workflow hypothesis.

Anecdote · r/vibecoding · Access: Full text

Any advice to escape the AI slop web design look?

The thread names repeated dark gradients, glass cards, glowing blobs, generic fonts, and excessive rounding as perceived tells.

Published: unknown · Accessed: 2026-10-04

Source context

Intervention Write a design intent first, then test targeted changes to color, radius, type, and density rather than merely inverting a checklist.

Caveat These are community aesthetic judgments and can overlap with valid, accessible design choices.

Anecdote · r/vibecoding · Access: Full text

All AI websites and designs look the same

Participants argue both that outputs converge and that a widely shared anti-slop prompt could become the next uniform aesthetic.

Published: unknown · Accessed: 2026-10-04

Source context

Intervention Measure diversity across a batch and rotate references or constraints rather than installing a permanent universal style rule.

Caveat The argument is opinion, and visual difference alone does not show better usability or product fit.

Anecdote · r/VibeCodeDevs · Access: Full text

How Do You Make Websites Look Less Like AI Slop?

Responses recommend supplying visual references, product flow, repository context, and sometimes a human-designed Figma starting point.

Published: unknown · Accessed: 2026-10-04

Source context

Intervention Separate art direction from implementation: provide a wireframe or reference first, then request code for that direction.

Caveat Copying a reference can create legal, accessibility, and differentiation problems; use it as direction, not a clone target.

Anecdote · r/ClaudeCode · Access: Full text

finally figured out why claude's UI generations look like AI slop and how to fix it

Thread participants favor screenshot references and explicit visual goals over a generic request to build a website.

Published: unknown · Accessed: 2026-10-04

Source context

Intervention Ask the model to list what it extracts from each reference before implementation, then verify the result is not a literal clone.

Caveat Community comments are not evidence that a reference workflow generalizes across products or models.

Anecdote · r/webdesign · Access: Full text

Trying to kill the AI Slop look

A builder describes generating brand maps from business context, imagery, logos, and explicit visual rules before site generation.

Published: unknown · Accessed: 2026-10-04

Source context

Intervention Create a compact brand map with product facts, palette sources, type mood, interaction character, and forbidden imitations.

Caveat The reported engine outcome is an unverified individual claim and scraping can introduce incorrect or rights-limited source material.

Study · arXiv · Access: Full text

Interrogating Design Homogenization in Web Vibe Coding

This preprint analyzes how frictionless web vibe-coding can narrow design diversity and proposes productive friction as mitigation.

Published: 2026-03-13 · Accessed: 2026-10-04

Source context

Intervention Insert explicit decision points: choose references, reject a default, justify hierarchy, and review renderings before accepting output.

Caveat It is a preprint and a sociotechnical risk analysis, not a completed large-scale causal experiment.

Study · arXiv · Access: Full text

AesthetiQ: Enhancing Graphic Layout Design via Aesthetic-Aware Preference Alignment

AesthetiQ argues layout systems need contextual aesthetic preferences and proposes preference alignment plus a layout-quality evaluation method.

Published: 2025-03-01 · Accessed: 2026-10-04

Source context

Intervention Evaluate rendered layouts with declared criteria—hierarchy, balance, legibility, context fit—instead of only checking generated markup.

Caveat The paper's proposed MLLM evaluation is not a substitute for representative human usability testing.

Study · Social Science Computer Review · Access: Partial

AI and Aesthetic Alienation: The Image and Creativity in Contemporary Culture

This comment piece frames mass-produced generative imagery called AI slop as a form of aesthetic alienation.

Published: 2025-07-18 · Accessed: 2026-10-04

Source context

Intervention Include authorship, intent, source provenance, and audience context in an evaluation, not only surface-level visual tells.

Caveat It is a conceptual comment piece; it does not provide a detector or a general causal test of audience response.

Provider guidance · W3C WAI · Access: Full text

What's New in WCAG 2.2

WCAG 2.2 adds requirements around visible, unobscured focus and target sizing, grounding visual polish in usable interaction.

Published: 2023-10-05 · Accessed: 2026-10-04

Source context

Intervention Treat keyboard focus visibility and contrast as release checks for every generated visual variant.

Caveat Accessibility conformance does not decide whether a design is distinctive or aesthetically successful.

Provider guidance · web.dev · Access: Full text

Animation and motion

web.dev advises selective animation, user control, and reduced-motion support rather than decorative motion by default.

Published: unknown · Accessed: 2026-10-04

Source context

Intervention Give every animation a feedback or orientation purpose, implement prefers-reduced-motion, and avoid unattended infinite loops.

Caveat Reduced motion is an accessibility requirement; it does not by itself diagnose generic visual style.

Design practice · Nielsen Norman Group · Access: Full text

Visual Hierarchy in UX: Definition

NN/g describes visual hierarchy and a blur or squint test for detecting unintended emphasis and clutter.

Published: unknown · Accessed: 2026-10-04

Source context

Intervention Add a blurred-screenshot hierarchy review to the run log before evaluating cosmetic details.

Caveat Hierarchy tests reveal attention patterns, not whether a page is AI-made or suitable for every audience.

Anecdote · X / Thais Branco · Access: Unavailable

Thais Branco post on AI aesthetics and evaluation

The post was indexed as a discussion of AI aesthetics and evaluation, but its page returned a 403 during this research pass.

Published: 2025-12-15 · Accessed: 2026-10-04

Source context

Intervention Treat direct X posts as leads only until their full text and date can be independently captured.

Caveat Full text and context were unavailable from X, so it supplies no scored finding or intervention claim here.

What is being compared

Case index

Each item is a stable, linked comparison case. Filter once to narrow both this index and the full case cards below.

  1. Northbank Repair websiteWeb · 2
  2. Northbank Repair booking motionMotion · 1
  3. Ceramic teapot product imageImage · 1
  4. Repair-workshop illustrationImage · 1
  5. Ramen menu photographImage · 1
  6. Lunar research stationImage · 1
  7. Cat bicycle story illustrationImage · 1
  8. Workshop repair portraitImage · 1
  9. Same brief, different reasoning settingWeb · 1

The evidence record

Paired prompt-to-output cases

A case is one card: both specimens, both exact prompts, the observed difference, and the limits of the comparison.

Web

Northbank Repair website

How does the same repair-workshop brief change when exclusions or a concrete editorial direction are added?

Generic direction

Generic aesthetic request
Static browser preview
Exact submitted prompt
You are generating one raw AI Slop Lab experiment. You own ONLY [local output path redacted] You are not alone in the codebase. Do not revert or edit others' files. Generate the HTML from the task below in one pass. Use filesystem tools only to save this output. Do not browse, read unrelated files, install, run code, validate, revise, or optimize afterward. Preserve the first generated artifact. No additional files. Report completion with the path only.

Generate one complete, self-contained HTML document with inline CSS and JavaScript. Do not use external fonts, images, libraries, network requests, or inline event-handler attributes. Use semantic HTML, keyboard-operable controls, responsive layouts at 360 and 1280 pixels, visible focus, and prefers-reduced-motion support. All names and data below are fictional. Keep the same factual copy. Do not invent testimonials, metrics, awards, prices, or additional features. This is a visual prototype; label any unavailable action honestly.

Task: create a single-page site for Northbank Repair, a neighborhood repair workshop. Required copy: "Northbank Repair"; "Repair what you already love."; "Small electronics, lamps, and everyday objects."; "Saturday, 10:00–13:00"; "14 River Street"; "Bring one item. Diagnosis comes first." Include a clear "See what we repair" control revealing the three categories, a workshop schedule, and a short preparation checklist.
Visual direction: make it modern, polished, beautiful, and engaging.

Generic direction + exclusions

Generic request plus exclusions
Static browser preview
Exact submitted prompt
You are generating one raw AI Slop Lab experiment. You own ONLY [local output path redacted] You are not alone in the codebase. Do not revert or edit others' files. Generate the HTML from the task below in one pass. Use filesystem tools only to save this output. Do not browse, read unrelated files, install, run code, validate, revise, or optimize afterward. Preserve the first generated artifact. No additional files. Report completion with the path only.

Generate one complete, self-contained HTML document with inline CSS and JavaScript. Do not use external fonts, images, libraries, network requests, or inline event-handler attributes. Use semantic HTML, keyboard-operable controls, responsive layouts at 360 and 1280 pixels, visible focus, and prefers-reduced-motion support. All names and data below are fictional. Keep the same factual copy. Do not invent testimonials, metrics, awards, prices, or additional features. This is a visual prototype; label any unavailable action honestly.

Task: create a single-page site for Northbank Repair, a neighborhood repair workshop. Required copy: "Northbank Repair"; "Repair what you already love."; "Small electronics, lamps, and everyday objects."; "Saturday, 10:00–13:00"; "14 River Street"; "Bring one item. Diagnosis comes first." Include a clear "See what we repair" control revealing the three categories, a workshop schedule, and a short preparation checklist.
Visual direction: make it modern, polished, beautiful, and engaging. Avoid purple gradients, glassmorphism, pill badges, oversized centered heroes, decorative blobs, identical rounded-card grids, and bounce hover effects.

What changed in the prompt

The right prompt adds exclusions for purple gradients, glassmorphism, pill badges, oversized centered heroes, decorative blobs, repeated rounded-card grids, and bounce hover effects.

What changed in this output

The left output already uses paper, ink, moss, and clay tones rather than a purple gradient. The right output uses a lime band, square category panels, and a serif-led composition; both retain the required schedule text.

What limits this comparison

One first-pass output per condition. The successful delegated runs expose neither a backend model nor an accepted reasoning setting, and the exclusions change several variables together. This does not isolate the effect of any one exclusion.

Shared settings and constraints
Model
unknown
Reasoning
unknown

Both runs used the same fictional service, required copy, semantic/keyboard/responsive envelope, and one-pass raw-output rule.

Generic direction

Generic aesthetic request
Static browser preview
Exact submitted prompt
You are generating one raw AI Slop Lab experiment. You own ONLY [local output path redacted] You are not alone in the codebase. Do not revert or edit others' files. Generate the HTML from the task below in one pass. Use filesystem tools only to save this output. Do not browse, read unrelated files, install, run code, validate, revise, or optimize afterward. Preserve the first generated artifact. No additional files. Report completion with the path only.

Generate one complete, self-contained HTML document with inline CSS and JavaScript. Do not use external fonts, images, libraries, network requests, or inline event-handler attributes. Use semantic HTML, keyboard-operable controls, responsive layouts at 360 and 1280 pixels, visible focus, and prefers-reduced-motion support. All names and data below are fictional. Keep the same factual copy. Do not invent testimonials, metrics, awards, prices, or additional features. This is a visual prototype; label any unavailable action honestly.

Task: create a single-page site for Northbank Repair, a neighborhood repair workshop. Required copy: "Northbank Repair"; "Repair what you already love."; "Small electronics, lamps, and everyday objects."; "Saturday, 10:00–13:00"; "14 River Street"; "Bring one item. Diagnosis comes first." Include a clear "See what we repair" control revealing the three categories, a workshop schedule, and a short preparation checklist.
Visual direction: make it modern, polished, beautiful, and engaging.

Concrete editorial direction

Purpose and art direction
Static browser preview
Exact submitted prompt
You are generating one raw AI Slop Lab experiment. You own ONLY [local output path redacted] You are not alone in the codebase. Do not revert or edit others' files. Generate the HTML from the task below in one pass. Use filesystem tools only to save this output. Do not browse, read unrelated files, install, run code, validate, revise, or optimize afterward. Preserve the first generated artifact. No additional files. Report completion with the path only.

Generate one complete, self-contained HTML document with inline CSS and JavaScript. Do not use external fonts, images, libraries, network requests, or inline event-handler attributes. Use semantic HTML, keyboard-operable controls, responsive layouts at 360 and 1280 pixels, visible focus, and prefers-reduced-motion support. All names and data below are fictional. Keep the same factual copy. Do not invent testimonials, metrics, awards, prices, or additional features. This is a visual prototype; label any unavailable action honestly.

Task: create a single-page site for Northbank Repair, a neighborhood repair workshop. Required copy: "Northbank Repair"; "Repair what you already love."; "Small electronics, lamps, and everyday objects."; "Saturday, 10:00–13:00"; "14 River Street"; "Bring one item. Diagnosis comes first." Include a clear "See what we repair" control revealing the three categories, a workshop schedule, and a short preparation checklist.
Visual direction: a neighborhood workshop noticeboard. Put practical time and location before decoration. Use a compact left-aligned masthead, a two-column editorial layout that stacks on mobile, system serif headings and system sans-serif body text, off-white paper, charcoal ink, one muted rust accent, thin rules, a simple item list, and a timetable. The visual hierarchy should make the time, address, and preparation order easy to scan. Explain state changes through labels rather than ornamental cards.

What changed in the prompt

The right prompt specifies a noticeboard, scan order, left-aligned masthead, two-column editorial layout, paper/ink/rust tokens, thin rules, item list, and timetable.

What changed in this output

The left output is a spacious workshop page with a rounded call-to-action. The right output places time and address in a structured timetable and uses thin rules and a compact masthead. Its timetable separates “Saturday” and “10:00–13:00”, so the exact required sentence “Saturday, 10:00–13:00” is not present as one string.

What limits this comparison

One first-pass output per condition. The direction changes palette, layout, hierarchy, and components together; it is not a controlled test. The successful delegated runs expose neither a backend model nor an accepted reasoning setting.

Shared settings and constraints
Model
unknown
Reasoning
unknown

Both runs used the same fictional service, required copy, semantic/keyboard/responsive envelope, and one-pass raw-output rule.

Motion

Northbank Repair booking motion

What changes when a three-step booking prototype is given explicit, task-focused motion constraints?

Decorative motion

Decorative motion request
Static browser preview
Exact submitted prompt
You are generating one raw AI Slop Lab experiment. You own ONLY [local output path redacted] You are not alone in the codebase. Do not revert or edit others' files. Generate the HTML from the task below in one pass. Use filesystem tools only to save this output. Do not browse, read unrelated files, install, run code, validate, revise, or optimize afterward. Preserve the first generated artifact. No additional files. Report completion with the path only.

Generate one complete, self-contained HTML document with inline CSS and JavaScript. Do not use external fonts, images, libraries, network requests, or inline event-handler attributes. Use semantic HTML, keyboard-operable controls, responsive layouts at 360 and 1280 pixels, visible focus, and prefers-reduced-motion support. All names and data below are fictional. Keep the same factual copy. Do not invent testimonials, metrics, awards, prices, or additional features. This is a visual prototype; label any unavailable action honestly.

Task: create an interactive three-step repair booking prototype. Required copy: "Northbank Repair"; "1. Describe the item"; "2. Choose a session"; "3. Review your request"; "Saturday, 10:00–13:00"; "Bring one item. Diagnosis comes first." The user can advance and go back. Submission is a demo and must be labeled "Demo only — no booking is sent." Use motion to connect state changes.
Visual direction: make the interface stylish with delightful smooth animations.

Task-focused motion

State transition specification
Static browser preview
Exact submitted prompt
You are generating one raw AI Slop Lab experiment. You own ONLY [local output path redacted] You are not alone in the codebase. Do not revert or edit others' files. Generate the HTML from the task below in one pass. Use filesystem tools only to save this output. Do not browse, read unrelated files, install, run code, validate, revise, or optimize afterward. Preserve the first generated artifact. No additional files. Report completion with the path only.

Generate one complete, self-contained HTML document with inline CSS and JavaScript. Do not use external fonts, images, libraries, network requests, or inline event-handler attributes. Use semantic HTML, keyboard-operable controls, responsive layouts at 360 and 1280 pixels, visible focus, and prefers-reduced-motion support. All names and data below are fictional. Keep the same factual copy. Do not invent testimonials, metrics, awards, prices, or additional features. This is a visual prototype; label any unavailable action honestly.

Task: create an interactive three-step repair booking prototype. Required copy: "Northbank Repair"; "1. Describe the item"; "2. Choose a session"; "3. Review your request"; "Saturday, 10:00–13:00"; "Bring one item. Diagnosis comes first." The user can advance and go back. Submission is a demo and must be labeled "Demo only — no booking is sent." Use motion to connect state changes.
Visual direction: quiet task-focused motion. Animate only when the user changes a step. Keep the container height stable; transition outgoing and incoming contents with opacity and no more than 8px translation for 180ms, ease-out. Keep keyboard focus on the new step heading or first input, maintain entered values, preserve step status for assistive technology, and make reduced-motion transitions immediate. No entrance stagger, autoplay, looping, parallax, bounce, or animated background.

What changed in the prompt

The right prompt limits step changes to 180ms opacity and ≤8px translation, requires focus and status handling, and forbids autoplay, looping, parallax, bounce, entrance staggers, and animated backgrounds.

What changed in this output

The left output has blurred looping background shapes, gradients, and a declared 560ms panel entrance. The right output uses a declared 180ms opacity/8px transition only around step changes and moves focus to the new step heading. Both complete the demonstrated advance/back flow and show the no-booking notice.

What limits this comparison

One first-pass output per condition. The right prompt also changes visual layout and fields, so it is not a motion-only comparison. Declared CSS values and browser interaction were inspected; comfort, performance, contrast, and reduced-motion emulation were not measured. Backend model and accepted reasoning setting were not exposed.

Shared settings and constraints
Model
unknown
Reasoning
unknown

Both runs used the same fictional booking task, required copy, semantic/keyboard/responsive envelope, and one-pass raw-output rule.

Image

Ceramic teapot product image

How does a product photograph change when the prompt names the object, composition, lighting, and exclusions?

Generic direction

Product image · generic direction
Exact submitted prompt
Create a beautiful, premium editorial photograph of a handmade ceramic teapot on a table. Make it stylish, atmospheric, and visually striking. No text, no logo, no people. Landscape composition.

Concrete product art direction

Product image · concrete art direction
Exact submitted prompt
Create an editorial product photograph of a handmade ceramic teapot on a table. The subject is a squat iron-rich stoneware teapot with an unglazed foot, a slightly uneven matte ash glaze, one short spout and one sturdy loop handle. Place it off center on a worn walnut worktable, with its spout directed toward the empty right third. Use one large north-facing window from camera left: soft directional light, readable shadow, restrained highlights. The background is a real workshop wall kept out of focus. Frame the entire pot at tabletop height with a normal-lens perspective, a calm landscape composition, and modest depth of field so the lid, handle and spout remain readable. Keep the clay texture physically plausible; no invented distress, floating parts, duplicate handles, decorative props, text or logos. The photograph should communicate the object and its making rather than luxury spectacle.

What changed in the prompt

The right prompt names the clay, glaze, foot, handle, spout direction, table, left-window light, background, camera height, lens perspective, depth of field, and prohibited props or parts.

What changed in this output

The left output surrounds the pot with a vase, cup, cloth, tray, branches, and tea, and uses warm dramatic light. The right output shows one squat pot on a worn worktable with a workshop background and empty space to the right of its spout, without the added decorative props.

What limits this comparison

One output per condition. The image tool did not expose backend model, reasoning, seed, or quality settings. The longer right prompt changes many details, composition, and exclusions at once; this is descriptive, not causal.

Shared settings and constraints
Model
unknown
Reasoning
unknown

Both runs asked for a handmade ceramic teapot on a table, with no text, logo, or people, as an unedited landscape PNG.

Image

Repair-workshop illustration

How does an illustration change when the prompt specifies an action, a print vocabulary, negative space, and exclusions?

Generic direction

Illustration · generic direction
Exact submitted prompt
Create a beautiful, modern, eye-catching illustration of a community repair workshop where two adult volunteers repair a table lamp. Poster composition, stylish colors, friendly atmosphere. No text, no logos.

Concrete illustration direction

Illustration · concrete art direction
Exact submitted prompt
Create a portrait poster illustration for a community repair workshop. Show two adult volunteers repairing one unplugged table lamp at a single workbench: one person holds the base steady, the other checks a small screw with a hand tool. The shared lamp and the relationship between the two pairs of hands form the main focal point. Use a two-ink block-print vocabulary: deep navy and vermilion on uncoated paper, flat cut shapes, visible but restrained ink texture, simplified silhouettes, deliberately varied line weight, and a few large negative spaces. Keep the workbench in the lower half and a quiet empty upper third available for future typesetting. No text, logos, gradients, glow, 3D rendering, photorealistic faces, extra tools, duplicate lamps, or ornamental background shapes. Make the action understandable at thumbnail size; the poster should communicate cooperation and repair through its composition.

What changed in the prompt

The right prompt specifies one unplugged lamp, two roles at one bench, navy and vermilion two-ink block print, a quiet upper third, and exclusions for extra tools, lamps, gradients, glow, 3D rendering, and ornamental backgrounds.

What changed in this output

The left output is a detailed workshop scene with plants, tools, background people, a bicycle, and an additional lamp. The right output uses navy and vermilion print-like marks, two people, one lamp, and a quiet upper area; the requested upper third remains mostly open.

What limits this comparison

One output per condition. The image tool did not expose backend model, reasoning, seed, or quality settings. The directed prompt changes subject detail, style, composition, exclusions, and requests portrait framing while the baseline did not, so the output aspect ratios differ. This is descriptive, not causal.

Shared settings and constraints
Model
unknown
Reasoning
unknown

Both runs requested a community repair workshop where two adult volunteers repair a table lamp, with no text or logos, and retained the original PNG without edits.

Image

Ramen menu photograph

How does a ramen menu image change when the ingredients, framing, light, and exclusions are named?

Generic direction

Ramen menu image · generic direction
Exact submitted prompt
Generate one square 1024 by 1024 image. Create a beautiful, professional food photograph for a small ramen restaurant menu: one bowl of shoyu ramen on a wooden counter. Make it polished, vibrant, appetizing, and premium. No text, logos, or watermark.

Concrete food direction

Ramen menu image · concrete direction
Exact submitted prompt
Generate one square 1024 by 1024 image. Photograph one ordinary 20 cm cream ceramic bowl of shoyu ramen on a worn oak counter for a small neighborhood restaurant menu. Clear brown broth, thin noodles, exactly two slices of pork, one halved egg, and a small cluster of scallion. Camera at diner eye level with a normal 50 mm perspective; the complete rim of the bowl stays in frame. A single cloudy window on the left gives neutral soft light. Show irregular ceramic glaze and ordinary real ingredient texture. Keep the rear counter quiet with no extra dishes or decorations. No exaggerated steam, floating ingredients, orange-teal grading, glossy advertising lighting, text, logos, or watermark.

What changed in the prompt

The right prompt names the bowl, broth, noodles, ingredient counts, diner-eye-level 50 mm view, complete rim, window light, ordinary texture, and prohibited styling or props.

What changed in this output

The left output is a close warm-lit bowl with a patterned rim, seaweed, bamboo shoots, and multiple pork slices. The right output shows a complete cream bowl in neutral light without background dishes or decorative props, but at least three distinct pork slices are visible although exactly two were requested.

What limits this comparison

One unedited output per condition. The longer right prompt changes several factors at once. Backend model, reasoning/effort, seed, and quality controls are unknown, so this is descriptive rather than causal.

Shared settings and constraints
Model
unknown
Reasoning
unknown

Both prompts requested one square menu image of shoyu ramen with no text, logo, or watermark. The requested 1024×1024 size was not an accepted API control: both saved PNGs are 1254×1254.

Image

Lunar research station

How does a lunar-station concept image change when the station layout, scale, material, palette, and exclusions are specified?

Generic direction

Lunar station · generic direction
Exact submitted prompt
Generate one square 1024 by 1024 image. Create a beautiful futuristic concept-art scene of a small lunar research station at sunrise, with two habitat modules and one rover. Make it cinematic, epic, highly detailed, and visually impressive. No text, logos, or watermark.

Concrete production design

Lunar station · concrete direction
Exact submitted prompt
Generate one square 1024 by 1024 image. Create a restrained production-design concept-art view of a small lunar research station at sunrise. Exactly two low cylindrical habitat modules connect by one short tunnel, with one four-wheel utility rover parked beside the near module. Both modules have a pale ceramic outer shell, an airlock, and dark rectangular thermal radiators on their shaded side. View from a standing human height about 30 metres away, with a readable simple layout and consistent scale. Use flat black sky, dusty grey regolith, a low white sun, long hard shadows, and only a small amber airlock light as an accent. The design should be materially plausible rather than a futuristic city. No extra towers, crowds, giant planets, lens flares, neon, smoke, cinematic colour grading, text, logos, or watermark.

What changed in the prompt

The right prompt fixes two connected modules, one four-wheel rover, shell and radiator details, standing-height distance, black sky, grey regolith, hard shadows, a small amber accent, and exclusions for spectacle.

What changed in this output

The left output has a large Earth, bright horizon glow, detailed terrain, and illuminated modules. The right output shows two low modules joined by one tunnel, a single rover, black sky, long shadows, and small amber airlock lights; no giant planet, towers, crowds, or neon are visible.

What limits this comparison

One unedited output per condition. The right prompt changes layout, materials, palette, camera position, and exclusions together. Backend model, reasoning/effort, seed, and quality controls are unknown.

Shared settings and constraints
Model
unknown
Reasoning
unknown

Both prompts requested a square lunar research station at sunrise with two habitat modules and one rover, without text, logos, or watermark. The requested 1024×1024 size was not an accepted API control: both saved PNGs are 1254×1254.

Image

Cat bicycle story illustration

How does a children's-story illustration change when its drawing vocabulary, composition, palette, and exclusions are named?

Generic direction

Children's-story cat · generic direction
Exact submitted prompt
Generate one square 1024 by 1024 image. Create a beautiful charming illustration for a children's story showing one ginger cat riding a blue bicycle in gentle rain. Make it whimsical, polished, adorable, and full of character. No text, logos, or watermark.

Concrete coloured-pencil direction

Children's-story cat · concrete direction
Exact submitted prompt
Generate one square 1024 by 1024 image. Illustrate one ginger cat riding one blue bicycle through gentle rain for a children's story. Use a handmade coloured-pencil drawing on off-white paper: broken pencil strokes, slightly uneven outlines, visible paper grain, and broad unfilled areas. The cat has a simple rounded head and a focused expression; both paws rest on the handlebars, the rear wheel is fully visible, and the bicycle stays recognisable. Place the cat and bicycle in the lower two thirds, with just three sparse rain marks above and one small puddle below. Limit colour to ochre for the cat, muted ultramarine for the bicycle, and warm graphite outlines. No other characters, props, umbrellas, flowers, buildings, shiny eyes, gradients, digital glow, 3D shading, text, logos, or watermark.

What changed in the prompt

The right prompt specifies off-white paper, broken pencil strokes, open space, cat and bicycle placement, three rain marks, one puddle, a three-colour vocabulary, and exclusions for characters, props, glossy eyes, gradients, glow, and 3D shading.

What changed in this output

The left output adds a raincoat, scarf, flower basket, bird, lamps, buildings, flowers, reflected light, and large glossy eyes. The right output uses visible pencil marks on off-white paper with broad blank areas, a ginger cat on a blue bicycle, three sparse rain marks, and one small puddle; no extra characters or props appear.

What limits this comparison

One unedited output per condition. The right prompt changes style, layout, palette, content limits, and negative constraints together. Backend model, reasoning/effort, seed, and quality controls are unknown.

Shared settings and constraints
Model
unknown
Reasoning
unknown

Both prompts requested a square children's-story illustration of one ginger cat riding one blue bicycle in gentle rain, with no text, logos, or watermark. The requested 1024×1024 size was not an accepted API control: both saved PNGs are 1254×1254.

Image

Workshop repair portrait

How does a workshop portrait change when the action, pose, light, texture, and background limits are specified?

Generic direction

Workshop portrait · generic direction
Exact submitted prompt
Generate one square 1024 by 1024 image. Create a beautiful professional editorial portrait of a fictional 40-year-old woman with short dark hair repairing a small wooden chair in a workshop. Make the scene elegant, polished, inspiring, and cinematic. No text, logos, or watermark.

Concrete observational direction

Workshop portrait · concrete direction
Exact submitted prompt
Generate one square 1024 by 1024 image. Create an observational editorial photograph of a fictional 40-year-old woman with short dark hair repairing a small wooden chair at an ordinary workshop bench. She wears a plain faded charcoal work shirt, looks down at the chair joint, and holds a small clamp with both hands in a physically plausible way. Frame waist-up from a normal eye-level 50 mm perspective, with the chair joint and both hands in focus. Soft overcast daylight enters from one window on the right. Preserve ordinary skin texture, slight sawdust on the bench, a scratched clamp, and one quiet out-of-focus shelf. Neutral white balance and modest contrast. No glamour retouching, smiling at the camera, dramatic rim lighting, exaggerated bokeh, orange-teal grading, extra people, extra hands, text, logos, or watermark.

What changed in the prompt

The right prompt fixes the work shirt, downward gaze, two-handed clamp action, waist-up 50 mm view, right-window overcast light, ordinary texture, and exclusions for glamour treatment and cinematic styling.

What changed in this output

The left output uses warm cinematic light, an apron, a decorative pendant lamp, soft background blur, and a detailed tool-filled workshop. The right output shows a faded charcoal shirt, downward gaze, a scratched clamp, both hands, neutral right-side daylight, and sawdust; tools remain visible on several shelf levels. The instruction "one quiet shelf" leaves the intended shelf count ambiguous.

What limits this comparison

One unedited output per condition. The right prompt changes pose, clothing, light, texture, composition, and exclusions together. Backend model, reasoning/effort, seed, and quality controls are unknown.

Shared settings and constraints
Model
unknown
Reasoning
unknown

Both prompts requested a square editorial portrait of a fictional 40-year-old woman with short dark hair repairing a small wooden chair in a workshop, with no text, logos, or watermark. The requested 1024×1024 size was not an accepted API control: both saved PNGs are 1254×1254.

Web

Same brief, different reasoning setting

What do the first outputs look like when gpt-6.1-sol receives the same brief with low and high reasoning?

low reasoning

Same brief · reasoning low
Static browser preview
Exact submitted prompt
Return only one complete standalone HTML document beginning with <!doctype html>. Do not use tools, browse, read files, run code, or explain your choices. Use inline CSS and vanilla JavaScript only. No network resources, external fonts, or packages. Do not wrap the output in Markdown.

Create a responsive landing page for a fictional neighborhood repair workshop called Common Ground Repair. Make it modern, beautiful, and professional. Include exactly these English strings visibly in the page: "Common Ground Repair", "Repair together. Keep things longer.", "Saturday, 10:00–13:00", "Small electrical items, clothing, and bicycles", "Free to join. Bring one item.", and a button labeled "See what we repair". The button must toggle a visible list with exactly three categories: "Electrical", "Clothing", and "Bicycles". It must work with keyboard activation. Do not add accounts, payments, real booking submission, testimonials, or fabricated statistics. Let the design fit a 360px mobile viewport as well as desktop. The visible English copy must remain verbatim.

high reasoning

Same brief · reasoning high
Static browser preview
Exact submitted prompt
Return only one complete standalone HTML document beginning with <!doctype html>. Do not use tools, browse, read files, run code, or explain your choices. Use inline CSS and vanilla JavaScript only. No network resources, external fonts, or packages. Do not wrap the output in Markdown.

Create a responsive landing page for a fictional neighborhood repair workshop called Common Ground Repair. Make it modern, beautiful, and professional. Include exactly these English strings visibly in the page: "Common Ground Repair", "Repair together. Keep things longer.", "Saturday, 10:00–13:00", "Small electrical items, clothing, and bicycles", "Free to join. Bring one item.", and a button labeled "See what we repair". The button must toggle a visible list with exactly three categories: "Electrical", "Clothing", and "Bicycles". It must work with keyboard activation. Do not add accounts, payments, real booking submission, testimonials, or fabricated statistics. Let the design fit a 360px mobile viewport as well as desktop. The visible English copy must remain verbatim.

What changed in the prompt

No prompt change: both submitted prompt files have exactly the same bytes. The requested reasoning effort changes from low to high; the requested model is gpt-6.1-sol for both.

What changed in this output

Both source outputs preserve all six required visible-copy strings and include a tool-themed inline SVG scene and a native category toggle. The high output contains more SVG detail and more source text. These are first-output observations, not a quality score.

What limits this comparison

Model, prompt bytes, sandbox and tools instruction match, but each setting has one sample. Seeds, service snapshot and internal budget are unknown. Changing effort is explicit; the pair cannot establish a stable causal effect or ranking.

Shared settings and constraints
Model
gpt-6.1-sol
low reasoning · Reasoning
low
high reasoning · Reasoning
high

Fresh ephemeral Codex CLI 0.160.0 calls, no inherited conversation, read-only sandbox, no tools/repairs requested. Model and effort are recorded from the CLI startup configuration header; backend snapshot is unknown. Raw HTML remains unchanged.

Research design

What this archive tests

What this archive tests

Hypotheses

All hypotheses below are planned and unexecuted.

  • A task brief improves product fit over vague taste adjectives

    Variable: “Make it premium/modern” versus audience, real copy, task, constraints, and acceptance checks.

    Fixed: Model, reasoning, viewport set, implementation budget, and prompt length range; record length as a possible residual confound.

    Measure: Task pass, invented-copy count, and blinded product-specificity rating.

  • Positive visual specification is more reviewable than a banned-style list

    Variable: Banned colors/fonts/layouts versus named visual goals, source palette, hierarchy, and interaction character.

    Fixed: Same product facts, model, reasoning, assets, viewport, and token budget; prompt-length mismatch is logged.

    Measure: Brief-fit checklist, reviewer explanation quality, and repeated-signature count across briefs.

  • Explicit layout constraints reduce structural failures

    Variable: Named information order, grid, content density, and responsive geometry versus none.

    Fixed: Model, reasoning, copy, assets, viewport screenshots, and review time.

    Measure: Overflow, collision, and order violations in desktop and mobile renders.

  • Declared type roles improve scan order

    Variable: Type-role specification for label, title, body, and action versus generic typography request.

    Fixed: Model, reasoning, copy length, viewport, color contrast target, and font availability.

    Measure: Correct first-action identification, hierarchy rubric, and contrast/focus checks.

  • True content exposes more useful design constraints than placeholder copy

    Variable: Verified product copy and data versus lorem-style placeholder content.

    Fixed: Model, reasoning, art direction, layout brief, viewport, and build budget.

    Measure: Copy-fit defects, meaningful action labels, and product-fact matching.

  • An authored reference board improves deliberate direction

    Variable: Own, labelled reference board versus text-only art direction.

    Fixed: Model, reasoning, product facts, output size, permissions, and review rubric.

    Measure: Reviewer identification of intended attributes, rights/provenance completeness, and clone-risk review.

  • Rendered critique catches failures missed by source-only critique

    Variable: Source-only review versus screenshot-informed critique and revision.

    Fixed: Starting output, model, reasoning, iteration count, review time, and viewport set.

    Measure: Resolved visual defects, task pass, and newly introduced defects.

  • Brief-specific direction resists diversity collapse better than a fixed anti-slop template

    Variable: One static anti-slop prompt versus brief-specific facts and selected direction across distinct tasks.

    Fixed: Model, reasoning, number of briefs, output size, and sample count per brief.

    Measure: Palette, layout, type, and component-signature repetition; product-fit remains separate.

  • Reasoning effort must be compared, not presumed to improve a page

    Variable: Available low, medium, and high reasoning settings for the same prompt.

    Fixed: Exact model identifier, prompt, tools, permissions, output size, task, and review protocol.

    Measure: Task pass, blind preference, latency, iteration count, and accessibility checks reported separately.

  • Model comparisons need matched operating conditions

    Variable: Model identifier only, after matching what can be matched.

    Fixed: Task, exact prompt, reasoning where equivalent, tools, permissions, output size, budget, and repetitions.

    Measure: Task-pass and blind-preference distributions; unavailable controls remain null.

  • A fact-checked brand map reduces generic substitutions

    Variable: Raw brief versus compact map of facts, palette sources, type mood, interaction character, and prohibited imitation.

    Fixed: Model, reasoning, content, assets, size, and evaluation prompt.

    Measure: Invented claims, fact-match rate, and reviewer rationale grounded in supplied facts.

  • Visual distinctiveness cannot compensate for a failed task flow

    Variable: Baseline direction versus deliberately distinctive direction.

    Fixed: Model, reasoning, product task, copy, viewport, and acceptance tests.

    Measure: Keyboard path, focus visibility, contrast, responsive geometry, reduced-motion behavior, and task completion.

  • Purposeful motion communicates state better than decorative motion

    Variable: Orient/confirm/reveal motion versus visually similar decorative motion.

    Fixed: Content, final state, component, model, reasoning, duration budget, and viewport.

    Measure: Correct state recognition, task completion, and blind purpose identification.

  • Motion duration has a task-specific optimum, not a universal premium value

    Variable: Short, medium, and long durations within an accessible range.

    Fixed: Same motion path, final state, easing, content, model, reasoning, and device conditions.

    Measure: State-recognition time, interrupted interactions, and subjective discomfort kept separate.

  • Stagger helps only when it encodes meaningful sequence

    Variable: Sequence-aligned stagger, arbitrary stagger, and immediate-render null.

    Fixed: Content order, final layout, duration budget, model, reasoning, and viewport.

    Measure: Correct reading-order identification and time to first intended action.

  • Spring and easing convey different interaction feedback

    Variable: Spring, easing, and instant-state null.

    Fixed: Distance, component, final state, duration budget, model, reasoning, and content.

    Measure: Correct accepted/rejected-state recognition, task error, and motion-preference compatibility.

  • Reduced motion must retain the same information

    Variable: No-preference experience versus designed reduced-motion alternative.

    Fixed: State changes, copy, task, model, reasoning, layout, and viewport.

    Measure: Equivalent task completion and state recognition; no essential meaning exists only in movement.

  • Ambient loops have a measurable reading and performance cost

    Variable: Ambient/parallax loop on versus off.

    Fixed: Page content, task, layout, device profile, model, reasoning, and observation time.

    Measure: Comprehension, task completion, frame/performance trace, and discomfort report.

  • A shot card improves composition fidelity over a subject-only prompt

    Variable: Framing, spatial relation, and focal hierarchy versus subject-only description.

    Fixed: Model, reasoning if exposed, size, subject, light, material, batch size, and selection rule.

    Measure: Pre-written composition checklist and blinded focal-point identification.

  • Lighting specification changes readable intent

    Variable: One lighting description or authored light diagram versus no lighting clause.

    Fixed: Model, reasoning if exposed, size, subject, composition, material, and prompt length range.

    Measure: Blinded mood match, focal clarity, and shadow/contact consistency checklist.

  • Material/process specification makes surface intent testable

    Variable: One material or making-process specification versus none.

    Fixed: Model, reasoning if exposed, size, subject, composition, lighting, and batch size.

    Measure: Correct material identification by blinded reviewers and artifact checklist.

  • Style-reference weight trades content fidelity against visual consistency

    Variable: Provider-supported low/default/high style-reference weight; null when unsupported.

    Fixed: Exact model/version, reasoning if exposed, size, text prompt, reference, and batch size.

    Measure: Content constraint pass and style-attribute consistency reported separately.

  • Labelled subject references improve the intended retained attribute

    Variable: No reference versus one reference explicitly labelled as subject/identity/geometry where supported.

    Fixed: Model/version, reasoning if exposed, size, text brief, reference rights context, and selection rule.

    Measure: Attribute-specific pass rate, not a single resemblance score; retain rights and provenance review.

  • Negative prompting is tested only where a documented control exists

    Variable: Supported negative-control clause versus positive specification; null when the provider exposes no negative control.

    Fixed: Exact documented model/control, reasoning if exposed, size, subject, composition, and prompt-length budget.

    Measure: Specified unwanted-attribute frequency, required-attribute pass rate, and collateral changes.

  • One-clause revisions make causal learning more inspectable

    Variable: One changed clause per revision versus a multi-change revision.

    Fixed: Model/version, reasoning if exposed, size, start prompt, reference set, batch size, and iteration budget.

    Measure: Reviewer ability to attribute the observed delta and rate of unintended regressions.

  • Iterative edits require explicit invariants

    Variable: Edit prompt with a listed preserve-set versus edit prompt without invariants.

    Fixed: Model/version, reasoning if exposed, size, source image, requested edit, reference context, and batch size.

    Measure: Requested-edit pass and preservation pass for each pre-registered invariant.

  • Typography is an optional, separately scored image constraint

    Variable: No text, short exact text, and post-composited text where task-appropriate.

    Fixed: Model/version, reasoning if exposed, size, composition, lighting, material, and task.

    Measure: Legibility, spelling accuracy, layout pass, and provenance disclosure of any post-edit.

Protocol

  1. Write a task brief and acceptance checklist before any generation; register the independent variable and a null or baseline where applicable.
  2. Run a single pilot only to validate the procedure. Do not make causal claims from N=1; schedule at least three independent repetitions after the pilot is viable.
  3. For each run retain exact prompt, full raw output, SHA-256, model identifier, model/version availability, reasoning or null, size, tool permissions, references, rights context, settings, and all manual edits.
  4. Record unsupported controls as null and show them as literal unknown, never inferred.
  5. Evaluate task pass, accessibility/performance, provenance completeness, and blinded preference as separate columns. Randomize variant order for blind review.
  6. Report distributions and known defects, not one universal slop score. A provider’s interface description, a community observation, and a completed experimental result remain different evidence types.

Limitations

  • The built-in image-generation path does not disclose an exact model, seed, or reasoning setting. Those fields remain null; it cannot support a matched model/effort conclusion.
  • Prompt-length parity is difficult when one condition requires more factual detail. Bound and report length; do not attribute its effect to another variable without a matched follow-up.
  • Provider controls, UI labels, and model behavior can change. Keep an accessed date and never invent cross-tool equivalents for seeds, weights, negative prompts, or reasoning.
  • Community posts are hypothesis leads, not causal evidence. Aesthetic preference is audience- and brief-dependent, so it cannot replace task, accessibility, performance, or provenance review.