Morphic animated logo
AI + Creative Intelligence

What Happens When You Ask AI to Design Before It Plans?

Short answer: We tested the same AI on the same company brief two ways: one homepage was designed immediately, while the other was planned before it was designed. Planning improved the page’s logic, factual discipline, and coherence, but it did not consistently improve the creative work. Some of the unplanned version’s individual decisions were stronger. The bigger lesson was that the plan itself needs evaluation before its judgments become constraints.
By Weston Baker · Founder, Morphic

AI makes it possible to see a designed page almost immediately. That creates an obvious temptation: why plan first when we can just generate the page and react to it?

We tested whether a planning step actually changes the quality of the work, rather than simply making the process feel more deliberate. We gave the same AI the same fictional company brief and asked it to create the same homepage two different ways for Meridian Peak Capital.

The setup

Condition A

Design immediately

The model received the Meridian Peak Capital brief and generated the homepage with no separate planning stage.

Condition B

Plan before design

Before designing anything, the model determined page objective, visitor state, positioning, message hierarchy, audience priorities, proof requirements, factual constraints, objections, narrative sequence, and section roles. It then designed the homepage using that plan.

We expected the planned version to be clearly better. It wasn't.

The planned homepage was more coherent, more disciplined with facts, and stronger at preserving the relationship between a claim and the information supporting it. But the design-first version produced some of the stronger individual creative decisions. That tension is the central story of this article.

What the planning step decided

Page job
Treat the homepage as a credibility filter and routing mechanism rather than a conversion funnel.
Primary audience
Prioritize founders while accounting for LPs and operators.
Order of ideas
Put investment criteria before philosophy so founders can determine fit quickly.
Factual constraints
Identify missing portfolio, team, performance, testimonial, and press information as unknowns rather than fabricate them.
Credibility logic
Combine philosophy, founder/operator origin, and functional scope so claims sit beside related proof.
Operating areas
Avoid a separate six-item operating-area section because the plan predicted it could become a generic card grid.

Results: design immediately

The design-first version produced the stronger individual headline:

“Your strategy doesn't need to change. Your decisions get harder.”

It was sharper and more distinctive than the planned version's safer opening.

It also elevated Meridian Peak's six operating areas into a dedicated “Where we get involved” section. Instead of becoming a generic SaaS card grid, the section was executed as a restrained editorial list that made the firm's operating role concrete and scannable.

Where we get involved
  • Go-to-market strategy
  • Leadership development
  • Pricing
  • Market expansion
  • Operating infrastructure
  • Strategic decision-making

The evaluators considered that a meaningful strength.

But the design-first version made weaker strategic decisions elsewhere. It placed the philosophy before the investment criteria, meaning a time-constrained founder encountered an abstract belief before learning whether Meridian Peak was relevant.

More importantly, it dropped the strongest credibility fact supplied in the brief: Meridian Peak was founded by investors and operators. The page asserted that strong companies need experienced partners but omitted the strongest available factual reason to believe Meridian Peak could provide that experience.

The design-first page also had production issues including mobile navigation disappearing and visible contact placeholders. These affect launch readiness but should not be attributed to the absence of planning. They are execution and QA failures.

Results: plan before design

The planned version was more disciplined. Its sequence was:

Positioning Investment criteria Philosophy and credibility Audience routing Invitation

It preserved the founder-and-operator fact and placed that fact directly beside the philosophy it supported. The source-aware evaluator identified this as the clearest planning-related improvement.

Planning also made missing proof explicit before design. Because portfolio companies, named team members, fund statistics, testimonials, and press had not been supplied, the plan excluded those elements instead of allowing plausible-looking evidence to be invented.

But planning did not improve everything. The plan predicted that giving the six operating areas their own section might create a generic SaaS-style card grid, so it compressed those ideas into prose. The design-first output demonstrated that this predicted failure was not inevitable: it created a restrained editorial list that evaluators preferred for specificity and scannability.

The plan identified a plausible failure mode, but then converted its prediction into a constraint before seeing an actual design.

Results: what planning actually changed

Strategic discipline
Improved
The planned version had a clearer objective, hierarchy, and narrative sequence.
Factual integrity and proof continuity
Improved
Planning made unknowns explicit and kept the strongest available credibility fact attached to the claim it supported.
Communication quality
Mixed
The planned page was more coherent, while the design-first page produced the stronger headline and more concrete treatment of operating areas.
Visual quality
No clear advantage
Both pages were restrained and credible. Differences were execution choices rather than effects traceable to planning.
Production readiness
Not attributable to planning
Both outputs contained implementation issues. Responsive behavior, working links, contact mechanisms, accessibility, and final QA require a separate evaluation layer.

“Planning helped most where it constrained facts, proof, objectives, and argument structure. It helped least when it tried to pre-decide creative form before the alternative had actually been designed.”

A plan can contain bad decisions too

The plan was also generated by AI. It can therefore contain strong judgments, weak judgments, assumptions, overcorrections, and reasonable-sounding predictions that have not been tested.

  • “Put criteria before philosophy” was a useful strategic judgment.
  • “Keep the credibility fact beside the philosophy it supports” was useful.
  • “Do not elevate these six operating areas because the result may become a generic grid” was a creative hypothesis, not a fact.

The actual design demonstrated that the six items could be elevated without producing the predicted failure.

The lesson is not simply that AI should plan before it designs. The plan itself needs judgment.

Constraints vs. hypotheses

Binding constraints

What should strongly constrain generation

Known company facts, factual unknowns, audience requirements, explicit objectives, proof relationships, and anti-fabrication rules.

Strategic hypotheses

What deserves strong consideration

Narrative sequence, section roles, message prominence, and where objections should be resolved. These remain open to evaluation.

Creative hypotheses

What should remain challengeable

Prose versus list, whether something gets its own section, composition choices, and predicted visual failure modes. These are inputs to exploration, not binding instructions.

The evolved workflow

Context
Plan
Evaluate the plan
Generate
Evaluate the result
Refine
  • Context defines the company and problem.
  • Planning makes communication decisions explicit.
  • Plan evaluation tests those decisions before they become constraints.
  • Generation explores the creative execution.
  • Output evaluation determines whether the work actually solved the problem.
  • Refinement applies what was learned.

“A plan should not become truth simply because the AI wrote it before it designed.”

Generation benefits from planning.

Planning benefits from judgment.

Methodology limitations

This was one fictional company, one homepage task, one model/version, and one pair of generated outputs. A blind evaluation was followed by a source-aware evaluation using the original brief and planning artifact. The findings describe what happened in this test and should not be generalized across all models, companies, or creative tasks without repeated testing.

Common questions

Frequently asked questions

What our experiment suggests about planning, generation, and where judgment belongs in an AI creative workflow.

01

Should AI plan a website before designing it?

Usually, but planning should not be treated as automatically correct. In our test, planning improved the homepage's strategic discipline, factual integrity, and argument structure. It did not consistently improve copy or creative execution. The strongest workflow evaluates the plan before turning its judgments into constraints.

02

What should an AI page plan include?

At minimum, it should define the page objective, audience priorities, visitor arrival and desired end states, message hierarchy, proof requirements, factual constraints, objections, narrative sequence, and the role of each proposed section. It should also distinguish known facts from strategic and creative hypotheses.

03

Can planning make AI-generated design worse?

Yes. A plan can contain plausible but overly restrictive judgments. In our experiment, the plan compressed six concrete operating areas into prose because it predicted that a dedicated section might become a generic card grid. The design-first version created a restrained editorial list instead, and evaluators preferred that treatment.

04

What parts of an AI page plan should be binding?

Facts, factual unknowns, audience requirements, important proof relationships, and anti-fabrication rules should be treated as strong constraints. Narrative and section decisions are strategic hypotheses. Specific composition or formatting choices should usually remain more flexible until the system can evaluate the actual execution.

05

How do you know whether planning improved an AI-generated website?

Evaluate the finished outputs against the same brief and rubric rather than assuming the planned version is better. Look separately at strategic discipline, factual integrity, communication quality, visual quality, production readiness, and correction cost. Blind evaluation can also help reduce bias about which workflow is supposed to win.

Morphic

Turn better thinking into better creative work.

Morphic combines company context, creative intelligence, design systems, and AI to help teams create better websites and brand materials—and improve them over time.