Morphic animated logo
AI + Creative Intelligence

Why Did ChatGPT Say the Website It Built for Me Isn’t Good?

Short answer: ChatGPT can successfully follow the instructions used to build a website without consistently applying the same standards it uses when later asked to critique the finished result. Generation, evaluation, and site-wide creative judgment are different tasks. A polished output can therefore satisfy the prompts that created it and still have weaknesses in positioning, hierarchy, consistency, or differentiation.
By Weston Baker · Founder, Morphic

There’s a strange experience becoming increasingly common with AI-built websites.

You ask ChatGPT to help you build a site. It writes the copy, suggests the structure, designs sections, generates code, and helps you make revisions. The result looks pretty good.

Then you ask a different question:

“Is this actually a good website?”

Suddenly, the same AI that helped create it starts finding problems.

The positioning could be clearer. The page feels generic. There’s too much repetition. The visual hierarchy is weak. The proof points aren’t strong enough. The design looks like other AI-generated websites.

So why didn’t it fix those things while it was building the site?

Because generating something and judging whether it is good are different tasks.

AI is very good at giving you what you ask for

Most AI website workflows are iterative. You make a reasonable request, get a reasonable result, and then keep building from there.

Prompt 01
Build a hero for this company.
Prompt 02
Make the headline bigger.
Prompt 03
Add a section explaining our services.
Prompt 04
Make this feel more premium.
Prompt 05
Add some animation here.

Each individual request can produce a perfectly reasonable result.

The problem is that a great website isn’t simply the sum of a series of reasonable requests. It needs an underlying strategy.

Message
What should someone understand within the first ten seconds?
Audience
What matters most to the people this page is for?
Proof
What evidence should appear before a claim is expected to feel credible?
Hierarchy
Which ideas deserve emphasis, and which should stay quiet?
Rhythm
Where should the experience accelerate, slow down, repeat, or intentionally break a pattern?

Those decisions are easy to lose when a website is assembled one prompt at a time.

Creation and evaluation put AI into different modes

When you ask AI to create something, the primary objective is usually to satisfy the request.

When you ask it to critique something, the objective changes. Now it is looking for weaknesses.

An AI can comply with your request to add another three-column card section and later correctly observe that the page has too many repetitive card layouts.

It can write a headline you approve and later tell you that the positioning isn’t differentiated enough.

It can make every section independently attractive and later notice that the overall page lacks rhythm.

We think of this as the generation–evaluation gap: the difference between what an AI system is willing to generate and what it recognizes as high quality when explicitly asked to evaluate the result.

A website is a system, not a collection of sections

AI is often remarkably good at local optimization. Give it one section and it can improve that section. But improving one section does not necessarily improve the website.

Make every headline larger and eventually nothing feels important. Make every section more visually interesting and the page becomes exhausting. Add more explanation everywhere and the story gets harder to understand. Give every section a unique layout and the site loses consistency.

Good design requires balancing local decisions against the whole. That’s why an experienced designer will sometimes make an individual element less interesting because it makes the overall composition better.

The missing step is judgment

The breakthrough in generative AI has made creating things dramatically easier. But creation was never the entire job.

Good creative work also requires deciding what should be created, what matters most, what should be removed, what should remain consistent, when consistency should be broken, whether the result actually communicates what it needs to, and whether the result is good enough to ship.

Those are judgment problems. And they become more important as generation gets easier.

How to get better results

If you’re building a website with ChatGPT, Claude, or another general-purpose AI tool, don’t use it only as a generator. Separate the process into stages.

01
Establish the company’s positioning, audience, objectives, proof points, brand characteristics, and website strategy.
02
Plan the site and determine the role of each page.
03
Plan individual pages before solving individual sections.
04
Establish a visual system for typography, spacing, color, imagery, interaction, and repeated patterns.
05
Generate, then stop and evaluate the work before continuing to build on top of it.

Ask whether the page is accomplishing its purpose, not simply whether it looks good. Ask what feels generic, what information is missing, whether the hierarchy is clear, whether the design has drifted from earlier decisions, and what should be removed.

AI becomes much more useful when it isn’t responsible only for producing the next thing.

This is a bigger problem than website generation

At Morphic, this question has shaped how we think about creative AI more broadly.

The future isn’t simply better generation. A useful creative system needs to understand, plan, create, evaluate, refine, and remember.

It needs to retain the decisions that led to good work so the next thing doesn’t start from zero.

That’s the difference between an AI that can make things and a system that can help a company consistently make good things.

Common questions

Frequently asked questions

Short answers to the questions that tend to come up once people start evaluating AI-generated websites more critically.

01

Can ChatGPT accurately judge a website it created?

It can provide useful criticism, but its critique should not be treated as an objective verdict. The important point is that critique is a different task from generation. When explicitly asked to evaluate a finished site, the model may apply criteria—such as differentiation, hierarchy, consistency, and proof—that were not enforced throughout the generation process.

02

Why can ChatGPT criticize problems it created itself?

Because recognizing a weakness in finished work and preventing that weakness during open-ended creation are different problems. Generation involves choosing among many possible outputs while satisfying instructions; evaluation starts with a concrete artifact and asks what is wrong with it.

03

Does a better prompt solve this problem?

Better prompts help, but they do not automatically create persistent strategy, site-wide awareness, design-system enforcement, or an independent evaluation loop. For complex creative work, the process around the model matters as much as the wording of the prompt.

04

How do I know whether an AI-generated website is actually good?

Evaluate it separately from the conversation that produced it. Review positioning, audience clarity, narrative, proof, hierarchy, brand consistency, distinctiveness, usability, mobile behavior, and whether the site accomplishes the business objective it was created to serve.

05

Is this only a ChatGPT problem, or does it happen with Claude and other AI tools too?

It is not unique to ChatGPT. The same underlying issue can appear with Claude, Gemini, and other general-purpose AI systems: generating an output, maintaining site-wide creative context, and independently evaluating the finished work are different tasks. The exact behavior varies by model and workflow, but the broader generation–evaluation problem is not specific to one product.

Morphic

Turn better thinking into better creative work.

Morphic combines company context, creative intelligence, design systems, and AI to help teams create better websites and brand materials—and improve them over time.