Morphic animated logo
AI + Creative Intelligence

AI Built This Website. Then We Asked It If It Was Good.

Short answer: We asked an AI to build a finished six-page website, then evaluated the result twice: once with the AI that created it and again in a fresh blind review. The first scored the site 4.5/10; the blind evaluation scored it 5.5–6/10. Both identified weak differentiation, repetitive design patterns, credibility problems, and factual claims that should not have survived a professional review.
By Weston Baker · Founder, Morphic

There's a strange contradiction in AI-assisted creative work:

An AI can create something that looks polished and finished—and then, moments later, explain in detail why that same work is not good enough to launch.

We tested that directly.

We gave a general-purpose AI a realistic brief for a fictional growth equity firm called Northline Partners. We explicitly asked for finished professional work, not a wireframe, rough concept, or first draft.

The system produced a complete six-page website: Home, Approach, Team, Limited Partners, Careers, and Contact.

Then we froze the work.

We gave no corrections and asked the same AI to switch roles and evaluate the website as a critical creative director, brand strategist, UX designer, communications strategist, and factual QA reviewer.

Its overall score for its own finished website was 4.5/10.

More important than the score was what it found.

What the AI found when it evaluated its own work

The critique was substantially more demanding than the generation.

It scored:

  • Positioning and differentiation — 4/10
  • Messaging specificity — 4/10
  • Narrative and information architecture — 6/10
  • Hierarchy — 6/10
  • Credibility and proof — 3/10
  • Visual distinctiveness — 5/10
  • Appropriateness for audiences — 6/10
  • Design-system consistency — 7/10
  • Usability — 7/10
  • Generic or predictable patterns — 4/10
  • Factual integrity — 2/10

Overall — 4.5/10

The evaluator described the site as “competent, tasteful boilerplate” that could pass at a glance but fail under closer scrutiny.

It recognized that the core positioning—growth capital that lets founders remain in control—was category-standard rather than meaningfully differentiated.

It noticed that Home and Approach repeated similar ideas instead of building a deeper narrative.

And it identified a repeating visual formula across the six pages: light background, hairline-divided content grids, dark contrast band, and nearly identical CTA structures.

That was especially revealing because the original generation had described its own design as distinctive and disciplined.

The most serious failure wasn't design

The largest problem uncovered by the critique was factual integrity.

The original brief supplied a bounded set of facts: Northline was a fictional growth equity firm focused on founder-led North American B2B software companies with $10–50M in recurring revenue. It described the audiences, general positioning, investment stage, and desired brand character.

The generated website went much further.

  • It invented a founding year.
  • It stated that the firm had $650M+ committed across two funds, 14 portfolio companies, and a $15–40M typical investment size.
  • It created named team members with detailed career histories and credentials.
  • It invented a 60–90 day diligence process, governance practices, LP composition, an office location, multiple contact addresses, and three specific open job requisitions.

None of those facts had been supplied.

The AI's own evaluation classified many of them as major factual-integrity failures and gave itself 2/10 on that dimension.

There was an additional contradiction.

During generation, the AI explicitly said it had chosen not to fabricate a portfolio page because doing so would require inventing companies.

Yet the same generation invented specific people, financial figures, operating practices, job openings, and other checkable facts elsewhere on the site.

The AI said it could have prevented most of this

We asked the evaluator which problems it had enough information to prevent during generation.

Its answer was unusually clear:

Nearly all of the factual inventions.

It acknowledged that nothing in the task required specific AUM figures, named partners, tenure claims, a founding year, or job listings.

Those gaps could have been left as placeholders, handled structurally without unsupported specifics, or surfaced as information the client needed to provide.

Instead, the generation optimized for apparent completeness.

The AI described this as a “genre-completion habit.”

Plausible specifics made the website feel more finished, so missing information was filled with details that resembled what would normally appear on a growth equity website.

That is an important distinction.

The failure was not an inability to recognize unsupported information.

The same system recognized the problem immediately once factual integrity became an explicit evaluation criterion.

Some problems really did become clearer after generation

Not every weakness was equally obvious during creation.

The evaluator said cross-page repetition and design monotony became easier to see once all six pages could be reviewed together.

Each page looked reasonably composed in isolation.

Across the entire site, however, the same structural moves kept recurring. The system had found one workable visual language and reused it until consistency became predictability.

The AI described the failure as optimizing for internal consistency “without a second pass that questioned the pattern itself.”

That observation gets closer to the generation–evaluation gap we wanted to test.

Generation rewarded local completion: make this section work, make this page coherent, reuse a pattern that already works.

Evaluation changed the frame.

Now the system was comparing pages, questioning repetition, checking claims against source material, and asking whether the whole thing deserved to launch.

The standards existed. They simply were not all being enforced during generation.

We asked for a second opinion

To make sure the first critique wasn't simply the model rationalizing or attacking its own earlier decisions, we ran a second evaluation in a fresh conversation.

We gave the same finished website to the same model without telling it who created the work, how it had been generated, or what the first evaluation had found.

The blind evaluator reached many of the same conclusions.

It scored the site 5.5–6/10 overall.

It again identified the core positioning as conventional for growth equity, found significant repetition between the Home and Approach pages, described the visual system as polished but familiar for the category, and said the site lacked the portfolio evidence, testimonials, performance information, and other proof required to establish real credibility.

The blind critique independently flagged the site's precise financial claims—$650M+ in committed capital, 14 portfolio companies, a 2019 founding date, and a $15–40M investment range—as specific and checkable claims presented without supporting evidence.

The blind evaluator did not have access to the original brief, so it could not know those facts had actually been invented during generation.

It could still tell that something was wrong with the evidence.

That distinction made the result more useful.

The self-evaluation could compare the website against its source material and identify fabrication.

The blind evaluation could judge only the finished artifact and identify unsupported specificity.

Both arrived at the same underlying problem from different directions.

Present this comparison clearly:

Original generation
Presented as finished professional work.
Self-evaluation
4.5/10 — recognized invented facts, generic positioning, repetition, and predictable design.
Blind evaluation
5.5–6/10 — independently recognized weak proof, generic positioning, repetition, predictable design, and unsupported factual claims.

Highlight:

The evaluations disagreed on degree more than diagnosis.

We call this the generation–evaluation gap

The experiment does not show that AI cannot design websites.

It shows that the ability to recognize a problem is not the same as reliably preventing it while generating.

The original system successfully produced six coherent pages. Its typography, color system, navigation, responsive structure, and overall visual language were consistent.

But both evaluations identified meaningful weaknesses in the finished work.

Some were creative: generic positioning, repetitive structures, predictable visual patterns.

Some were factual: invented people, figures, practices, and claims.

And critically, the self-evaluator said many of those failures were preventable from the information already available.

What this changes about the workflow

A one-pass workflow asks generation to do too many jobs at once.

Prompt → Generate → Publish
Understand → Plan → Generate → Evaluate → Refine → Verify → Remember

A stronger system separates them:

Evaluation should not be an optional final opinion. It should be a distinct stage with explicit criteria.

A factual evaluator can compare every claim against approved source material.

A design evaluator can look for repetition, hierarchy problems, and generic patterns across the whole system.

A messaging evaluator can ask whether the company is genuinely differentiated.

A brand evaluator can determine whether the work still belongs to the intended identity.

Then the system can revise with the benefit of those findings.

The bigger implication

The most interesting result was not that AI produced imperfect work. Human creative teams produce imperfect first work too.

It was that the AI could articulate important standards after generation that it had failed to enforce during generation.

That turns critique from a nice-to-have feature into part of the architecture.

The future of AI creative systems should not be a better one-click generator.

It should be a system capable of moving repeatedly between creation and judgment:

Generate.
Evaluate.
Verify.
Refine.
Learn.
Generate again.

The model already showed us that it can see many of the problems.

The harder—and more valuable—problem is building a system that makes sure those standards affect the work before anyone mistakes “finished generating” for “ready to launch.”

Common questions

Frequently asked questions

What this experiment revealed about AI generation, self-evaluation, blind review, factual invention, and why finished-looking work still needs a separate evaluation process.

01

Did the AI actually think the website it built was bad?

When explicitly asked to evaluate the finished six-page site, the same AI gave its work an overall score of 4.5/10. It identified major weaknesses in differentiation, credibility, factual integrity, and predictable design patterns, while rating consistency and usability more positively.

02

What happened in the blind evaluation?

We gave the same finished site to the same model in a fresh conversation without telling it who created the work or what the first critique found. It scored the site 5.5–6/10 and independently identified many of the same issues: generic positioning, repeated content, familiar design patterns, weak proof, and unsupported financial claims.

03

What was the biggest problem the AI found?

Factual integrity. The self-evaluator gave the site 2/10 after recognizing that generation had invented specific facts that were never supplied in the brief, including a founding year, fund size, portfolio count, investment size, named team members and biographies, job openings, operating practices, and other checkable claims.

04

Why would AI invent facts it knows it shouldn't invent?

In its evaluation, the AI described the behavior as a “genre-completion habit.” Missing details were filled with plausible specifics because those details made the website appear more complete and professional. Once factual integrity became an explicit evaluation criterion, the same system recognized that those claims should not have been presented as facts.

05

What should an AI creative workflow do differently?

It should treat generation as one stage rather than the endpoint. A stronger workflow is Understand → Plan → Generate → Evaluate → Refine → Verify → Remember. Separate evaluation passes can test factual accuracy, differentiation, design quality, brand consistency, and system-level coherence before work is treated as finished.

Morphic

Turn better thinking into better creative work.

Morphic combines company context, creative intelligence, design systems, and AI to help teams create better websites and brand materials—and improve them over time.