There are two common explanations for disappointing AI creative work.
One is that the prompt was not good enough.
The other is that the AI did not have enough context.
Both explanations sound plausible. But they describe different problems.
A prompt tells a model what to do. Context gives it information it may need in order to do that work well.
We wanted to see what happened when we changed those variables independently.
So we gave the same AI the same homepage assignment for the same fictional private equity firm four different ways: basic prompt with basic context, strong prompt with basic context, basic prompt with rich context, and strong prompt with rich context.
The result was more useful than a simple winner between prompting and context.
Better prompting and richer context solved different problems. Neither substituted for the other.
The Experiment
We created a fictional private equity firm called Alder Ridge Partners.
The basic context contained only two sentences: Alder Ridge invests in founder-owned and family-owned business services companies in North America, and it invests in established, profitable companies while working with management teams to help them grow.
The rich context added the information a strong creative partner might reasonably know before designing the firm's website: investment criteria, an approximately $5M–$20M EBITDA range, recurring or repeat revenue characteristics, the firm's approach to control and retained founder ownership, six areas where it helps portfolio companies, audience priorities, positioning, brand character, design direction, and explicit factual constraints.
We then generated four homepages in fresh conversations using the same model.
The four conditions were:
K — Basic prompt + Basic context
M — Strong prompt + Basic context
R — Basic prompt + Rich context
T — Strong prompt + Rich context
We did not critique the outputs between generations.
After all four existed, we evaluated them twice. The first evaluation was blind to the generation conditions. The second received the full source material and compared each homepage against what was actually known about Alder Ridge.
Because this was one generation per condition, the experiment is illustrative rather than statistically conclusive. But the differences were large enough to expose several useful failure modes.
Basic Prompt + Basic Context: Specificity Without Truth
The weakest condition was also, in some ways, the most convincing at first glance.
With very little company information and no strong instruction about how to handle missing facts, the AI filled the page with plausible private-equity specificity.
It invented a $10M–$150M revenue range. It invented a $3M–$20M EBITDA range. It invented a 5–8 year hold period, six subsectors, a three-stage deal process, 14 platform investments since 2011, 100% management retention, a $40–$150M enterprise-value range, a Chicago office, and two company email addresses.
None of those facts had been provided.
The blind evaluator still found the page persuasive in several respects. It scored 9/10 for messaging specificity and 9/10 for audience understanding. It felt complete because it contained the kinds of details a real private-equity website might contain.
But its factual-integrity score was 2/10.
That distinction matters.
AI can create the appearance of specificity by filling information gaps with plausible details. More specificity is not necessarily more understanding.
The page looked more complete partly because the model had quietly created the missing company.
Better Prompting Changed What Happened When Information Was Missing
The second condition kept the same sparse company context but replaced the basic request with a much stronger instruction.
The prompt asked the model to make deliberate decisions about objective, audience, positioning, hierarchy, narrative, proof, composition, and section roles. It also explicitly prohibited inventing company facts and instructed the model to design around missing information instead.
The difference was immediate.
The fabricated facts disappeared.
The resulting homepage used only what the model actually knew: sector, ownership type, company stage, geography, and the fact that Alder Ridge works with management teams after investing.
The source-aware evaluator gave it 10/10 for factual integrity.
But another limitation became obvious: the page was thin.
Its messaging-specificity score was only 3/10. It did not contain an EBITDA range, detailed investment characteristics, operating capabilities, ownership nuance, or multiple audience paths because none of those things existed in the context it had received.
The stronger prompt helped the model handle uncertainty responsibly. It could not create company knowledge that was not there.
A better instruction raised the discipline of the work. It did not raise the information ceiling.
Rich Context Did Not Automatically Produce a Rich Result
The third condition reversed the experiment.
We returned to the basic prompt but supplied the complete rich company context.
Now the model had substantially more to work with.
It knew Alder Ridge typically looked for companies with approximately $5M–$20M of EBITDA. It knew about durable customer relationships, recurring or repeat revenue, fragmented markets, control investments with possible retained founder ownership, six specific operating-support areas, and three distinct audiences.
The output remained factually safe. It received 9/10 for factual integrity.
But it used surprisingly little of what it knew.
The EBITDA range disappeared. The recurring-revenue criteria disappeared. The fragmented-market acquisition thesis disappeared. The ownership nuance disappeared. All six operating capabilities disappeared. Executives and limited partners disappeared as audiences.
The source-aware evaluator gave the page only 4/10 for use of available context and 3/10 for avoidance of harmful omission.
This may be the most important result in the experiment.
Giving AI information does not guarantee that the information will meaningfully influence the work.
The context existed. The generation process did not reliably determine which parts deserved to become prominent decisions in the homepage.
That is a different failure from hallucination. The model was not making things up. It was leaving valuable truth unused.
The Strongest Result Combined Both
The final condition paired the strong prompt with the rich context.
This produced the clear winner in both evaluations.
The homepage used the real EBITDA range, recurring-revenue characteristics, fragmented-market acquisition rationale, ownership nuance, all six operating-support areas, and all three audiences. It preserved founders as the primary audience while giving executives and the investment community appropriate secondary roles.
It also maintained 10/10 factual integrity.
The source-aware evaluator scored it 9/10 for positioning, messaging specificity, narrative coherence, information architecture, audience understanding, appropriateness, and overall launch readiness. It scored 10/10 for both factual integrity and use of available context.
This was not simply because the model had more information. The basic-prompt version had access to the same rich context and left much of it unused.
And it was not simply because the prompt was better. The strong-prompt version with basic context was disciplined but could not become equally specific.
The strongest output required both.
Prompt Quality and Context Appear to Govern Different Failure Modes
The four conditions make the interaction easier to see.
When we improved the prompt while keeping context sparse, fabrication disappeared and the structure became more deliberate, but the page remained thin.
When we improved context while keeping the prompt basic, fabrication also disappeared, but much of the valuable information never made it into the page.
When we added rich context to the strong prompt, specificity, positioning, audience understanding, and context utilization all increased sharply while factual integrity stayed perfect.
When we added the strong prompt to the already-rich context, the model began surfacing information it had previously ignored.
So the useful distinction is not prompt versus context as though they were competing techniques.
Context determines what the system has available to understand. Orchestration helps determine what it should do with that understanding.
There is also a third layer: judgment. Neither having the information nor following a sophisticated instruction guarantees that the resulting creative decisions are actually good. The output still needs to be evaluated.
What This Means for AI Creative Systems
A common response to weak AI output is to focus on prompt engineering.
Another is to build larger knowledge bases and retrieve more context.
This experiment suggests that either approach in isolation is incomplete.
A creative system needs access to the right company knowledge. But it also needs a process for deciding which knowledge matters for the assignment, which audiences matter most, what the page needs to accomplish, which claims require proof, what should be prominent, what should be omitted, and how unknown information should be handled.
That changes the workflow from something like:
Prompt → Generate
to something closer to:
Understand → Select relevant context → Plan the work → Generate → Evaluate → Refine
The context layer and the orchestration layer have different jobs.
A strong company knowledge layer without good orchestration can produce work that knows more than it communicates.
Strong orchestration without company knowledge can produce work that is disciplined but generic.
And weak versions of both can produce something particularly dangerous: polished work whose apparent specificity comes from information the model invented.
More Context Is Not the Same as Better Context
There is another implication worth separating from the experiment itself.
The rich context condition was structured and task-relevant. It was not simply a larger pile of documents.
That distinction matters.
Giving a model thousands of tokens does not ensure it has the right information, knows which information is authoritative, or understands which facts matter to the current assignment.
Useful context needs selection, structure, provenance, and relevance.
For creative work, that might include company facts, audience priorities, positioning, proof, brand rules, strategic priorities, previous decisions, approved patterns, rejected directions, and task-specific constraints.
The objective is not maximum context.
It is the right context, available at the right moment, used for the right reason.
The Prompt Is Still Not the Product
None of this makes prompting unimportant. The experiment showed the opposite.
The stronger prompt had a substantial effect on factual discipline, structure, context utilization, audience coverage, and launch readiness.
But the prompt was one component of a larger system.
Someone still had to determine the company context. Someone had to decide which instructions belonged in the workflow. Someone had to evaluate the outputs. Someone had to know which claims were actually supported.
For one creative task, a sophisticated user can manually assemble all of that into a long prompt.
For repeated professional work, rebuilding the entire operating environment every time becomes the problem.
The more durable opportunity is a system that already understands the company, assembles the relevant context for the assignment, applies an appropriate process, evaluates the result, and remembers what should carry forward.
What This Experiment Does and Does Not Prove
This was one fictional company, one homepage task, one model and version, and one generation per condition.
We did not run enough repetitions to claim that the measured score differences represent universal effects of prompt quality or context depth. Different models, tasks, companies, and random generations could produce different rankings.
The experiment also does not establish that one variable matters more in general.
What it does show, in this instance, is a useful interaction pattern.
The basic prompt with basic context produced confident fabrication.
The strong prompt with basic context produced disciplined but thin work.
The basic prompt with rich context produced accurate but underpowered work.
The strong prompt with rich context produced the most complete, specific, grounded, and launch-ready result.
The question is not whether prompts or context matter more. The more useful question is what job each should do.
Context gives AI something real to understand.
Orchestration determines how that understanding should be used.
Judgment determines whether the result deserves to ship.
