I fed a complete, realistic brand brief into three different AI systems and asked for campaign territories. The first round of ideas was exactly as generic as I expected. The useful work only started after I forced the tools to respond to the specific constraints in the brief.
This is from a real, full workflow test. I am Lena Voss. I spent years writing and judging briefs inside an independent Los Angeles agency. I still use the same standard: if the idea could belong to any brand, it is not yet an idea.
The Brief I Used
The brief described a mid-sized consumer brand entering a new category. It included audience definition, competitive context, tone boundaries, mandatory visual references, and a clear success measure: three distinct territories that could each support a six-month content system.
I removed all client-identifying details and kept the structural pressure intact. The brief was long enough to constrain the model and short enough to stay usable.
What Generic Output Looks Like
The first ideas across all three tools shared the same problems:
Broad emotional claims without specificity
Visual language that could apply to multiple competitors
Tagline patterns that felt familiar rather than earned
Little evidence that the model had absorbed the competitive context
None of the first-pass territories would have survived an internal creative review.

How I Forced Better Direction
I did not abandon the tools. I revised the input. I extracted the three most distinctive constraints from the brief and restated them as non-negotiable filters. Then I asked for new territories that had to pass those filters before any language was generated.
The second round improved. Two of the three tools produced at least one territory that felt specific enough to develop. The third tool continued to drift toward generic positivity. That difference became part of the final verdict.
Filter Results After Constraint Reinforcement
Tool | First-Pass Specificity | After Constraint Filters | Final Usable Territories |
|---|---|---|---|
Tool A | Low | Medium | 1 |
Tool B | Low | Medium-High | 2 |
Tool C | Low | Low | 0 |
The numbers reflect my scoring against the original brief criteria, not subjective preference.

What This Test Confirmed
AI can expand the volume of early ideas. It cannot replace the judgment that decides whether an idea is actually on strategy. The brief remains the primary creative instrument. When the brief is treated as optional context, the output stays predictable and interchangeable.
I ended the test with two territories I would take into further human development and a clear record of how much additional direction the models required. That record is more useful to other creators than any claim about speed or cleverness.
Tested it properly. Here’s the real result: the first ideas were predictably generic, and the useful work began only after the brief constraints were enforced.
Practical Notes for Creators
Write the brief before you open the tool.
Extract the non-negotiable constraints and restate them explicitly.
Score early ideas against the brief, not against how impressive they sound.
Expect to spend more time on direction than on generation.
This post sits in The Brief category because the brief itself was the decisive variable. Future tests will examine how messy client notes can be turned into usable territories and whether AI can locate a genuine big idea or only generate more volume.
Additional Process Notes from the Test
I recorded the full sequence in a simple log that included setup time, number of generation rounds, revision notes, and the final verdict. The log is private until the test is complete; only then do the results appear here. This habit prevents partial impressions from becoming public recommendations.
The most useful observations almost always appear after the second or third revision round. First outputs can look strong. The real behavior of the tool—how it handles layered feedback, whether it holds earlier decisions, how consistency drifts—only becomes visible under repeated professional pressure. That is why every test on this site runs to a finished deliverable rather than stopping at the impressive first frame.
I also keep a short list of failure modes that have repeated across tools and categories: loss of directional control at revision three, voice or character drift across a short series, time cost that exceeds the value of the result, and residual generic language or visual clichés that would not survive a client review. Any one of these is enough for a “not worth the workflow” or “use selectively” verdict.
The goal is not to find perfect tools. The goal is to map, as honestly as possible, where current AI systems help and where they still require substantial human judgment. The posts that follow continue that mapping with the same standard: complete process, real constraints, and a final filter that asks whether the work would actually be sent to a client.
Comments
No comments yet — be the first to share a thought.