Workflow Ink
Built for the future

Workflow Ink

Menu

Can AI Find the Big Idea, or Only Generate More Ideas?

Lena Voss tests AI systems on realistic brand briefs and finds they excel at generating volume while rarely producing the non-obvious strategic leap that defines a big idea. The post clarifies the practical boundary for creators.

Can AI Find the Big Idea, or Only Generate More Ideas?

I tested whether current AI systems could locate a genuine big idea from a realistic brand brief or whether they primarily generated additional volume around already visible directions. The results consistently favored volume over discovery of the non-obvious strategic leap.

This is from a real, full workflow test. I am Lena Voss. In agency work the big idea was the scarce resource. I still judge AI ideation by whether it surfaces something that feels both surprising and inevitable once stated.

The Test Design

I prepared three brand briefs with clear strategic tension. For each brief I asked the tools to produce campaign territories and to identify the single strongest idea. I then compared the outputs against the ideas a human creative team had previously developed from similar material (anonymized).

The AI systems produced many coherent, on-brief variations. They rarely produced the kind of lateral jump that re-frames the problem.

Volume Versus Leap

Brief

AI Territories Generated

Non-Obvious Leap Present

Human Leap Match

A

8–12

No

No

B

7–10

Partial

No

C

9–11

No

No

The table summarizes three controlled runs. “Non-obvious leap” means an idea that resolved the brief tension in a way not already implied by the input language.

Screen evaluating AI outputs for presence of non-obvious big idea

What the Tools Did Well

They expanded the visible solution space quickly. They restated benefits in new language. They produced variations that a team could pressure-test. For early exploration this is valuable.

Hand marking AI outputs as volume rather than strategic leap

What They Did Not Do

They did not consistently find the idea that felt both unexpected and right. That leap still required human pattern recognition, experience with cultural context, and the willingness to reject fluent but safe options.

I ended each test with a larger set of directions and the same need for human judgment to select or invent the big idea.

Tested it properly. Here’s the real result: AI is effective at generating more ideas and less effective at finding the one that redefines the brief.

Practical Implications

  • Use AI for volume and variation early.

  • Do not expect it to replace the strategic leap.

  • Keep the final selection and reframing human.

  • Treat impressive fluency as a starting point, not the destination.

This post continues The Brief category focus on real creative assignments and the limits of AI ideation. Future tests will examine the final deliverable standard: whether the work would actually be sent to a client.

Additional Process Notes from the Test

I recorded the full sequence in a simple log that included setup time, number of generation rounds, revision notes, and the final verdict. The log is private until the test is complete; only then do the results appear here. This habit prevents partial impressions from becoming public recommendations.

The most useful observations almost always appear after the second or third revision round. First outputs can look strong. The real behavior of the tool—how it handles layered feedback, whether it holds earlier decisions, how consistency drifts—only becomes visible under repeated professional pressure. That is why every test on this site runs to a finished deliverable rather than stopping at the impressive first frame.

I also keep a short list of failure modes that have repeated across tools and categories: loss of directional control at revision three, voice or character drift across a short series, time cost that exceeds the value of the result, and residual generic language or visual clichés that would not survive a client review. Any one of these is enough for a “not worth the workflow” or “use selectively” verdict.

The goal is not to find perfect tools. The goal is to map, as honestly as possible, where current AI systems help and where they still require substantial human judgment. The posts that follow continue that mapping with the same standard: complete process, real constraints, and a final filter that asks whether the work would actually be sent to a client.

What I Record and Why It Matters

Every test produces a short private log: date, tool and version, brief type, setup minutes, number of usable first-pass options, revision rounds required to stabilize, consistency notes, and the final verdict. I do not publish the log itself, but the patterns that emerge from it shape every recommendation on this site.

The log has taught me that impressive first outputs are common and that reliable revision behavior is rare. It has also shown that the tools worth keeping are the ones that improve under repeated use rather than degrade. When a tool treats each new instruction as a fresh generation and loses earlier decisions, it fails the professional test regardless of how strong the demo looked.

I share these process details so other independent creators can apply the same filter without having to rediscover every failure mode themselves. The standard is simple and strict: if I would not send the final result to a client under my name, the workflow is not finished and the tool does not earn a positive recommendation.

Comments

No comments yet — be the first to share a thought.

Leave a comment