I took one street photograph in Silver Lake and asked five different AI image tools to produce five distinct art directions from it. The first results looked promising. The problem nobody mentions showed up later: consistency of intent, control over revision, and whether the final set could survive a professional review.
This is from a real, full workflow test. I am Lena Voss. I treat AI image generation the same way I treated art direction in agency life: the brief stays fixed, the directions must be distinct, and the final images must still feel intentional.
The Starting Point
The source image was a straightforward street photograph: late afternoon light, a single figure crossing a quiet intersection, Los Angeles residential architecture in the background. I wrote a short art-direction brief that required five different visual territories while preserving the original emotional register and practical usability for a campaign-style set.
I did not ask for “make it better.” I asked for five controlled directions with clear constraints on palette, mood, and composition.
The Five Directions I Requested
Clean editorial with restrained color grade
High-contrast graphic treatment
Soft documentary with film grain
Slightly elevated commercial product-style framing
Minimal geometric crop emphasizing architecture
Each direction had its own short prompt block derived from the same brief.

What the First Round Produced
Four of the five tools returned images that looked polished on first glance. One direction stayed close to the source. The others drifted in lighting, subject scale, or emotional tone. The real problem appeared when I tried to revise any single direction toward the brief without losing the others.
Revision control was uneven. Some tools accepted detailed feedback and moved closer. Others treated every revision as a new generation and lost the previous intent. That is the problem nobody mentions in most demos: the ability to hold a direction across multiple edits.
Revision Behavior Observed
Tool Behavior | Held Direction | Lost Intent | Required Full Restart |
|---|---|---|---|
Accepts layered feedback | Yes | Rarely | No |
Treats each edit as new | Sometimes | Often | Yes |
Strong first pass, weak edit | Yes initially | After round 2 | Often |
The table summarizes patterns I recorded across the five tools in this single test.

The Consistency Gap
Even when individual images looked strong, the set as a whole rarely held a coherent art direction. Character scale, light direction, and color temperature shifted enough that the five images no longer felt like one campaign. For professional use that gap matters more than any single beautiful frame.
I ended the test with two directions I would consider usable after further human work and three I would not send forward. The time cost of reaching that point was higher than the first impressive outputs suggested.
Ran the full process myself. This is what it looked like when the brief stayed fixed and the tools were forced to work inside it.
Practical Takeaways for Creators
Lock the art-direction brief before generating.
Test revision behavior early, not after you fall in love with a first image.
Evaluate the set, not the single frame.
Record time and control loss so the real cost is visible.
This Image Lab test sits inside the larger practice of treating AI as a controlled collaborator rather than a magic generator. Future posts will examine character consistency across longer sequences and the specific points where human art direction remains essential.
Additional Process Notes from the Test
I recorded the full sequence in a simple log that included setup time, number of generation rounds, revision notes, and the final verdict. The log is private until the test is complete; only then do the results appear here. This habit prevents partial impressions from becoming public recommendations.
The most useful observations almost always appear after the second or third revision round. First outputs can look strong. The real behavior of the tool—how it handles layered feedback, whether it holds earlier decisions, how consistency drifts—only becomes visible under repeated professional pressure. That is why every test on this site runs to a finished deliverable rather than stopping at the impressive first frame.
I also keep a short list of failure modes that have repeated across tools and categories: loss of directional control at revision three, voice or character drift across a short series, time cost that exceeds the value of the result, and residual generic language or visual clichés that would not survive a client review. Any one of these is enough for a “not worth the workflow” or “use selectively” verdict.
The goal is not to find perfect tools. The goal is to map, as honestly as possible, where current AI systems help and where they still require substantial human judgment. The posts that follow continue that mapping with the same standard: complete process, real constraints, and a final filter that asks whether the work would actually be sent to a client.
Comments
No comments yet — be the first to share a thought.