#Tools·4 min read

Two questions about any idea. You can only answer one of them yourself.

One judgment used to cover two questions. It no longer does.

You ask AI for ideas and one of them is good. Better than you would have reached alone, and you can tell — it clears the bar your own drafts usually clear. That judgment is the problem. It is the only one you are in a position to make.

Six researchers put this to a proper test over the summer, and the result comes in two halves. On quality the tool wins and it is not close: its ideas were rated seven times more likely than human ones to land in the best tenth. On distinctiveness it loses by about the same margin. Measure how far apart the ideas within one batch sit, and human batches were roughly twice as spread out.

The narrow-task objection is worth taking, and then worth looking at twice

The obvious response is that this was one narrow task — consumer product ideas under a fixed price, judged on whether people would buy them. Worth taking seriously. And then worth looking at twice, because a narrow brief with a clear measure of success is not a rare case at work. It is most of what gets asked for.

The paper is careful about its own scope. It does not claim these figures describe strategy, taste or writing, and neither does this post. What travels is the mechanism.

The split you have to keep in mind now

You have two questions about any idea, and you can only answer one of them yourself. Is this good? — you can answer that. You compare it against your own past work, which is the standard you carry with you. Is this different? — you cannot. That is a comparison against everybody else's answer to the same question, and you have never seen those.

Until recently you did not need to. Good ideas and unusual ones came out of the same process, so the first was fair evidence of the second: if something felt sharper than your usual, it was probably also not what the person down the hall had written. One judgment covered both.

The tool raised quality by the same thing that makes it converge

The tool raises quality by drawing on what has worked before. Drawing on what has worked before is also what makes its answers converge — so the better it gets, the more reliably the same brief hands the same good idea to whoever asks for it. The feeling of having landed on something has not changed at all. What that feeling predicts has.

None of which is a reason to use it less. It is one habit that has stopped working. The moment an AI idea felt good used to be the moment to stop looking. It is now the moment to ask what everyone else putting this question to the same tool is being handed — because there is a fair chance you are holding it.

The step where a team goes past the first good answer is a discipline you can build, not a reflex you can lean on: Build the step where you go past the first good answer →

Keep reading: The most common use is not the best one · The AI deskilling risk

This is the thinking behind a Trusted AI Culture.


Source: Christian Terwiesch, Lennart Meincke, Karan Girotra, Ethan Mollick, Gideon Nave and Karl Ulrich, "Artificial intelligence and its impact on creativity and diversity: An empirical study of large language model-generated product ideas," Production and Operations Management, July 2026. Figures describe a consumer product ideation task and should not be read as effect sizes for other kinds of work.

Written with AI in the loop: my idea, AI drafted and sharpened, my judgment on the way out. Every word is mine to stand behind. — Darren

← Back to Insights
Work with AIFueledCulture

Not sure where your team actually stands with AI?

The AI Blind Spot Assessment is free and takes under 3 minutes — 18 questions that reveal exactly where your team is gaining or losing ground in how they work with AI.

Take the assessment →
Newsletter

Stay current on AI in the workplace.

Practical insights on AI adoption, culture change, and what's actually working in organizations.

AIFueledCulture
© 2026 AIFueledCulture