Brainstorming was invented by advertising, and its central claim was false from the start. Alex Osborn, the O in BBDO, popularised the method in his 1948 bestseller and sold it on one promise: a group produces more ideas than the same people working alone. Yale researchers tested the promise in 1958 and found the opposite. Individuals working separately, their ideas pooled afterwards, beat the interacting group. The result replicated for decades. Diehl and Stroebe traced the main cause in 1987, and it is almost embarrassing: only one person can talk at a time, so everyone else holds their ideas until the ideas die waiting. A 1991 meta-analysis added the insult: the bigger the group, the wider the gap.
So, the brainstorm, judged as an idea machine, has been a documented failure for nearly seventy years. And yet we kept booking the room. Seventy years of contrary evidence, and the sticky notes never stopped. That is usually the sign of one of two things: a profession too incurious to read its own literature, or a practice doing a job the idea count never measured.
Ten years of running product teams taught me it was the second.
What the room was actually for
The economist Friedrich Hayek wrote in 1945 that the knowledge a society needs never sits in one place. It exists only as dispersed fragments: incomplete, and frequently contradictory, scattered across individuals. He was arguing about markets, but he described every product decision I have ever sat in. The engineer knows why the last migration failed. The support lead knows what customers actually say when they cancel. The new hire knows what the industry does elsewhere. Nobody knows all of it, and the fragments can disagree.
That is the job of the room. Not to manufacture ideas: the literature killed that claim in 1958. But to surface knowledge that exists nowhere else, force the contradictory fragments into contact, and leave the people who must build the thing committed to it, because they were there when it took shape. When researchers embedded at IDEO in the nineties to study its famous brainstorms, they reached the same verdict long before I did. The sessions earned their keep in other currencies: memory, and status settled by ideas rather than titles. The third currency, the one that mattered most in my rooms, their study didn't need to name: ownership of the problem. The idea count was the receipt, not the product.
Shopping for your home
I watched this work for years at Amazon before I could name it. In April 2021 I ran an ideation session for Amazon's Home business, over Chime, mid-pandemic, focused on the search and discovery stage of the shopping funnel: our team's global ownership area. After context-setting, everyone wrote silently before anyone spoke. Ideas were pooled without names. Votes were plus signs; the loudest voice counted exactly once. My speaker notes from that morning say why: it lets everyone's ideas shine, not just the loud and the confident.
The wall filled fast, and with good ideas: bigger images, 3D models in the results, filters that felt like play, a quick product overview. Then one unsigned note caught everyone's attention. It quoted our own research back at us: customers don't have the vocabulary to describe what they want. They type modern sofa and mean a feeling; two-thirds of them are squinting at a phone. Every idea on the wall was a better answer. The note said our customers couldn't ask the question well. I still don't know who wrote it: the session was built so I wouldn't, yet it quietly reframed our roadmap: we stopped polishing answers and started building ways to show customers what they had no words to request.
The inversion
The classic brainstorm lost ideas to interaction: blocked, self-censored, dead while waiting their turn. AI deletes that loss and installs a new one.
Two peer-reviewed findings describe the trade. AI-assisted short stories were rated more creative one by one, and more similar as a pool. Then Wharton researchers put a number on it. Asked to invent a toy from a brick and a fan, people working without AI produced ideas that were entirely unique. Among people using ChatGPT, only 6% were: nine participants independently proposed the same toy, down to the name: Build-a-Breeze Castle.
Read that again. Separate people, separate sessions, the same castle.
This is not a bug to be patched in the next release. It is what a model-led room produces. A language model is a compression of everything everyone has already written: the most articulate average in history. Ask it first and you receive the centre of the distribution, beautifully phrased. Your competitor, asking the same question, receives the same centre. The old brainstorm produced too few ideas. The AI brainstorm produces a hundred versions of one. AI makes ideas abundant and independence scarce.
Correlated error
There is a statistical name for what just happened to the room: correlated error. A room earns its keep because ten people arrive carrying independent observations and independent mistakes. Give them all the same collaborator before they speak and the room still contains ten people: but it behaves like one predictor with ten keyboards. The studies measure the sameness of what rooms produce; statistics tells you what sameness implies about what rooms will miss. And the cost is not blandness. Correlated rooms do not merely produce similar ideas; they make their mistakes in unison, and unanimous mistakes are the ones nobody catches.
Which means AI did not break brainstorming. It ran the experiment that exposes what brainstorming was for. Generation is now free; the machine does it better than the room ever did. What the machine cannot generate is the thing Hayek named: the fragment only you hold. The failed migration, the cancelled customer, the contradiction between your model of the problem and your colleague's. That knowledge enters the pool at exactly one moment: before you read the model's answer. Anchoring research is blunt on this point: once a fluent suggestion is on the table, people converge on it, and first ideas pull hardest. After the model speaks, your first thought is no longer fully yours.

Bring yourself
So, the rule: bring yourself before the model speaks. Your first thought, every session, every participant, in writing, vaulted before anyone opens a chat window. And vault observations before ideas: what you have seen, not what you would build. An idea can be regenerated; an observation can only be witnessed. The customer call you sat in. The exception you have seen that the data denies. The product that delighted you last month, the article you cannot stop citing, and the reason either one interests you. The datapoint from another industry, the anecdote you keep retelling, the principle you quietly run your decisions by.
The vault is not a shrine. A first thought can be conventional, wrong, anchored by yesterday's meeting. It is captured not to be admired but to be preserved: a sample taken before the samples are mixed. Expect most vaulted thoughts to be destroyed, combined, or superseded. The room needs their information, not their survival.
Ideas the model can pre-empt: increasingly it already has, because half the room consulted it before arriving. But it was not there when you were delighted. Write-first is not new: nominal group technique and brainwriting have enforced it since the seventies, as the patch for loudmouths and waiting. Everyone knows to discount the vice president's pet idea, and the vice president eventually leaves the room. There is no discount for fluency. Dissent against a boss feels like courage; dissent against the model feels like being wrong. The loudest voice in the room used to be a person you could doubt. Now it is the most articulate average of the entire internet. And it speaks first unless you stop it.
Do we really need to do this?
There is a serious case against keeping the humans in the loop at all, and it recently cleared peer review. Looking further at Procter & Gamble’s field experiment which put 791 professionals to work on real product problems and found that one person with AI matched the output quality of a two-person team without it. The AI even dissolved silos: R&D people proposed commercial ideas, commercial people technical ones. Participants felt better, too. If AI can replace not just the ideas but much of what we thought the team was for, the first-thought ceremony looks like nostalgia.
A deeper look shows the story is more complex. These were one-day flash teams solving assigned problems: commitment, the currency my rooms ran on, was never on the scoreboard. No one was scored on living with the answer, shipping it, or defending it in the third quarter when it wobbled.
The study's own decomposition cuts deeper. Participants generated five ideas, then chose one to develop. AI raised the quality of what they generated but lowered their ability to pick their own best: teams without AI selected their strongest idea about half the time; participants with AI, about 37%. The machine improved the searching and quietly dulled the choosing, even though nobody asked it to choose: exposure was enough. The authors' suggested explanation should sound familiar by now: human teammates introduce friction and dissent; the model tends to affirm. In the paper's own words, AI is a quality amplifier, not a decision enhancer. And note the study's other finding: AI preserved each person's range of quality. That is variance within one mind; the toy study's collapse was sameness across minds. AI can widen everyone's spread while handing everyone the same spread.
The model also bridged only the silos it was given. Put that beside the toy study and the shape emerges: AI widens the individual and narrows the collective. It cannot surface the fragment nobody typed: the support lead's sense of why customers really leave is what the great thinker Polanyi meant when he said we know more than we can tell. It reaches the pool through a person or not at all.
The honest position is narrower than either camp wants: AI substitutes for generation and some connective tissue, not for dispersed knowledge, independent judgment, or the commitment of the people who must build the thing.
Running it now
My working recipe, offered as operator testimony rather than settled science:
First thoughts, vaulted: observations before ideas. Everyone writes before the model speaks: what they have seen, used, loved, and learned. A customer verbatim, a product that delighted them, a principle they trust, a number from somewhere the room doesn't look. Bring links, images, references and reports to be scanned, synthesized and shared. This is the best-evidenced move on the list: sequencing effects and the diversity-collapse studies both point the same way. And the vault applies to reactions too: before the model comments on the pool, let people record their first response to it: the note they would back, the insight they can extend, the customer problem they recognise from their own desk.
The model argues, asks, and attacks: after the vault. Use it to build connections, find threads, validate and deepen research. Ask it to challenge, as red team, pre-mortem voice, and question-asker rather than oracle. A 2026 study and its preregistered replication found that a model that asks questions preserves both the diversity of the pool and people's sense of owning their ideas; I'd call that promising, not proven, and it matches everything I have seen in practice.
Judge blind first. Strip names: human and machine alike, before anything is ranked. Anonymity's effect on candour is old, with strong evidence; hiding whether an idea came from a person or a model also removes a bias in both directions. Reveal provenance after the ranking, when feasibility and accountability need it.
Watch the pool, not the ideas. The new failure mode is sameness, and sameness is measurable. Count distinct frames, not cards: twenty suggestions clustered around three assumptions are three directions, not twenty. If every idea sounds like a cousin of the first, the model spoke too early, or people are not trying. Diverge again before you converge.
This does not slow the machine down, it sequences and guardrails it. Vault what you know before it generates; vault what you think before it judges. The model gets the middle. It still produces its hundred directions: after the room has produced the six that exist nowhere else on earth.
The model will give you the eloquent average of every room it has ever read; your first thought and instinct are the only contribution the internet has not already made. Protect them, and your judgment at the far end. The model gets the middle and always speaks second.
Postscript. I built this room in two days. First Thought is a small working prototype of the recipe above: brief, sealed vault, model second, blind judging. Its one non-negotiable is enforced in product rather than promise: the model cannot speak until every first thought is committed. Try a session, and tell me where it breaks in the comments: first-thought.lovable.app
A Hundred Versions of One
How to preserve collective intelligence, and build better brainstorms and ideas in an AI world