When I led Home Innovation at Amazon, I once played a cooking game with my PM team, stolen from Chopped: standard burger ingredients for everyone, plus one compulsory wildcard: cornflakes. The category leaders for Food and Home judged the entries blind; then we sat down and ate the failures with the successes. It left two things in the work. The constraint forced sideways thinking: the odd ingredient is where the interesting answers hide. And the shared meal and memories made failure recoverable and built a shared lexicon. A team that eats its own bad answers together stops being afraid of producing them. Most of it was, in fairness, delicious.

Innovation lore celebrates play: hackathons, skunkworks, twenty-percent time; then treats it as a perk. A break from real work, scheduled once or twice a year, justified as team bonding, or hurried along to meet annual planning timelines. That framing gets play wrong in a way that makes it easy to de-prioritise, and the test for getting it right is simple. At work, play is not judged by fun. It is judged by residue: what it leaves in the work after the laughing stops. Practice-play leaves capability: shared language, calibrated judgement, trust you can draw on later. The cornflake burgers passed.

The right kind of play

Next quarter, I mailed every PM a blank canvas and the same paint set, with a one-line brief: paint your Day One. Our lead UX, a painter himself, judged blind. Identical materials; no two paintings alike; none analytical. The residue was a reminder the team needed: product sense is not only a spreadsheet skill, and how important it is to create.

Another time, I handed each PM a random Lego set: a Harry Potter scene, a Sesame Street character. I then asked them to build their dream home from it. Entries went in anonymously; the Director of Home EU judged. What came back was not a row of toy houses. It was a gallery of metaphors for what home meant, and this residue is easy to name: months later, the team was still using those metaphors to argue about the roadmap.

None of these were perks. They were work, woven into the operating weeks. Across them, my recipe became four things: low stakes, blind judging, strange constraint and anchoring in the human. Low stakes let a bold idea surface without attaching it to a roadmap. Blind judging mutes status: when entries are anonymous, the quietest PM can beat the loudest leader, and everyone watches it happen. The strange constraint makes the obvious answer unavailable. And every brief reached into the person: your Day One, your idea of home, your culinary taste. Material anchored in someone's own life is material no one else could have brought, and its owner will defend it.

Residue has to survive two tests of its own: persistence and transfer. Does it outlast the session, and does it appear in the real work? The painting afternoon could have been memorable and still failed, if nobody's judgement changed. The Lego passed because the vocabulary did.

Two neighbouring frames do a different job. Amy Edmondson's psychological safety describes the climate; the residue test examines the consequence. A team can feel safe and still learn nothing. Michael Schrage's Serious Play showed that a prototype's value lies in the interactions it provokes; the residue test asks what remains after the prototype disappears.

Older than management

Johan Huizinga's claim in Homo Ludens was that play does not decorate culture; it precedes it. Ritual, law, poetry and sport begin as forms of play, and animals played long before anything human existed to name it. Many serious capabilities are rehearsed somewhere failure is cheaper: the cub play-fights before it hunts. Workplace play matters for the same reason. It gives judgement somewhere cheap to learn before a real decision makes error expensive.

The enemy of all this is not seriousness. It is play-as-perk: the annual permission slip, prototypes that never ship, applause that never compounds. Rigour gets a team to the answer faster. Play widens what counts as an answer. A company that schedules the second as a holiday has decided, without noticing, to find only what it already expected.

The strongest objection

The case against everything above got harder this year. A field experiment at Procter & Gamble put 791 professionals on real product problems and found that one person working with AI matched the quality of a two-person team working without it. The model even bridged functional silos: technical people proposed commercial ideas, commercial people technical ones, and participants reported enjoying the work more. So the serious objection is no longer that idea generation has become cheap. It is that AI may reproduce much of what we thought the team itself was for. Add that in a year of displacement anxiety mandated whimsy can read as cruelty, and the perk-cutters have their brief.

But look at what the experiment measured: the quality of output from teams assembled for a day. The authors themselves flag it: these were flash teams, and what the humans could do together the following week was never on the scoreboard. That is precisely the residue test's territory.

What those afternoons actually produced was not ideas but a team calibrating taste together: judging blind, arguing about what good meant, failing in public at low cost and recovering by lunch. The anonymous vote was an eval, run on humans, for humans. The session was training the team's evaluation function. A model can widen the option set for almost nothing. It cannot give a team a shared standard for choosing among the options: knowing not just what I think is good, but what we mean by good and where we honestly disagree. That standard is built together or not at all, and it is exactly the muscle the AI era strains hardest. The P&G study itself found the seam: AI lifted the quality of the ideas, but the teams working without it were better at picking their own best one.

An output can be excellent and leave the people who produced it unchanged. The residue test asks whether the people changed too. As for the whimsy: what reads as cruelty is mandated fun with your job on the line; practice-play is anonymous, costs nothing to fail, and is judged on nothing that follows you.

The AI-era versions design themselves once the purpose is clear. Give everyone the same brief and judge the outputs blind: not to find the best prompt, but to surface how differently the team defines good, and argue it out at low stakes. Stage the model's most plausible wrong answers and compete to catch them, an hour of theatre that trains the exact eye the confident failure defeats. The point is never just the artifact. The point is the calibration.

Make it a mechanism

Play survives calendar pressure only if it stops depending on enthusiasm. Put it in the operating cadence with an owner, the way any real practice lives. Keep a residue log, not a score: one sentence per session on what changed in the real work. If nothing transferred, it was a perk.

Make play the habit, not the holiday. It earns its place the way everything else at work does: by what it leaves behind.


Next: A Hundred Versions of One. Diversity of thought used to be free: people couldn't help thinking differently. Now it has to be designed.

What Play Leaves Behind

The value of workplace play is what survives it: judgement, language and trust. Models generate options, yet teams still choose; play is where choosing together gets trained.