No modern product is built without AI now, and building well with it takes a different eye. AI is taking over the middle of the work: the generation, execution, production, volume. This middle is growing fast: a quarter of the code at Google two years ago, a third at Microsoft last year, now 75%+ at Google, and 80%+ at labs like Anthropic, up to 100% for power users at frontier labs. You would expect the human to disappear as the middle swallows the work. The opposite happens. The judgement concentrates at the two ends.
Humans define the AI in product: both how you build with AI, and how you build AI into the product - the line where the machine's freedom stops. This article is about the first, in a two-part series covering building with AI. Underneath both is a plainer aim: keeping what we build, and how we build it, human.
Confidently, invisibly wrong.
AI isn't ready for full automation. AI is extraordinary, and on narrow tasks it has cleared any bar you would set for a person. Set it loose unsupervised though, and the failures are not the ones you brace for. They are confident, well-reasoned, and often invisible.
I saw the everyday version looking after Amazon's shopping experience, working with the world's biggest brands in 2025. Our models could build a brand's entire product environment at a scale no studio could match. They would also, with total confidence, invent a detail no product had: a feature, a finish, a colour the brand never made. That is not cosmetic for shopping. Misaligned product media moves return rates, customer trust, and conversion; inaccuracy was a business risk. The model was never unsure. It was wrong the way a confident colleague is wrong, which is the hardest kind to catch.
The pattern is general. OpenAI's own researchers concluded that models hallucinate because training rewards the confident guess over the honest blank: on the benchmarks that decide which model wins, admitting uncertainty scores zero, and a wrong answer sometimes scores. The model that invented a finish no product had was not malfunctioning. It was doing what it had been graded to do.
A model can write almost all the code in a system. That system will not, on its own, understand the customer's problem, stay safe, notice when it is wrong, or know when to pivot. A working product needs customer understanding to shape the experience, fail-safes, a way to notice it is wrong, and the judgement to pivot: exactly what the confident-and-wrong failure mode removes. The human is not there because the model is weak. The human is there because the model is convincing. That holds while you build the product, and again in what you ship. Take the building first.
Building with AI: the PM's touch, first and last.
Across my work, the interfaces with AI that mattered most were always the first and the last.
The first touch is framing. Before the model generates anything, a person sets the conditions: the vision, the customer outcome, working backwards from what good looks like to the brief that gets there. Structured prompts, a defined tone, the domain rules the model does not know it is breaking, and underneath all of it, the right question. The model cannot do this for you, because it has no stake in the outcome and no view on which outcome is worth having. Give it a vague brief and it compounds the vagueness. A sharp, opinionated first touch makes a difference, needing clarity before the model is asked. Nothing downstream adds it back.
This matters more now, not less, because the cheaper the middle gets, the faster a wrong first touch builds the wrong thing at scale. Meta's metaverse is the monument to it: more than eighty billion dollars of committed execution poured into a vision built backward from a technology bet rather than from anything people wanted. By the time the spending peaked, fewer than one in ten of its user-built worlds had ever been visited by more than fifty people. Alexa found the opposite failure: hundreds of millions of devices and no settled answer to what the assistant was actually for. Unable to convert novel technology into true utility.
The last touch is validation and improvement: reviewing the edge cases, applying judgement, deciding what ships. Two things make it harder than it sounds. It is fractal: the same sign-off recurs at every altitude, and the skill is knowing which one you are on. You validate a prompt, a template, a workflow, a system, an agent. A threshold that is right for a single prompt is reckless for an autonomous agent. And it is not the last thing you do; it is the thing you never stop doing. You ship, you watch, you feed what you learn back in.
Calling it a touch undersells it. Validation at any scale is a system, and most teams build it badly because they treat every failure the same. They are not the same. Deterministic. A broken format, a missing field, a schema that will not parse; code catches these, and you do not need a model to judge them. Subjective. Did the system understand what the user meant, is the output any good. These need a human-defined standard before any automated judge can apply it. Drift. Right today, wrong next month, as the world and the users move underneath you. Reach for the same expensive check on all three and you get a system that looks rigorous and is not.
Tuning Judgement.
Where you set both touches is not one decision but many, scaled to what is at stake. On Amazon's shopping surfaces the same asset carried wildly different tolerances. Product-page imagery had to be what-you-see-is-what-you-get: get it wrong and the gap between the picture and the parcel comes back as a return. So the line sat strict: exact product, high grounding, logos pixel-correct. A category page could show a representative product rather than the exact one, and the line could loosen.
Upper-funnel marketing could loosen further, but only by importance: we tiered placements one to four, and a Tier 4 slot ran almost fully on templates and light QA while a Tier 1 hero stayed a human-orchestrated process end to end. Quality never moved. What moved was where the human stood. The craft was matching the validation to the cost of being wrong, surface by surface.
When and where you check matters as much as how. Before you ship, validation prevents regressions; after, it shows you reality and catches drift; periodically, it asks the question everyone skips: are we still measuring the right thing, and does our definition of good still hold?
Because under all of it sits one act that does not scale and cannot be handed off: a person deciding what good means. Every automated judge, every monitor, every threshold is scaffolding around that judgement. Skip it and you can build an eval suite that passes cleanly while the product quietly gets worse: one that looks like it works, until it doesn't.
The failure mode is a sign-off with no judgement inside it. Three lawyers suing Walmart in 2025 filed a brief citing nine cases, eight of them invented by a model; all three e-signed it, and two admitted they had not checked a single citation. The last touch happened. It was just hollow. It is not one profession's lapse: Deloitte refunded the Australian government after a report shipped with invented citations, and a researcher's running database of AI-fabricated court filings now counts them in the hundreds.
There is a structural reason the sign-off went hollow, and it is not laziness. Review used to be social by accident. Work passed through many hands on its way out the door: read, argued over, changed, corrected. AI collapsed that often inefficient chain. Losing the middle also takes the crowd out of the work. One person now faces confident, endless output alone. Three signatures with two unread is not three failures of character; it is what happens when redundancy that used to come free disappears, and nothing is built in its place. The last touch is that rebuilding. The edges are not just where the judgement concentrated. They are where it got lonely.
The rebuilding starts at the level of one person's rules. Sean Goedecke, an engineer at GitHub, lets AI write his code but not his pull-request descriptions or design decisions: code is verifiable, so a deterministic check catches its errors, while a document carries implicit judgement the model will smooth into something fluent and wrong.
What endures.
Judgement moving to the edges is not a disruption. It is a pattern, and an old one. The arrangement has survived automation before. When the sewing machine took the stitching, the tailor's craft did not vanish; it re-formed: the vision and measuring before, the fitting and finishing after. The machine only ever got the stitches. When the camera took the middle of image-making, a whole new craft grew that was nothing but the two ends: the frame chosen before the shutter, the print and development chosen after it. Painters still exist. Hand-sewn dresses still exist. But the worlds of pictures and fashion have been transformed, in good ways and bad. AI is the largest automation of middles in history, and the stations have not moved.
The through-line is simple. The value was always in the judgement, never the volume or capability. This, it raises it to the ceiling and leaves the judgement where it has always been. Everything you build this way still has to ship, though, and in the shipping the same judgement stops being something you do and becomes something the product is: a line drawn through it that the customer lives inside. That line is the next essay.
The model fills the middle. The judgement moved to the edges.
Next: Drawing the Line. What's actually different about an AI-native product.
Of Middles and Edges
Building with AI: Part 1/2. As AI takes the middle of the work, the judgement does not shrink: it concentrates at the edges.