Where it helps, where it makes you worse, and why knowing the difference is now most of the job
Most takes on AI and product management land in one of two camps. Either the tools are about to delete the role, or they're a slightly better autocomplete for Jira tickets. Both are lazy, and both come from the same mistake. They treat "product management" as one big undifferentiated pile of work.
It isn't. In any given week I'm doing four different jobs, and AI performs very differently on each one. Once I started sorting my work that way instead of by calendar, the question changed. It stopped being "can AI do this?" and became "is this the kind of thing AI is good at, and what does it cost me if it's confidently wrong?"
Nobody asks that second question. It's the one that separates PMs who got faster from PMs who just got sloppier.
Everything I ship falls into one of four buckets. The split matters because each bucket fails in a different way.
What the product becomes if we're right, and why anyone should care. It's one artefact. It's also the smallest, slowest and most important thing I produce.
How we win a specific market against specific alternatives. Where we play, what we refuse to build, and what has to be true for the bet to pay off.
The bulk of the craft. Customer insights, data insights, roadmaps, specs, prototypes. This is where the vision meets a real screen and a real user having a bad day.
Tasks, decisions, OKRs. The boring machinery that turns intent into shipped software. It quietly eats more than half of a PM's week.
Look at what happens when you lay the role out this way. The two buckets that decide whether the product deserves to exist are the two smallest. Most of a PM's week is spent downstream of the thinking that actually matters. That imbalance is the real opportunity, and it's the one AI is genuinely good at fixing.
Here's my honest scorecard after running most of my product work through these tools for about two years. Filled boxes are where I now go to AI first. The rest I still do by hand, on purpose.
Filled = AI writes the first draft. Outlined = I do it, and I don't hand it over.
The pattern isn't subtle. AI is great at squeezing a thousand messy inputs into something structured, and great at turning one clear intent into a hundred concrete pieces. What it can't do is have conviction. It can't pick which of several reasonable directions is worth two years of your team's life.
This was the first place AI changed what I actually do, not just what tools I open.
Every product team sits on a pile of qualitative data it never reads. Support tickets. NPS comments. Sales call recordings. Twelve user interviews someone ran last quarter that live in a folder nobody opens. It goes unread not because people are lazy, but because reading four hundred verbatims properly takes three days, and no roadmap review waits three days.
So PMs do what humans do under time pressure. They read fifteen responses, find the two that agree with what they already believed, and put those in the deck. I've done this. Everyone reading this has done this.
AI takes away the excuse. I now feed in the whole set and ask it to group complaints by the underlying job the user was trying to do, not by keyword. Then I ask for the themes that contradict my current roadmap. That last instruction is the whole trick. Left alone, these models will happily agree with whatever assumption is baked into your question.
Before
Quarterly insight review meant one afternoon, fifteen responses, my own bias, and a slide.
After
The full set gets grouped in twenty minutes. Then I spend two hours arguing with the output and pulling the raw quotes behind every theme I plan to act on. Same time spent. It now goes into judgement instead of transcription.
That last line is the point of this whole essay. AI didn't save me time. It moved my time from the low judgement half of the task to the high judgement half. Anyone selling you pure hours saved is measuring the wrong thing.
For most of my career, the gap between a product question and its answer was measured in analyst availability. You'd want to know whether users who finished the onboarding checklist stuck around longer at day thirty. The answer arrived Thursday, by which point you'd already decided on instinct.
The interesting part isn't that a model can write SQL. It's what happens to the number of questions you're willing to ask. When a query costs four minutes instead of four days, you stop rationing your own curiosity. You check the hunch you'd have skipped before because it wasn't worth an analyst's afternoon.
Three rules stop this from turning into confident nonsense:
In a regulated space, and I work in one, this matters more rather than less. A made up number in a clinical context isn't an embarrassing slide. It's a compliance problem. The tools speed up the lookup. They don't take the number off your name.
Quick definition, in case you sit outside product. A spec is the document that tells engineering and design exactly what to build. Some teams call it a PRD. It covers the problem, the user, the behaviour screen by screen, every edge case, what's deliberately out of scope, and how we'll know it worked. If the roadmap says "we're building a review queue this quarter", the spec is the part that says what that actually means.
The usual framing puts specs on the human side of the line, next to roadmaps and decisions. I think that's wrong. It's the one place where my map differs from most frameworks I've seen.
Here's the reasoning. A spec isn't a creative document. It's a completeness document. Its job is to leave no gaps for the engineer building it at 11pm on a Thursday. And specs written by humans almost never fail for being unimaginative. They fail because someone forgot the empty state, or didn't say what happens on a partial network failure, or never defined what the screen does when the user has zero records and also no permissions.
Listing every edge case is exactly what a machine does better than a tired human. So my spec workflow is flipped. I write the intent, the user, the job they're trying to do, and what success looks like. Then I hand it over with one instruction: act as a skeptical staff engineer who has to build this, and list every state, boundary and failure I haven't defined.
What comes back is usually thirty questions. Six are noise. Twenty four are real gaps I'd have found in code review three weeks later, at roughly forty times the cost.
The human writes the intent. The machine stress tests the completeness. Do it the other way round, asking AI to invent what the feature should be and then polishing the prose, and you end up with a spec that reads beautifully and describes a product nobody needed.
There's a second reason I moved specs across the line, and it has nothing to do with saving my own time.
Most engineers I work with now write code with AI in the loop. Not autocomplete. Whole features, from a description. And a coding model is only as good as the description it's handed. Give it a vague ask and it will invent the missing details, confidently, in a way that looks finished and passes review until someone hits the empty state in production.
That's the same failure mode as a human engineer working from a thin spec, except it happens in ten minutes instead of ten days, and at ten times the volume. Ambiguity used to cost you a Slack message. Now it costs you a pull request that has to be thrown away.
So the spec has quietly become the most valuable document on the team. It's the input that decides what the code turns out to be. A spec that names every state, every boundary and every failure isn't bureaucracy any more. It's the brief the machine is actually building from.
Which makes the pairing work in both directions. AI is good at listing edge cases exhaustively, and AI coding tools need edge cases listed exhaustively. One tool's best output is the other tool's required input. The PM's job in the middle is the part neither of them can do: deciding what should be built and for whom, then checking that what came back is the thing we wanted.
The biggest change to how I work is that I don't defend ideas in prose any more.
Written proposals invite vague objections. Someone reads "a unified review queue", builds their own mental version of it, and then argues with that version. You spend the meeting reconciling two imaginary products.
A clickable prototype ends that in ninety seconds. People stop debating the concept and start reacting to the thing in front of them. This column is in the wrong place. I'd never click that. Where's the bulk action. Those are useful objections. They're specific, they're testable, and they show up before an engineer has written a line of code.
Being able to build a working prototype in an afternoon instead of borrowing a designer for a week also changed which ideas reach the table at all. Showing an idea got cheap, so the bar for suggesting one dropped with it. More odd ideas get tested. Most die fast, which is the point.
One warning. A prototype is a way to communicate, not a design system. The urge to ship what the model generated because it looks finished is real, and resisting it takes discipline.
Meeting notes. Status summaries. Ticket drafts. The third stakeholder update this week explaining the same delay in a slightly different tone. None of this is hard. All of it is required. Together it's the tax you pay for being the connective tissue of a team.
This is the least interesting use of AI and it has the best return. It's boring, it's safe, and the worst case is a slightly wrong summary that someone fixes in a reply. It hands back hours every week. If you're a PM who hasn't automated this layer yet, start here rather than with strategy work. It's the safest place to build the habit.
The line I hold is simple. Tasks yes, decisions no. A model can draft the update explaining that we cut scope. It shouldn't be the thing that decided to cut scope. Those two look similar in a document and are nothing alike inside a company.
Now the part that matters most, because this is where I've watched good PMs get worse.
Ask a model for a product vision and you'll get something clear, well structured and completely forgettable. It reads like every vision statement ever written, because that's what it is: an average of all of them. That's exactly the wrong output. A vision has to be non obvious. If it sounds like consensus, it either isn't a vision or it isn't yours.
Strategy is similar, with a twist. AI is genuinely good at the inputs to strategy. Mapping competitors. Surfacing pricing models from a nearby market. Pulling the regulation I hadn't thought about. It's also a very good critic. I get more value from "attack this strategy the way our sharpest competitor would" than from anything it writes on its own.
What it can't do is make the bet. Strategy is an act of saying no. Picking this segment means giving up that one. Picking this architecture means we can't serve those customers for two years. You make those calls with incomplete data, and the consequences land on real people. A model predicting a plausible next word has no stake in the outcome and no taste. Taste comes from having been wrong in specific, expensive ways.
AI is a great researcher and a brutal critic. It's a weak visionary and a risky strategist. Risky because the writing is smooth enough that a tired PM will read fluency as rigour.
Everything above depends on one thing, and it isn't clever prompt wording.
A generic prompt gives you the median of the internet. Ask for a competitive analysis and you get an average competitive analysis, because average is what you asked for.
The fix is to set the reference point before asking for the work. I keep a small library of context I paste in. The framework I want it to apply. Two or three examples of work I think is excellent. My team's operating principles. The specific ways things go wrong in our domain. Then I ask.
The gap between "write a product strategy for X" and "here's how our team defines a winning strategy, here are two strategy docs I think are exceptional and why, here's what our market punishes, now draft one for X and then argue against it" is not a small quality bump. It's the difference between output I delete and output I build on.
Which leads somewhere a bit uncomfortable. What you get out is capped by the taste you put in. These tools don't level the field between a thoughtful PM and an average one. They widen the gap, because the thoughtful PM has better references to feed them.
If AI takes over a big share of research, synthesis, drafting and admin, what's left is what a PM is really for. It's a smaller job than the old description and a harder one:
I don't think AI replaces product managers. I think it removes the parts of the role that were never really product management, the transcribing and summarising and formatting and queueing, and leaves the parts that always were. Judgement, taste, conviction and ownership.
That's a harder job. It's also a better one.
The PMs who fall behind won't be the ones who refused to use these tools. They'll be the ones who used them on the wrong half of the job, handing over the thinking and keeping the typing.