The math we're not doing enough of

Giri Ram

In the first two posts in this series, I wrote about UX Guardians and Guardrails - the shift in our role from producing design work to curating the product knowledge that AI tools across the team depend on, and what it actually takes to build a Guardrail that’s worth the name. Both posts made a case for using AI more deliberately.
This one is different. This one is about when not to and how to actually know, rather than guess.
Because here’s what I’m seeing, and I’d guess you’re seeing it too: we’ve gotten good at reaching for AI. We haven’t gotten good at asking whether reaching for it was actually the cheaper option.
The math we skip
When we use AI for a piece of work, we almost always measure the win against the old way of doing the whole task manually. That’s the wrong comparison. It flatters AI every time, because it ignores the cost we’ve added on the other side: prompting, reviewing, correcting, and - per the Guardrails process - curating the output before it goes anywhere near a source of truth.
The comparison that actually matters is:
Cost with AI = time spent prompting + time spent reviewing/correcting + Guardian curation time
Cost without AI = time the work would have taken done manually, start to finish
If the first number isn’t meaningfully smaller than the second, we didn’t save time. We just spent it differently - and probably added a review step we wouldn’t have needed otherwise.
I want us thinking in this comparison as a habit, not an occasional audit. Here’s roughly how I’d want us framing it, using illustrative numbers - not measured data, just the shape of the thinking:
Task | Manual (est.) | With AI: doing + reviewing + curating | Verdict |
|---|---|---|---|
First-pass research synthesis from an existing transcript set | ~4 hrs | ~45 min drafting + ~30 min Guardian review = ~1.25 hrs | Strong win |
Competitor pattern scan for a new feature area | ~3 hrs | ~30 min + ~45 min verifying claims (AI got two competitors wrong) = ~1.25 hrs | Solid win, but the verification step nearly ate the gain |
Novel interaction pattern for an ambiguous edge case | ~2 hrs of focused design thinking | ~20 min generating options + 90 min figuring out none of them actually solve the real constraint + doing it manually anyway = ~2 hrs 40 min | Net loss |
That third row is the point of this whole post. It happens more often than the productivity narrative admits, and it’s invisible unless you’re actually counting.
Worth flagging plainly: this model is deliberately incomplete. It doesn’t price in risk exposure - what it costs if sensitive context ends up somewhere it shouldn’t. That’s a different kind of calculation, one I go into properly later in this series, and it’s why some categories of work need a security gate before the ROI question is even worth asking. A task can win on time and still be the wrong call.
One variable moves every row of this table more than any other: how good the Guardrail was going in. A well-built Guardrail - the kind I described in the last post, specific rather than sprawling, carrying judgment and not just facts - gets you closer to a usable answer on the first attempt, which is most of the difference between the “strong win” row and the “solid win, but nearly ate the gain” row above. A thin or sprawling one means every person who uses that context pays a correction-cycle tax, repeatedly, for as long as the weak Guardrail stays in place. The time invested getting the context right once is time every future user of it doesn’t have to spend fixing it. That’s not a small multiplier - it’s often the entire difference between a task landing in the top row of that table or the bottom one.
When to stop mid-task
The first stopping condition is the one that happens inside a single piece of work, and it’s the easiest to miss because it creeps up gradually.
You prompt. The output’s close but not right. You refine the prompt. Still not right it’s missing context, or it’s technically correct but doesn’t sound like your product, or it keeps re-introducing an assumption you’ve already corrected twice. You refine again.
Here’s the discipline I want us building: count the correction cycles, and set a personal limit before you start. For most UX tasks, if you’re past three rounds of correction and the output still isn’t usable, you have almost certainly crossed into net-loss territory and the fourth attempt rarely fixes what the first three didn’t.
A few signals it’s time to stop and do it yourself:
You’re now spending more time explaining the problem than solving it would take. If you could describe the fix in two sentences to a teammate and they’d get it immediately, you probably don’t need four more prompt iterations.
The AI keeps “fixing” the wrong thing. This usually means the Guardrails context is thin or the task genuinely needs judgment the AI doesn’t have access to — not a prompting problem, a fit problem.
You’re rewriting more than you’re keeping. If your “review and correct” step has quietly become “rewrite from scratch while glancing at the draft,” you’ve absorbed the full manual cost and the AI cost. That’s the worst outcome on the table, and it’s common enough that it deserves a name: double-paying.
Worth being honest about where this usually traces back to: in my experience it’s rarely the model being unreliable. It’s far more often a thin Guardrail - the tool doing its best with sparse or unstructured context, and the sparse context is a choice someone made upstream, not a limitation of the tool itself. Before concluding a task simply isn’t a fit for AI, it’s worth asking whether the real fix is a better-built Guardrail rather than abandoning the attempt.
That said, sometimes it genuinely is the wrong fit, and the discipline is knowing the difference rather than sunk-costing your way through another round.
When to stop at the category level
The second stopping condition is bigger than any single task. Some categories of UX work shouldn’t default to AI at all not because AI can’t produce an output, but because the output isn’t the expensive part of the work.
A few categories I’d want us treating this way:
Work involving ambiguous user harm. If a design decision touches something where getting it wrong has real consequences for a vulnerable user - safety-critical flows, anything adjacent to financial or medical stakes, dark-pattern-adjacent decisions - the judgment call is the whole job. AI can help you gather context faster. It cannot make the call. Don’t let the speed of the draft imply the decision is settled.
Genuinely novel problems. AI is excellent at pattern-matching against what exists. It’s weak at the thing that doesn’t have precedent yet - the interaction that no one’s designed before because the problem itself is new. If you’re working at that edge, AI-assisted drafts will anchor your thinking to existing patterns exactly when you need to escape them. Sometimes the right first move is a blank page, not a prompt.
Trust-building conversations. Stakeholder alignment, conflict resolution, the actual work of getting a sceptical team to buy into a UX process - the war stories I wrote about in the Ecolab case study. AI can help you prepare. It cannot have the conversation for you, and drafting your talking points too heavily can make you sound like you’re reading a script when the moment calls for you to actually be present.
Where users are telling you the AI-shaped answer isn’t landing. If you’re using an AI tool to communicate with users or stakeholders - draft responses, explain a decision, write something meant to build trust - and you’re noticing pushback, confusion, or a sense that something feels off, that’s signal, not noise. Some users respond worse to AI-flavoured communication even when the content is accurate. When you see that pattern, it’s a category-level flag: this type of interaction needs a human voice, not a faster draft.
Where AI clearly earns its place
None of this is an argument against using AI - it’s an argument for using it deliberately. The categories where the math consistently works in AI’s favour:
First-pass synthesis of large, structured inputs - transcripts, survey data, ticket volumes - where the manual alternative is genuinely tedious and the review step is fast because you’re checking synthesis quality, not generating original judgment
Pattern-matching against precedent - competitor scans, “how have others solved this,” accelerating the research phase before the actual design thinking starts
First drafts you were always going to heavily edit anyway - if your process already involves a human rewrite pass, AI doing the zero draft is close to free time saved
Anything the Guardrails have made genuinely well-contextualised - this is where Post 2 pays off directly. The better the Confluence/Jira/JPD context and the more deliberately it was built, the fewer correction cycles, the more the math tilts toward AI
The habit I want us building
Not a formal ROI spreadsheet for every task - that’s its own overhead, and we’d be trading one inefficiency for another. But a quick, honest gut-check, especially on anything that’s taking longer than expected:
Am I past three correction cycles? Am I rewriting more than I’m keeping? Is this a category where the judgment is the expensive part, not the draft? And is this actually an AI-fit problem, or a Guardrail-quality problem?
If the answer to the first three is yes, stop and do it the other way. If the answer to the last one is “Guardrail-quality,” the fix might not be abandoning AI on this task - it might be going back to Step 1 and building the context properly before trying again.
That’s not a failure of the AI experiment. That’s the experiment working exactly as it should: telling you where the tool fits, where it doesn’t, and where the fix is upstream of the tool entirely.
We’re Guardians of the product knowledge. Part of that job is being honest about when the new tools are actually serving the work - and when we’re just using them because they’re there.
Next in this series: the first real experiment write-up, with what we tried, what broke, and what it actually cost us to find out.
This is the third post in the UX Guardians series. Start with Post 1 → · Read Post 2 →
