Tell us why.

Every AI product gets told it's wrong dozens of times a day, at the exact moment the user knows precisely what's wrong with it. Almost every AI product throws that away.

Whirl AI - CR-4471

Whirl found a problem

The expense system will silently drop records

Adding a second approver produces a list instead of a single name. Concur expects one name and discards the rest - without showing an error.

INT-Concur-PO-Sync · 3 prior incidents · 94% confidence

Whirl found a problem

The expense system will silently drop records

Adding a second approver produces a list instead of a single name. Concur expects one name and discards the rest - without showing an error.

INT-Concur-PO-Sync · 3 prior incidents · 94% confidence

What did we get wrong?

One tap. This goes straight to the team.

Whirl found a problem

The expense system will silently drop records

Adding a second approver produces a list instead of a single name. Concur expects one name and discards the rest - without showing an error.

INT-Concur-PO-Sync · 3 prior incidents · 94% confidence

Noted.

Confirmations are stored too. Without them there's no denominator - you'd know how often the product was called wrong, but not out of how many.

Finding
FeatureOne-tap rejection feedback
Built forWhirl AI - enterprise IT agents
ByKaran Chulliparambil

01 - The moment

Catch them
while they're
annoyed.

Whirl AI reads a company's systems and warns you what will break if you change something. Sometimes it's wrong. When it is, the person reading it knows exactly why - right then, for about four seconds.

That is the only moment they will ever be able to tell you precisely what went wrong. A survey next week gets you nothing.

Try it - click "This is wrong"

Whirl AI - CR-4471

Whirl found a problem

The expense system will silently drop records

Adding a second approver produces a list instead of a single name. The connection to Concur expects one name, and discards anything it doesn't understand - without showing an error.

INT-Concur-PO-Sync · 3 prior incidents · 94% confidence

What did we get wrong?

One tap. This goes straight to the team.

Logged.

02 - The decision

Four buttons
instead of a
text box.

The obvious design is a box that says "tell us more." It feels generous. It is the wrong answer, for two reasons.

Almost nobody types in it - they're mid-task and annoyed, not writing a bug report. And the handful of paragraphs you do get can't be counted, which means you can never prove anything got better.

The obvious version
Tell us what went wrong…

3%

type anything at all. What you get back is prose - impossible to count, impossible to compare week to week, impossible to prove you fixed.

This version
Already fixed Wrong system Not a risk

61%

tap something. Every tap is the same shape as every other tap - which means it becomes a number, and a number can go down.

03 - The same feature, elsewhere

It's a pattern,
not a one-off.

The same feature works on anything the AI produces. Here it is on the current-state map - where the AI drew a step that doesn't exist anymore.

The reasons are different because the mistake is different. That's the point: the shape stays the same, the vocabulary follows the surface.

Try it - click "This is wrong" on step 03

Whirl AI - Procure-to-Pay · as-is

Whirl mapped the current process

Purchase order approval - today

4 steps · 3 systems

STEP 01
Requisition raised
SAP Ariba
STEP 02
Manager approval
SAP Ariba
STEP 03
Finance review
manual · email
STEP 04
PO issued
S/4HANA

What's wrong with this step?

Different surface, same one-tap. Reasons match the thing being questioned.

Logged.

04 - The payoff

A complaint
you can count
is a complaint
you can kill.

Once every dismissal is one of a handful of shapes, the team gets something they've never had: a measure of how wrong the product is, by reason, over time.

And when they fix something, they can watch the reason disappear. That's the point. Not the chart - the proof.

"Out of date" - times tapped per week

Eight weeks

Fix shipped
31
38
44
50
19
6
2
1
W1W2W3W4W5W6W7W8

People kept saying findings were stale. One data source was refreshing weekly instead of hourly. Nobody ran a study - the product told on itself, because someone made it one tap instead of a paragraph.

05 - This already worked once

"When a customer is clicking a button in anger, they are more than happy to give you a piece of their mind."

Snowflake rebuilt their console and customers said it was slow. Engineers assumed a speed problem and started optimising code.

But they had put a button in the new version that let people go back to the old one - and asked why on the way out. They stored the answers and counted the words.

One word kept spiking: "tabs." Nothing was slow. Users just had to click back to a list page every time they wanted to switch between two things they were working on.

They added tabs. The word disappeared from the complaints. A performance problem that was never a performance problem, found by a button.

James, front-end engineer at Snowflake - on rebuilding the Snowflake console

06 - Decisions

Five choices,
and why.

01

It fires at rejection, never on a timer

The prompt only appears when someone actively says the product is wrong. Never on a schedule, never on page load, never after a task went fine. Feedback asked for at a random moment is noise. Feedback asked for at the moment of friction is a diagnosis.

02

Tapping, not typing

Four fixed reasons instead of an open box. This trades richness for countability, on purpose. You lose the eloquent paragraph from the one person who writes essays. You gain a number you can watch move.

03

Skipping is one tap and always visible

There is a plain skip that dismisses without answering, and it is never hidden or greyed out. A feedback prompt that traps you produces angry garbage - people will click anything to get out. Making it easy to leave is what keeps the data honest.

04

You get a receipt

After tapping, the person is told what happened to their answer - that the finding is gone, and that this reason has come up before. Most feedback UI is a black hole, which teaches people it isn't worth using. One sentence back is what makes them do it a second time.

05

Confirming is captured too

"Looks right" is also a tap, also stored. Without it you only ever measure failure and have no denominator - you would know forty people said something was wrong, but not whether that is out of fifty or five thousand.

07 - What I'd want to be wrong about

The parts I'm
not sure about.

The reasons are a guess. I picked them by imagining the ways an AI finding can be wrong, not by watching anyone. In practice the right list is discovered in the first month, and the honest version of this ships with a rough list and a fast way to change it.

The 3% and 61% are industry-typical for open-text versus tap-to-answer prompts. They are not measured on this product, and I would want to verify them here before quoting them to anyone.

One more thing I'd want to work through with the Whirl AI team is which reasons are too broad. If a reason is too broad, every rejection lands in the same bucket and the chart tells you nothing. "Wrong system" is doing a lot of work in my version and would probably need splitting.