← Back to portfolio

Note · June 2026 · 4 min read

First, Prove It

Why I refuse to start a redesign before I can measure it. A note on the most uncomfortable weeks of my favorite project.

In 2024 I joined Solving AI to redesign their automation workflows. Users were abandoning multi-step AI flows, everyone on the team could feel it, and the cure had been chosen before I walked in: a redesign, as soon as possible. The only thing missing was the screens.

I spent my first weeks refusing to draw them.

Not out of stubbornness. Two problems were tangled together, and only one of them was visible. Comprehension was failing: people could not tell what the AI was doing, what it needed from them, or how to read its output. That problem you can see. The second one was quieter and worse: every design decision in the product's history had been argued from opinion, because nothing was measured. A redesign would have treated the visible problem while leaving the machine that produced it untouched. Six months later we would have been arguing about the next redesign, with the same empty hands.

What it actually costs

So my first deliverable was research infrastructure: a Maze prototype test to get a baseline on record before anyone touched the interface. It sounds reasonable written down. Living it was a different thing. Progress reviews where I had nothing visual to show. A roadmap that looked stalled from the outside. The growing suspicion in the room that the new designer was all process and no product.

I defended it with the only argument I had: without a baseline, we would ship a prettier product and never know whether it was a better one.

When the data came in, it paid for those weeks all at once. Only 2 of the 10 participants could tell what to do when they landed on the canvas where models get wired together. The failure was not buried somewhere deep in the flow, which is where most of the internal debate had been pointing. It was the first thirty seconds, before anyone got far enough to configure anything. Every opinion-driven redesign would have polished the wrong rooms of the house.

Design against evidence, then make it stick

The redesign that followed was almost the easy part: an onboarding path for the first arrival, clearer tool selection, tooltips for anyone who skipped the walkthrough, AI states that narrate what the system is doing instead of hiding it, and outputs accessible to screen readers and skeptics alike. In the follow-up test, 8 of the 10 could tell what to do. The two who still could not had skipped the onboarding and then found the tooltip copy itself unreadable on hover, which sent me back for a UX-writing pass on the tool descriptions. That pair taught me more than the eight did.

And because honesty is the whole point of this note: ten participants is a directional test, not a statistical study, and I say so wherever I cite it. A number you cannot defend is worth less than no number at all.

The principle I kept is the one I now apply to everything, including this portfolio: I would rather ship a smaller design I can defend with a result than a bigger one I can only defend with taste. It is why the case studies here lead with what can be verified, why the claims carry their measurement context, and why every flagship screen has a Spec toggle that exposes the reasoning behind it. If I ask stakeholders to trust evidence over opinion, the least my own portfolio can do is show its work.

AI products raise the stakes on all of this. When the interface is a model's behavior, taste cannot tell you if it works. Only users can, and only if you built the instruments to hear them.

The case study behind this note

First, Prove It: A Validation Culture for AI Products →