Getting AI coaching feedback to 85% agreement with real trainers.
Role:
AI Design Lead
Timeline:
10 Weeks
Team:
Company:
Product Manager Accelerator
Overview:
People who strength train at home can't afford for a personal trainer and generic fitness apps don't watch how anyone actually moves. We designed an AI powered app that gives feedback like a real coach on their actual form and helps them progress in their strength training journey.
My responsibilities:
Designed the core flow of version 1 of the app.
Built and maintained a design system using Claude Code.
Studied and used AI-UX patterns like stream-of-thought and quality-gate UI.
Impact:
People training strength at home had no way to know when to progress.
YouTube never told her whether her form matched what she was watching.
No one corrected her. leading to her progress stalling for long stretches.
One bad set without guidance would take away all her confidence to go heavier.
20 interviews, people wanting correction while they were lifting in real time.
Understanding form and stance.
Knowing when can I increase my weights
Giving them a well analyzed verdict of what they need to do next.

One clear signal against a majority. We built for the majority

"Even pressing start is too much effort for me when lifting."
"Feedback should come after the set, not during."
"A correction mid-lift breaks focus instead of helping."
Google stitch helped design quick solutions but lacked most of the context.
I had my two designers take the PM's basic concept, upload a video and get analysis back, and run it through Google Stitch to generate a flow fast.
It followed the idea literally but missed the context underneath it and generating more variations from that same flow would have meant polishing something built on a gap, not fixing the gap

I Stepped back to build the entire architecture of the app using OOUX.
Instead of spending more cycles refining a flow with a hole in it, I stopped the team to map out:
Every core object in the app.
How they relate to each other.


Back to Stitch, now with something rules, context and a design system to refer to.
It produced variations we could actually judge against something, rather than against taste:
How many steps it took to upload a video outside a workout flow.
How many steps inside one.
How logging weight, reps, and picking which set to film actually worked.
Studying how AI products talk to people who don't trust them yet
I went through Shape of AI and Microsoft's AI UX guidance, looking specifically at products shipped in the last two years and how they handled someone with no AI background. Three patterns from that research mattered enough to build in.

The wait, made visible
Each step reports as it finishes: uploading, counting reps, reading form, writing it up. The rep count shows the moment it's known.

Telling the user when it can't do the job
If footage can't be graded honestly, the app says so, names what went wrong, and tells the user how to fix it next time. The set still counts in the log

Naming what it can't do yet
The app tells users directly that its reads are weaker for body types the training data
covers thinly, and that feedback ships at 85% coach agreement or it doesn't ship at all
Real-time feedback broke the moment people used it
We built v0.1 the way the majority of research pointed, spoken feedback during the set itself.
It didn't land as coaching. It landed as an interruption. People were mid-lift, focused, and a voice correcting them broke the exact concentration they needed to lift well
What they actually wanted was feedback after the set, when they had the attention to use it.

I used the remaining four weeks to rebuild almost everything.
The meant rebuilding how a video got uploaded, how a set got logged, and replacing spoken, mid-lift correction with a written read delivered after the set was done

Legs is scored. The rest are shown as still
in validation.

Every unavailable lift says why, with its coach-agreement number.
Launched the first version in a span of 10 weeks.

The live set takes weight and reps,
then upload or log.

Mark which sets gets a video
to analyze.
Score, what happened, best and worst rep,
four areas, weight call.
Rep-level and session-level views with trend chip
following the slope.

Every day, every set: weight, reps, score and session volume.

Units, storage, and how the score gets validated against coaches.
However, the upload still makes people leave the app
You film in your own camera app, come back to Kinetic, then pick the clip from your gallery. It's a step that shouldn't exist, and it creates real errors: the wrong clip, or the same one uploaded twice
I'd build recording into the app itself. The filming guidance we wrote as an onboarding screen would become a camera overlay, which is where it belonged from the start.

One statement, one primary action,
one link to the last read.
Shown before every analysis, from
both entry points.
Validating against with 10 real coaches.
The subject performed a goblet squat on camera, filmed on an ordinary phone. Right there, the coach gave the same feedback he'd give training that person in person, and we wrote down everything he said.
Agreement was measured two ways: did the app flag the same issues the coach flagged, and how close was the actual language to what the coach said. Both together produced the 85% figure, repeated across 10 coaches over the 10-week build.

85% agreement. Shipped and live.
Matched against real coaches, live, across 10 coaches and a 10-week build.
Feedback ships at that threshold or it doesn't ship, which is why version one scores
legs and nothing else yet.
85%
Coach agreement.
10
Week build.
4
Weeks to rebuild post-pivot.
See my other projects.


Enterprise
User Research
Information Architecture
Redesigned and launched a feedback collection system for service desk employees
reducing workflow time by 67%


Retail
Onboarding UX
Redesign
Designed a new loyalty-program flow with 70% of the non-member shoppers agreeing to sign-up.
SAN FRANCISCO, ca
20
°C




