Getting AI coaching feedback to 85% agreement with real trainers.

Role:

AI Design Lead

Timeline:

10 Weeks

Team:

2 Designers, 3 Engineers

2 Designers,

3 Engineers

Company:

Product Manager Accelerator

Overview:

People who strength train at home can't afford for a personal trainer and generic fitness apps don't watch how anyone actually moves. We designed an AI powered app that gives feedback like a real coach on their actual form and helps them progress in their strength training journey.

My responsibilities:

  1. Designed the core flow of version 1 of the app.

  1. Built and maintained a design system using Claude Code.

  1. Studied and used AI-UX patterns like stream-of-thought and quality-gate UI.

Impact:

85%

85%

Match with real trainers' judgement.

Match with real trainers' judgement.

90%

90%

Core-flow task completion rate.

Core-flow task completion rate.

People training strength at home had no way to know when to progress.

  1. YouTube never told her whether her form matched what she was watching.

  1. No one corrected her. leading to her progress stalling for long stretches.

  1. One bad set without guidance would take away all her confidence to go heavier.

20 interviews, people wanting correction while they were lifting in real time.

  1. Understanding form and stance.

  1. Knowing when can I increase my weights

  1. Giving them a well analyzed verdict of what they need to do next.

One clear signal against a majority. We built for the majority

One interview didn't fit that pattern. A different person,

further along in their own thinking about form, said the opposite

One interview didn't fit that pattern. A different person, further along in their own thinking about form, said the opposite

"Even pressing start is too much effort for me when lifting."

"Feedback should come after the set, not during."

"A correction mid-lift breaks focus instead of helping."

Google stitch helped design quick solutions but lacked most of the context.

I had my two designers take the PM's basic concept, upload a video and get analysis back, and run it through Google Stitch to generate a flow fast.

It followed the idea literally but missed the context underneath it and generating more variations from that same flow would have meant polishing something built on a gap, not fixing the gap

I Stepped back to build the entire architecture of the app using OOUX.

Instead of spending more cycles refining a flow with a hole in it, I stopped the team to map out:

  1. Every core object in the app.

  1. How they relate to each other.

  1. what action a user can take on each one and what

    the call to action is.

  1. what action a user can take on each one and what the call to action is.

Back to Stitch, now with something rules, context and a design system to refer to.

It produced variations we could actually judge against something, rather than against taste:

  1. How many steps it took to upload a video outside a workout flow.

  1. How many steps inside one.

  1. How logging weight, reps, and picking which set to film actually worked.

Studying how AI products talk to people who don't trust them yet

I went through Shape of AI and Microsoft's AI UX guidance, looking specifically at products shipped in the last two years and how they handled someone with no AI background. Three patterns from that research mattered enough to build in.

The wait, made visible

Each step reports as it finishes: uploading, counting reps, reading form, writing it up. The rep count shows the moment it's known.

Telling the user when it can't do the job

If footage can't be graded honestly, the app says so, names what went wrong, and tells the user how to fix it next time. The set still counts in the log

Naming what it can't do yet

The app tells users directly that its reads are weaker for body types the training data

covers thinly, and that feedback ships at 85% coach agreement or it doesn't ship at all

Real-time feedback broke the moment people used it

  1. We built v0.1 the way the majority of research pointed, spoken feedback during the set itself.

  1. It didn't land as coaching. It landed as an interruption. People were mid-lift, focused, and a voice correcting them broke the exact concentration they needed to lift well

  1. What they actually wanted was feedback after the set, when they had the attention to use it.

I used the remaining four weeks to rebuild almost everything.

The meant rebuilding how a video got uploaded, how a set got logged, and replacing spoken, mid-lift correction with a written read delivered after the set was done

Legs is scored. The rest are shown as still

in validation.

Every unavailable lift says why, with its coach-agreement number.

Launched the first version in a span of 10 weeks.

The live set takes weight and reps,

then upload or log.

Mark which sets gets a video

to analyze.

Score, what happened, best and worst rep,

four areas, weight call.

Rep-level and session-level views with trend chip

following the slope.

Every day, every set: weight, reps, score and session volume.

Units, storage, and how the score gets validated against coaches.

However, the upload still makes people leave the app

You film in your own camera app, come back to Kinetic, then pick the clip from your gallery. It's a step that shouldn't exist, and it creates real errors: the wrong clip, or the same one uploaded twice

I'd build recording into the app itself. The filming guidance we wrote as an onboarding screen would become a camera overlay, which is where it belonged from the start.

One statement, one primary action,

one link to the last read.

Shown before every analysis, from

both entry points.

Validating against with 10 real coaches.

The subject performed a goblet squat on camera, filmed on an ordinary phone. Right there, the coach gave the same feedback he'd give training that person in person, and we wrote down everything he said.

Agreement was measured two ways: did the app flag the same issues the coach flagged, and how close was the actual language to what the coach said. Both together produced the 85% figure, repeated across 10 coaches over the 10-week build.

85% agreement. Shipped and live.

Matched against real coaches, live, across 10 coaches and a 10-week build.

Feedback ships at that threshold or it doesn't ship, which is why version one scores

legs and nothing else yet.

85%

Coach agreement.

10

Week build.

4

Weeks to rebuild post-pivot.

See my other projects.

Enterprise

User Research

Information Architecture

Redesigned and launched a feedback collection system for service desk employees

reducing workflow time by 67%

Retail

Onboarding UX

Redesign

Designed a new loyalty-program flow with 70% of the non-member shoppers agreeing to sign-up.

CT