Training Sessions: Moving from Rep-by-Rep to Full-Session Intelligence
2026-07-09
Player 1 used to be a rep-by-rep tool. You hit record, did one rep, stopped, and got feedback on that single rep. Then you did it again. And again.
It worked. It was also the slowest, most annoying part of training with the app — for you and for the model.
The problem with rep-by-rep
Recording every rep by hand meant you spent half your session playing camera operator. You missed reps because you forgot to hit record. You stopped and started so much you never found a rhythm. And when it was over, you had a pile of disconnected clips instead of an actual workout.
The bigger problem was hiding behind that. Because the model only ever saw one rep at a time, it could only ever tell you about one rep at a time. It couldn’t tell you which reps were your best. It couldn’t tell you your mechanics fell apart in round four. It couldn’t see the session — because no session existed.
What changed
Training Sessions fixes that. You hit record once and train. While you work, the app watches the whole thing and finds every rep on its own — swings, throws, shots, runs, whatever the movement is.
When you’re done, you don’t have a pile of clips. You have a session.
Once the model understands a whole session, you can ask about any scope — one rep, a few reps, or the whole thing — and get the same coach-level quality either way.
That wasn’t possible before. A rep-by-rep tool can answer “what went wrong on that swing?” It cannot answer “which three reps were my best?” or “am I getting tired by the end?” Those are session questions. They need a model that sees across reps, not just within one.
I tried to build this on Gemini first. It’s a strong model, but on longer video it couldn’t reliably identify reps, and its per-rep analysis drifted — same swing, different feedback on different days. The workarounds (frame-by-frame stitching, chunking) were fragile and slow.
Gemini
- Strong general model — not built for long video
- Couldn’t reliably find reps inside a full session
- Per-rep feedback drifted: same swing, different notes
- Needed frame-by-frame workarounds (slow, fragile)
TwelveLabs Pegasus
- Native video model — ingests a whole clip as one piece
- Finds every rep on its own
- Coach-level accuracy, consistent across sessions
- Faster and cheaper per session
Pegasus is a native video model — built to take in a whole clip and reason about it as one continuous thing, not as a stack of frames. Side by side, it matched coach-level accuracy on individual reps and was more consistent, faster, and cheaper. And because it understands the entire video at once, the session layer finally became possible.
What this means
The old app could critique a rep. The new app can read a workout.
That’s the foundation everything else builds on.
— Jon