Web application

Pitchback

Voice sales trainer. Practise a call against a buyer who pushes back

What I built

  • Built the missing speech-to-text half of a voice loop: Whisper, a bounded audio endpoint and a browser recorder
  • Measured the voice loop at 4.0s median and 5.3s p95 over 12 turns, retiring an unmeasurable inherited claim
  • Designed a 7-state machine constraining an LLM buyer: it proposes a move each turn, the machine rejects illegal ones
  • Rebuilt scoring so three of four competencies are arithmetic on the transcript, leaving one quoted model judgement
  • Removed a silent fallback scoring 50 whenever the grader failed, which made broken evaluations look real
  • Drove real push-to-talk in a headless browser on a synthetic capture device, asserting the loop without a microphone
  • Made a paid voice demo cheap to run: rate limits in Postgres, TTS cached by content hash, runs capped at 12 turns

How the numbers were measured

Measured the voice loop at 4.0s median and 5.3s p95 over 12 turns, retiring an unmeasurable inherited claim
Measured 2026-09-09, run cmtuibsii0006a89hfbust544, n=12 turns of the 'discovery' scenario against the production build on localhost with Postgres in Docker. Per-turn totals (ms): 2754 2762 2909 3361 3390 3518 4004 4092 4298 4317 4737 5349. The figure quoted is the app's own, as /api/run/:id/grade returns it. MedianMs 4004, p95Ms 5349. Deliberately, so the resume and the running product cannot disagree; note the app takes the upper of the two middle values, where the conventional median (their mean) would be 3761. Stage means: transcription 1230ms, model 928ms, synthesis 1633ms. CAVEATS, all load-bearing: the rep audio was TTS-synthesised rather than a human recording, so Whisper had unusually clean input; every line was unique because synthesis is cached by content hash and repeats returned in ~5ms, which silently deflates any average taken over repeated text; and this is localhost against a local database, not Vercel serverless, so production will differ
Designed a 7-state machine constraining an LLM buyer: it proposes a move each turn, the machine rejects illegal ones
Seven states and the constraint rules are in lib/sim/buyer-state.ts, covered by 21 unit tests in tests/buyer-state.test.ts. The original 7 states and 4 metrics were verified in the pre-rebrand code first (lib/ai/emotion-engine.ts, lib/types/session.ts).
Rebuilt scoring so three of four competencies are arithmetic on the transcript, leaving one quoted model judgement
lib/sim/scoring.ts, 26 unit tests in tests/scoring.test.ts plus tests/ai-contracts.test.ts covering quote verification. Thresholds follow the shape of Gong's published call analyses; their raw 11-14 question count is deliberately not transplanted, and the reason is written into the module docstring.

Built with

Next.jsTypeScriptReactPostgreSQLPrismaOpenAIWhisperTTSState MachinesTailwindVitestPlaywrightRate LimitingCachingPrompt EngineeringGitHub Actions CIDockerEslintSatori OgFfmpegPillowShadcn Ui

Open it