Machine learning
SignSpeak
ASL fingerspelling recognised in the browser, measured on signers it never trained on
What I built
- Rebuilt an ASL fingerspelling recogniser on MediaPipe and PyTorch, 92.2% across 24 letters on held-out signers
- Measured 98.7% random against 92.2% signer-disjoint and published the 6.5-point gap, not the flattering number
- Shipped inference client-side with MediaPipe Tasks and a hand-written forward pass, so no video leaves the device
- Designed the front end on design tokens, two ten-step colour scales and one type scale, with four load-time charts
- Built a 114-test suite across three browser engines, finding seven defects a green single-browser run had missed
- Cut time-to-demo from 3.8s to 55ms by starting the 8.7MB model download on intent rather than on click
How the numbers were measured
Rebuilt an ASL fingerspelling recogniser on MediaPipe and PyTorch, 92.2% across 24 letters on held-out signers
92.2% mean accuracy under leave-one-signer-out cross-validation over the 5 signers of the Surrey ASL Fingerspelling dataset (Pugeault & Bowden, ICCV 2011): train on four people, test on the fifth, five times, so each of 65,522 landmark samples is held out exactly once and no frame of the test person appears in training. Per-fold 87.5%-95.4%, sd 2.6%. Reproducible from the repo with `ml/evaluate.py`; protocol, per-letter accuracy and confusion matrix in docs/EVALUATION.md, which is generated from eval/results.json rather than written by hand.
Measured 98.7% random against 92.2% signer-disjoint and published the 6.5-point gap, not the flattering number
Both protocols run in `ml/evaluate.py` with the same training code and the same seed. The dataset is continuous recording sessions, so a random split puts near-duplicate frames on both sides of the boundary; the 98.7% is reported specifically as the number not to trust. That claim is measured rather than asserted: under a random split a held-out hand sits a median 0.10 from its nearest training neighbour with 3.4% inside 0.05, and splitting by signer doubles the distance and takes that share to zero. The same measurement rules out one recording session being labelled as two people, which would make the honest protocol leak while looking careful. Both figures are rendered on the live page from eval/results.json, so the site cannot claim an accuracy the repository did not measure. The same protocol also measures calibration. The model is under-confident by 3.3 points, ECE 3.4%, which is the label smoothing in the loss showing up from outside. And the confidence floor the demo commits at, which discards 12% of frames to buy five points of accuracy.
Built a 114-test suite across three browser engines, finding seven defects a green single-browser run had missed
33 pytest tests and 36 browser tests, the browser suite run on Chromium, Firefox and WebKit. 94 green runs and 14 skipped on WebKit, where a camera cannot be faked, for a reason that was measured rather than assumed. No flakes over repeated runs. The suite reconciles the published accuracy against its own confusion matrix, checks the checkpoint, the ONNX export and the browser weights against each other, and asserts 240/240 browser-versus-Python agreement on real hands. An independent full re-run of the evaluation reproduces eval/results.json byte for byte, same SHA-256. The defects it found are listed in the repository's commit history. 54 Python and 112 browser assertions in total, plus a strict content-security policy enforced locally and in CI rather than only in production.
Cut time-to-demo from 3.8s to 55ms by starting the 8.7MB model download on intent rather than on click
Measured against the deployed site with Playwright and CDP throttling. The idle page is 0.36MB with first paint at 308ms; starting the camera pulls 8.7MB (the MediaPipe WebAssembly build plus a 7.8MB hand-landmark model) and took 5.2s on fast wifi, 38s throttled to 4Mbps, behind one unchanging label. Prefetching on hover, focus or touch of the button makes the warm path 55ms, while a visitor who never engages pays nothing. Progress is computed against the real uncompressed size, stamped at build time, because Content-Length is the compressed length on the wire and would run the counter past 100%.
Built with
PythonPyTorchMediaPipeComputer VisionMachine LearningNeural NetworksFeature EngineeringModel EvaluationCross ValidationHand Landmark DetectionNumPyOpenCVONNXJavaScriptES modulesHTMLCSSCanvas APIgetUserMediaStatic SiteModel CalibrationAblation StudiesNearest Neighbour SearchData Leakage AnalysisContent Security PolicyWeb security headersCross Browser TestingAccessibility TestingTest DesignDesign SystemsData VisualizationDesign TokensResponsive DesignWeb accessibilityVariable fontspytestPlaywrightGitHub ActionsVercelMultiprocessingBatch NormalizationData AugmentationConfusion MatrixSoftmaxMLPScikit LearnOnnxruntimeMatplotlibPillowFfmpegWebassemblyCss Custom PropertiesCss GridIntersection ObserverSvgAdamwCosine AnnealingLabel SmoothingDropoutgitleaksHttp Range RequestsAxe CoreWcagFirefoxWebkit
Why this stack
mediapipe/feature-engineering. The design decision the project rests on is classifying 21 hand landmarks rather than pixels, then stripping position, scale, rotation and handedness out of those landmarks so only the pose survives. That is what makes the model generalise across people, and it is also what makes P and Q confusable, since those two signs differ largely by orientation. Cross-validation/model-evaluation. Leave-one-signer-out over five people, reported beside the random-split number it invalidates. Onnx/onnxruntime. The model is exported to ONNX and the export is verified against PyTorch before it ships; the browser runs a hand-written forward pass instead, checked against PyTorch by a test, so the page ships no inference runtime. Playwright. Real dataset images are pushed through Chromium end to end and required to reach the same letter as Python. Webassembly. MediaPipe's wasm build is vendored and cache-headed rather than pulled from a CDN. Http-range-requests. The 2.2 GB dataset download was parallelised into eight verified byte ranges after a single stream ran at 700 KB/s. Design-systems/design-tokens. The page is built on two ten-step colour scales with fixed semantic roles rather than hand-picked hexes, so elevation and state are a step change instead of a new colour. Data-visualization. The evidence section is four charts built from the evaluation output at page load, and the two decisions worth defending are plotting per-letter *error* rather than accuracy, because 24 bars between 70% and 99% of full discriminate nothing and truncating the baseline to fix that is dishonest, and greying the confusion matrix's own diagonal, because a single ramp over the whole matrix is dominated by it and hides every error the chart exists to show.