Web application
Verity
Compliance assessment whose citations are verified, and whose retrieval is measured
What I built
- Built a 118-query retrieval eval with a held-out split; LLM reranking hit recall@10 90.6% against 84.4% for dense
- Measured equal-weight RRF cutting nDCG@10 from 0.666 to 0.383, and published the negative result
- Built span-level citation verification labelling each quote exact, near or unsupported, excluding unsupported ones
- Designed a five-stage pipeline over a 413-section FERPA, HIPAA, GDPR and 508 corpus, ~25s at ~$0.008 a document
- Implemented BM25, dense retrieval, RRF and int8 quantization from scratch; quantization cut the index 4x
- Swept chunk size fourfold and kept the shipped config, the best margin being smaller than the held-out set resolved
- Rebuilt the pipeline as an async generator streaming NDJSON, first result at 20 seconds rather than 30
- Replaced a cold-start-resetting rate limiter with a Redis fixed window, and cached repeats from 30s to 0.1s
- Deleted routes that could not run in production, removing about 60% of the source and 28 of 40 dependencies
- Rejected zip decompression bombs after a 217 KB upload took the server from 613 MB to 1,013 MB
- Added a CSP and five missing security headers, with browser tests that fail the build if any is dropped
- Made retrieval degrade honestly without the embedding service, serving BM25 under a banner naming what is missing
- Audited five pages with axe across two themes and viewports, fixing WCAG AA contrast and unreachable tables
How the numbers were measured
Built a 118-query retrieval eval with a held-out split; LLM reranking hit recall@10 90.6% against 84.4% for dense
docs/retrieval-eval.md, generated by eval/run-eval.mts against the committed index and the committed gold set at eval/gold-set.jsonl. 118 labelled queries over 413 sections of regulation text: 76 dev queries phrased close to the regulation's language, 32 held-out paraphrased queries that deliberately avoid the target section's vocabulary, and 10 bare citation lookups. The figures quoted are the 32-query held-out slice. Relevance is judged at section granularity; labels are the author's own, which the report states.
Measured equal-weight RRF cutting nDCG@10 from 0.666 to 0.383, and published the negative result
Same harness and same 32-query held-out slice as vty-eval. The weight sweep is in docs/retrieval-eval.md and selects a lexical weight of zero on the development split. The citation-lookup slice is reported alongside it, where equal-weight fusion does win (nDCG@10 0.950 against dense retrieval's 0.874), because the conclusion is only honest with the counter-case on the page.
Designed a five-stage pipeline over a 413-section FERPA, HIPAA, GDPR and 508 corpus, ~25s at ~$0.008 a document
Latency and cost are read from the pipeline's own trace, which records per-stage timings and takes token counts from the provider's usage metadata rather than estimating them; cost is computed from published Gemini rates held as a constant in src/lib/telemetry/trace.ts. Measured on the sample document against the live deployment; the trace is shown on the page. Corpus size is a fact about data/corpus.json.
Implemented BM25, dense retrieval, RRF and int8 quantization from scratch; quantization cut the index 4x
Implementations are in src/lib/retrieval/. The quantization claim is the Dense and Dense (int8) rows of docs/retrieval-eval.md, which agree at 84.4% recall@10 on the held-out slice and differ by 0.001 nDCG@10. Below the resolution a 32-query slice supports, which the report says. 4x is the ratio of one float32 to one int8 per dimension.
Swept chunk size fourfold and kept the shipped config, the best margin being smaller than the held-out set resolved
docs/chunking-experiment.md, generated by scripts/chunking-experiment.mts. Four full index rebuilds (160/32, 320/64, 320/0, 640/128) each evaluated by the same harness on the same 32-query held-out slice. Dense nDCG@10 across all four spans 0.662 to 0.685; the best alternative beats the shipped configuration by 0.019, against a slice where one query is worth roughly 3.1 recall points.
Rebuilt the pipeline as an async generator streaming NDJSON, first result at 20 seconds rather than 30
Timings read from a live run against the deployed site, printed by the stream itself: classification done at 10.1s, first framework at 20.6s, all five complete at 28.7s. The pipeline is src/lib/pipeline/assess.ts; the route serialises its events in src/app/api/assess/route.ts.
Replaced a cold-start-resetting rate limiter with a Redis fixed window, and cached repeats from 30s to 0.1s
Both figures measured locally against the deployed schema: an uncached run of the sample takes 30.1s, an identical repeat returns in 0.125s. The namespacing was added after test runs were found to be incrementing the same counters as the public deployment.
Deleted routes that could not run in production, removing about 60% of the source and 28 of 40 dependencies
Commit d561a84 in the Verity repository. 278 files changed, 96,098 deletions against 11,692 insertions. Dependency counts are the diff of package.json across that commit. That no environment had DATABASE_URL or the Firebase variables was confirmed against `vercel env ls` and the local .env.local before deleting anything.
Rejected zip decompression bombs after a 217 KB upload took the server from 613 MB to 1,013 MB
Memory figures measured with process.memoryUsage().rss before and after a crafted .docx whose document.xml expands to 200 MB; the same file now returns 422 in 21 ms with no measurable allocation, because the guard reads the zip central directory rather than the archive. Src/lib/documents/archive-guard.ts and its tests. The disclosure was found by running the live endpoints with the API account's credits depleted, which returned the provider's billing message verbatim to any caller; classification now happens in src/lib/errors/public-error.ts. A second copy of the same leak lived in the trace summary, which is also returned to the browser, and was not fixed by fixing the routes. Src/lib/telemetry/__tests__ pins it.
Audited five pages with axe across two themes and viewports, fixing WCAG AA contrast and unreachable tables
e2e/accessibility.spec.ts, 26 checks green against production. The contrast failures were computed rather than eyeballed, against the worst surface each colour renders on: all four light-theme semantic colours failed against their own tinted backgrounds. Verified 3.51, near 3.79, unsupported 4.33, accent 4.49. Which is the case that matters for a badge and the one that passes if you only test against the page. Axe had flagged one of the four; the rest came from doing the arithmetic.
Built with
Next.jsTypeScriptRAGInformation RetrievalBM25Vector SearchEmbeddingsGeminiLangChainEvaluation HarnessPrompt EngineeringLLM rerankingPostgreSQLpgvectorApp RouterReact Server ComponentsStreamingNDJSONLCELStructured OutputConstrained DecodingZodRank FusionQuantizationnDCGMRRRecall@kHallucination DetectionRedisUpstashCachingRate LimitingSentryRechartsPDF parsingDOCX parsingNeonHNSWJestPlaywrightDockerDocker ComposeGitHub ActionsVercelTailwindDesign SystemsCost InstrumentationEcfr ApiEur LexTsvectorTs Rank CdPdfjsMammothgitleaksFfmpegPillowTsxEslintOpengraphInstrument SerifPlaywright TraceServer ActionsData Retention
Why this stack
bm25/rank-fusion/quantization. Okapi BM25, reciprocal rank fusion and symmetric int8 vector quantization are implemented in the repository, not imported, and each is measured against the same gold set. Langchain/lcel. Prompts, LCEL composition and withStructuredOutput are load-bearing; VerityRetriever implements BaseRetriever over the same store the evaluation harness scores, so a chain and the harness rank through one implementation. Embeddings deliberately bypass LangChain, because its Gemini wrapper accepts outputDimensionality and does not send it. Postgres/pgvector/neon/hnsw. 1,147 chunks live in Neon behind an HNSW index as the scale-out path, populated rather than aspirational, and the store is also exercised against a real Postgres in Docker. Redis/upstash. A fixed-window limiter namespaced by environment, replacing an in-memory one that reset on every cold start. Streaming/ndjson. The pipeline is an async generator and the route streams its stages. Ecfr-api/eur-lex. The corpus is fetched from primary sources, so it is reproducible rather than pasted.