Engineering writeups on AI systems, product architecture, and the operational details that make learning platforms reliable.
AI EvalsLLM EngineeringObservabilityQuality
AI Evals 101: How to Actually Know If Your AI Is Any Good
August 4, 202610 min read
A practical guide to tasks, datasets, scorers, offline and online evals, human review, production spans, and the quality flywheel that turns failures into regression tests.
Building a Multimodal Agentic Pipeline for Educational Content Ingestion
July 16, 202611 min read
An engineering case study on turning publisher PDFs into DoK-tagged questions for Wayground, printable as worksheets and playable as formative assessments, built with OCR, LangGraph, SymPy verification, and quality gates.