work · Boston, MA
DMSB AI Strategic Hub (DASH)
AI Software Engineer
A year of shipping GenAI that people actually used: an AI grading platform faculty adopted, a voice coach with 500+ daily users, and the GPU infrastructure and observability underneath both.
Measured
- Grading API response time
30–60s→ 100ms - Faculty grading time
8h+→ 30 min - Voice coach response latency
90s→ 40s - Mean time to recovery
2h→ 30 min - Deploy time
30 min→ 3 min
- $30K
- Infrastructure cost saved per year
- 20x
- Grading capacity
- 90%
- RAG accuracy, A/B tested on 500+ essays
- 500+
- Voice coach daily active users
Inside this span
- 01Built an AI grading platform with LangGraph agents, adopted by 67% of 200+ faculty
- 02Found the I/O bottleneck and rebuilt grading as queue-based microservices on Amazon MQ, ElastiCache and ECS
- 03Engineered a RAG pipeline on LlamaIndex, Qdrant and BGE-large with metadata filtering and drift detection
- 04Built an AI voice coach with real-time speech analysis and parallel multi-agent workflows
- 05Extracted Whisper into a shared GPU microservice, taking memory from linear to O(1) and workers from 3 to 40
- 06Deployed FERPA-compliant AWS infrastructure and self-hosted vLLM, with a Grafana, Loki and Promtail observability stack