Selected work

Projects

A few representative builds. Filter by stack to narrow down.

Internal AI Agent Evaluation Lab
2026
Synthetic evaluation lab for AI-agent reliability: 350 golden cases, 60 red-team cases, RAG evaluation, safe refusal, safety classifiers, prevalence estimation, human-review simulation, mitigation impact, release gate reporting, FastAPI, OTel tracing, and CI.
- Python
- FastAPI
- Streamlit
- Pydantic
- pytest
- ruff
- Docker
- GitHub Actions
- GitHub Pages
View source
LLM Red-Team Evaluation Harness
2026
Reproducible benchmark measuring how published adversarial prompts perform against 2026-era LLMs and whether prompt-only defences move the needle — with cross-judge validation and bootstrap confidence intervals.
- Python
- Claude Sonnet 4.6
- Llama 3.1 8B
- Inspect AI
- GitHub Actions
- pytest
- ruff
- mypy
View source