Lead SRE / AI Platform Engineer
Feb 2022 — PresentRemote
- › Shipped a production multi-agent system that runs thousands of deployment investigations daily, 24/7: agents on a durable workflow engine route each task to the right Claude model, verify failures against live evidence, and escalate low-confidence cases.
- › Cut LLM cost per investigation 44-57% (~62K to ~30K tokens) with a deterministic gate that skips a redundant model call on ~70% of investigations, 10-15x tool-output compression, and per-run cost attribution built from scratch.
- › Own reliability for an internal developer platform deploying 300+ services across commercial and GovCloud (FedRAMP) environments, holding a 99%+ deployment success rate on AWS (EKS, Lambda, Step Functions).
- › Raised SLI accuracy from a 42% noise rate to actionable signal by classifying pipeline failures as platform vs service-team issues and calibrating SLO alert thresholds against 60-day statistical baselines.
- › Designed the evaluation methodology for AI code review at enterprise scale, calibrating security, compliance, performance, and maintainability checks against thousands of production pull requests before rollout to developers.