Built the agent harnessEngineered the production runtime behind an autonomous sourcing team: specialized agents, Zod-validated tool contracts, shared workflow state, checkpointed execution, retries, fallbacks, and human escalation across Gemini, OpenAI, and Anthropic.
Made quality measurableBuilt reliability loops around every stage with constraint and schema validation, evidence-grounded outputs, hallucination checks, confidence gates, regression evals, and model-based scoring—sustaining 95%+ spec accuracy across 1,200+ monthly requests.
Optimized every tokenReduced AI cost and latency through task-aware model routing, context compaction, token budgets, caching, batching, parallel fan-out, and early exits—reaching 100+ specs per minute and 70% faster turnaround.
Productionized the intelligenceRan the system on AWS Lambda, Step Functions, SQS, CDK, Python, TypeScript, and PostgreSQL with per-run traces, cost/latency telemetry, replayable jobs, and recovery paths powering AutoQuotation, supplier ranking, and product mapping.