Home/About
I build production AI systems and the infrastructure underneath them — and I publish what breaks, because that's the part nobody writes about.
Book a consultation callTen years in distributed systems and cloud infrastructure, currently owning end-to-end architecture for a multi-tenant GenAI platform on AWS. Rust, Go, Python and Java across LLM serving, RAG, agentic workflows and MLOps.
RAG on Amazon OpenSearch Serverless and Bedrock Knowledge Bases, with a Rust ingestion path chunking and embedding at 8M records/hour, plus an evaluation harness scoring retrieval precision, groundedness and hallucination rate against a curated golden set.
12M+ documents · 8M records/hour ingestion
Qdrant with HNSW tuning and payload filtering, hybrid retrieval combining BM25 and dense similarity with cross-encoder reranking, open-weight models served via vLLM with continuous batching, evaluation gates wired into CI.
p95 under 25ms at 1.5K QPS · 3.2× serving throughput
Centralised LLM gateway fronting Bedrock — Claude, Llama, Titan — with model routing, token streaming, per-tenant rate limiting, request tracing and cross-region failover. Consolidated three fragmented team integrations into one.
40K+ concurrent connections · sub-100ms overhead
Token-level cost attribution per team and use case via CloudWatch and Cost Explorer tagging, prompt caching and model right-sizing. Separately, moved high-volume batch inference to EKS with vLLM and Karpenter across G6 and Inferentia2.
$180K annual spend removed · 55% lower per-token cost on batch
Multi-step agent systems on Bedrock Agents and Step Functions: tool-calling contracts, MCP server integrations to internal APIs, human-in-the-loop gates on high-risk actions, and replay and audit trails on every decision path.
In production
Bedrock access policies, EKS clusters, OpenSearch collections, VPC endpoints, IAM boundaries — with ArgoCD GitOps delivery and observability through Prometheus, Grafana and OpenTelemetry. Bedrock Guardrails, SageMaker Clarify, CloudTrail audit logging.
SOC 2 aligned · all model traffic on PrivateLink
Engineering Manager at a European telecommunications group — led a team of engineers across platform and applied AI, set engineering standards, ran architecture review, and drove adoption across five business units. Built a Rust and Apache Arrow pipeline processing 10M+ records/minute, and a production HTTPS server handling 50K+ concurrent connections at sub-millisecond response.
Technical Lead at a cloud consultancy — migrated AWS CloudFormation to Terraform HCL across 10 environments, and rewrote the CloudFormation automation tooling in Rust: 3× faster execution, and it eliminated the memory bugs behind 15% of failed deployments.
I publish the architecture and the failures at medium.com/@aiinfra — agentic RAG in Postgres with the raw routing payloads, why 95%-accurate agents finish 100-step jobs 0.6% of the time, and building a semantic cache in Rust with the one metric that tells you it's returning wrong answers.
Code at GitHub. Shorter notes at @c199benzene.
Working together starts with a call. Thirty minutes, no charge — you describe what's breaking and I tell you whether I can help. If an audit is the right next step, it's five days at a fixed price.
Book a consultation callLast updated 13 September 2026