AI GovOps · AI FinOps · AI Performance Engineering · AI Development

I make enterprise AI governed, cost-controlled and fast — in production.

I'm John D'Esposito, founder of Amalfi AI. For 20+ years I've built production infrastructure that has to be fast, observable and auditable — API gateways, cloud platforms, performance engineering. Now I apply that discipline to AI: agents, LLM traffic and the data platforms they run on.

Databricks Snowflake Amazon Bedrock LangGraph Kong AI Gateway Nginx / OpenResty Lua Claude Code Gemma & local LLMs MLflow
United Airlines

Ran the Kong API platform for mobile

Kong Gateway and Kong Mesh at airline scale, automated across every AWS region, with custom Lua plugins for authorization, request validation and routing — plus early Kong AI Gateway adoption.

Goldman Sachs

Apple Card cloud infrastructure

Multi-account, multi-region infrastructure for customer-facing banking transactions, with CIS-tested AMI pipelines.

Production AI

Multi-agent triage in production

A LangGraph system that classifies, validates and diagnoses operational tickets against live warehouse data. Case study ↓

Patents

3 patents in performance & load balancing

Real-time network and HTTP monitoring, server clustering for load balancing, and web-performance testing.

Featured Case Study

TriageBot: the agent proposes, deterministic code decides.

Client: a national healthcare payments company · Multi-agent AI · Snowflake · Amazon Bedrock

Operational tickets arrived through several service desks and issue trackers. Many were underspecified, routing was manual, and diagnosis meant senior engineers reconstructing system state by hand across the data warehouse.

I designed and built TriageBot, a multi-stage agent pipeline that became the single front door. It kicks back incomplete tickets before they consume engineering time, classifies the rest, and diagnoses root cause from live warehouse state — reconstructing what the data looked like when the problem occurred — then proposes a remediation for an engineer to approve.

The model never has the last word. It adjudicates over a closed registry of known failure modes, recommendations come from a deterministic rule table, and when the evidence isn't there the system abstains instead of guessing. Diagnosis that once meant an engineer reconstructing system state by hand now arrives as an evidence-backed brief, ready for review.

88–93%

Classification accuracy on live weekly holdouts — above the 84% rate at which human agents agreed with each other.

95%

Accuracy on high-confidence predictions — the ones safe to automate.

27→90%

Recommendation agreement, from early development to production holdouts.

0

Fine-tuning runs. It learns from human-verified precedent in retrieval, not from retraining.

The six-stage pipeline
  1. Hydrate

    Pull the ticket and its context from the source system.

    Code
  2. Classify

    Hybrid semantic + lexical retrieval against resolved precedent.

    Model proposes
  3. Validate

    Required fields checked; incomplete tickets kicked back.

    Code decides
  4. Diagnose

    Point-in-time warehouse queries, adjudicated against known failure modes.

    Model proposes
  5. Recommend

    Deterministic rule table; invariants stop contradictions.

    Code decides
  6. Act

    Write back to the tracker for a human to approve.

    Human approves

Evaluation before shipping

No prompt or code change ships without passing regression gates over frozen evaluation sets and blind weekly holdouts, traced end to end in MLflow.

Cost discipline by design

Frontier models only where judgement is needed; small open-weight models run off-peak for extraction and evaluation. Bounded queries per ticket — no runaway loops.

Auditable end to end

Every incoming event is kept in an immutable audit store for replay, and every decision is traceable to the evidence that produced it.

LangGraphAmazon BedrockOpenSearch ServerlessSnowflakeMLflow TracingKubernetes (EKS)TerraformOpen-weight SLMs
What I Help With

Four disciplines. One production standard.

Every engagement starts with a measurement you can act on, and ends with controls running in your own infrastructure. Consulting is delivered through Amalfi AI.

01 — Governance

AI GovOps

Governance as a continuous flow, not a quarterly checklist.
The problem
AI usage is growing faster than anyone can review it, and audit evidence is reconstructed after the fact.
The outcome
Risk classification, guardrails and audit evidence produced at the point of every AI interaction, mapped to ISO 42001, NIST AI RMF, the EU AI Act and HIPAA.
Typical first engagement
Governance baseline — where your AI traffic flows, what is exposed, and the controls that close the gap.
02 — Cost

AI FinOps

Know what every agent, team and model costs — then stop the waste.
The problem
AI spend arrives as one monthly bill with no owner, and agents can burn budget faster than anyone notices.
The outcome
Every AI call — LLM, MCP, agent-to-agent — metered, attributed and reconciled, with costs reduced in-line and budgets enforced where the spend happens.
Typical first engagement
AI spend baseline — your real spend per team, agent and model from live traffic, with no application changes.
03 — Performance

AI Performance Engineering

Fast, reliable AI in production — engineered, not hoped for.
The problem
Latency, rate limits and provider outages surface as user complaints instead of being designed for.
The outcome
Latency, throughput and resilience engineered into the AI path — down to the Nginx, OpenResty and Lua layer of the gateway — and held to SLOs with real telemetry, across every model and provider.
Typical first engagement
Performance baseline — SLOs, tail latency and failure modes measured, with the fixes ranked by impact.
04 — Development

AI Development

Agents that are accurate because they are constrained.
The problem
Prototypes impress in demos, then stall: no evaluation, no guardrails, no way to prove they are right.
The outcome
Production agents and LLM applications grounded in your data — on frontier, open-weight or local models — with deterministic guardrails, evaluation gates and tracing from day one.
Typical first engagement
Production pilot — one real workflow, an evaluation set built from your history, and a go/no-go backed by measured accuracy.
Where I Build

On the platforms you already run.

AI governance, cost and performance live in the data platform that holds the evidence, the gateway that carries the traffic, the observability stack that watches it — and the models themselves, hosted or local.

Databricks

Medallion architecture with Unity Catalog lineage from raw data to model prediction; MLflow for experiments, tracing and evaluation; Mosaic AI Gateway and Lakehouse Monitoring for governed model serving.

Unity CatalogMLflowMosaic AI GatewayDelta LakePySpark

Snowflake

Agents grounded in warehouse truth: guarded, read-only, point-in-time SQL that reconstructs system state, and query-history mining that turns undocumented expert practice into reusable playbooks.

Snowflake SQLData-grounded agentsForensics

Gateway & observability

Kong AI Gateway, Kong Mesh and the Nginx / OpenResty engine beneath them as the enforcement point for AI traffic — custom Lua plugins for authorization, prompt validation and routing, and the data plane tuned for latency. OpenTelemetry and Dynatrace put spend per call, latency per model and governance SLAs on one pane.

Kong AI GatewayKong MeshNginxOpenRestyLuaOpenTelemetryDynatraceAWSTerraform

Models & AI tooling

Model-agnostic by design. Frontier models — Anthropic Claude, OpenAI GPT, Google Gemini — through Amazon Bedrock and private endpoints; open-weight models such as Gemma and Qwen served locally with vLLM where cost, privacy or latency call for it. An AI gateway like Kong puts every model, hosted or local, behind the same governance, metering and SLOs. Day to day I build with Claude Code, custom MCP servers and multi-model review panels.

ClaudeClaude CodeOpenAIGeminiGemmaQwenvLLMAmazon BedrockMCP
How I Work

Measure. Engineer. Operate.

I work with a small number of enterprises at a time, close to the engineering team. No slide-deck strategies — controls, metering and agents that run in your infrastructure.

01

Measure

A short baseline from your live AI traffic and data: what it costs, how it performs, how accurate it is, and where the governance exposure sits.

Outcome — a baseline you can act on

02

Engineer

Guardrails, metering, performance policies and agents built as code, tested like software, and embedded in the platforms you already run.

Outcome — controls in the request path

03

Operate

SLOs, budgets, accuracy and audit evidence reviewed continuously as new models, agents and regulations arrive.

Outcome — an operating model, not a one-off project

Background

Twenty years of production engineering, now applied to AI.

  1. 2024 — present

    Amalfi AI

    Founder

    Consultancy operationalizing AI GovOps, AI FinOps and AI Performance Engineering; creator of PromptClassify.

  2. 2021 — 2024

    United Airlines

    API platform, MLOps & AI traffic management — mobile cloud transformation

    Moved mobile services to ECS Fargate, introduced Kong Gateway and Kong Mesh with full Terraform automation and custom Lua plugins, built OpenTelemetry tracing to Dynatrace, and introduced Kong AI Proxy for AI traffic.

  3. 2020 — 2021

    Nomi Health

    DevOps & automation architect

    AWS infrastructure automation across business units; GitHub Actions and serverless delivery.

  4. 2018 — 2020

    Goldman Sachs

    DevOps & automation architect — Apple Card

    Multi-account, multi-region AWS infrastructure for Apple Card transactions; AMI pipelines with CIS security testing; HashiCorp Vault and Consul.

  5. 2016 — 2018

    Liberty Mutual · GE Digital

    DevOps & automation architect

    Compliance as code, secrets management and configuration automation for mission-critical cloud systems.

Patents

  • Remote and real-time network and HTTP monitoring with a real-time predictive end-user satisfaction indicatorUS 8,972,569 B1 · issued 2015 · sole inventor
  • System and method for clustering servers for performance and load balancingUS 6,965,938 B1 · issued 2005 · co-inventor, IBM
  • Method and computer program for web site performance monitoring and testing by variable simultaneous angulation

Publications

  • Creating an Effective DevOps Strategy
  • Creating ChatOps & Incident Response Solutions

Education & certifications

  • Lafayette College — B.S., Electrical Engineering
  • Advanced Dynatrace Certification
  • Service Mesh Certification
Start the Conversation

Book a 30-minute AI operations review.

Bring one question — what your AI really costs, whether it would pass an audit, why it's slow, or whether an agent is ready for production. You'll leave with a clear next step, whether or not we work together.

or write directly: johndesp@amalfi.ai