Skip to content

Limited availability — one build per quarter

AI agents that run your operations — not just your demos.

I design and ship production AI agents for incident response, observability, and internal knowledge. Four of them run today inside a Fortune 500 pharmaceutical enterprise — on a multi-application Kubernetes estate, behind an enterprise incident queue. One cuts mean-time-to-acknowledge by roughly half.

30 minutes · video call · no pitch

4

agents live in enterprise production

~50%

MTTA reduction on live incidents

10

production systems shipped

6

MCP servers & integrations built

The short version

Most AI projects die between the demo and the on-call rotation.

A working prototype proves a model can do something once. It says nothing about what happens at 3am when the ticket is ambiguous, the logs are noisy, and the query times out on high cardinality.

That gap is the whole job. Closing it means meeting your systems where they already are — so I write the tool layer myself, including the MCP servers that let an agent reach Grafana, ServiceNow or Confluence reliably rather than approximately. That layer is usually the difference between an agent that demos and one that gets deployed.

Give me the problem statement. I build the system that solves it reliably, in production, where it creates measurable value.

Selected work

Shipped, not sketched.

Three systems in production. Each one shown the way an engineer would want to see it: the problem, the architecture, the guardrail, and what actually changed.

How I build

Capability is easy. Restraint is the skill.

Most agent projects ask how much autonomy they can grant.
I start from what the system must never be allowed to do.

Human-in-the-loop by default

Every agent I have put into production retrieves, reasons and recommends — a person decides. The incident copilot posts a recommended fix onto the ticket; the SRE acts on it. That boundary is why these systems were allowed anywhere near production in the first place.

Scoped, non-destructive access

My DevOps diagnostic agent has SSH into a production EC2 host with a hard allowlist of read-only commands. It can tell you exactly what is broken and what the fix should be. It cannot apply it. Capability is easy; deciding what an agent must never be able to do is the engineering.

Built on the stack you already run

No rip-and-replace. I wrap what you have — Grafana, Loki, Prometheus, Jaeger, ServiceNow, Confluence, GitHub Actions — behind MCP servers and tool interfaces. When I needed a Grafana-stack MCP server that didn't exist, I wrote one.

Failure modes before features

High-cardinality PromQL that times out. Ambiguous tickets that retrieve the wrong runbook. Log noise that poisons a summary. I design for these first, because they are what actually decides whether an agent is trusted six months in.

Working together

Engagements

Two ways in. Both start with the same 30-minute call about the problem you actually have.

Most engagements

4–8 weeks

Fixed-scope agent build

One agent, scoped to one painful workflow, shipped into your stack and running in production.

Who it is for

Platform, SRE and DevOps teams who know exactly which manual loop is burning their week.

  • Discovery: the workflow, the data sources, the failure modes, the success metric
  • Architecture with an explicit guardrail and human-in-the-loop design
  • Tool/MCP layer over your existing systems — no platform migration
  • Evaluation harness so you can prove it works before you trust it

Also available

Ongoing, no minimum

Hourly consulting

Architecture review, unblocking, and a second pair of eyes from someone who has run agents in production.

Who it is for

Teams already building who need to pressure-test a design, fix an agent that works in demo and fails in prod, or decide whether to build at all.

  • Agent and RAG architecture review
  • Evaluation strategy — how you will actually know it works
  • Guardrail, permission and human-in-the-loop design
  • Model, cost and latency trade-off decisions

The process

How a build runs

01

Problem statement

A 30-minute call. You describe the manual loop that hurts. I tell you honestly whether an agent is the right answer — sometimes it is a script and I will say so.

02

Scope and success metric

We agree on one workflow, one measurable outcome, and what the agent is explicitly not allowed to do. Fixed price from here.

03

Build against your stack

Tool layer over your existing systems, evaluation harness alongside it. You see working software early, not a slide deck at the end.

04

Ship and hand over

Deployed in your environment, observable, documented, owned by your team. Measured against the number we agreed in step two.

What you get

A running system with a measured before and after — deployed in your environment, observable, documented, and owned by your team. Not a proof of concept that dies in a branch.

Limited availability

Tell me the problem.
I’ll tell you if an agent is the answer.

A 30-minute call, no pitch. Describe the manual loop that hurts and I’ll give you a straight read on whether this is worth building — including when the honest answer is a script, not an agent.