AI systems engineer / TypeScript / Mexico

I build reliable AI systems that ship.

I design bounded agent loops, MCP infrastructure, evaluation systems, and human approval surfaces. My work turns difficult AI behavior into typed, observable software a team can operate.

MCP surface
43 tools
Test suite
6,000+
Public proof
1 pkg + 3 repos

Triage

Classifies the request and selects the narrowest expert that can safely handle it.

Private system

43 tools

MCP tool surface

Private commercial system

Private system

6,000+

Automated tests

Combined private platform suite

Baseline pending

v2 contract

Reliability benchmark

Reviewed live baseline pending

Selected systems

Production architecture, with the tradeoffs exposed.

Each case focuses on the problem, system boundary, safeguards, and what I personally owned.

SAT-MCP interface01

Fiscal infrastructure

Private system

SAT-MCP

Problem
Make rule-heavy Mexican tax operations available to AI agents without giving them unbounded access.
Architecture
A typed 43-tool MCP server with resources, prompts, UI surfaces, multi-tenant credentials, and five PAC integrations behind circuit breakers.

Safety and reliability

  • Domain tool allowlists
  • Schema validation at every boundary
  • HITL for irreversible actions

Ownership: Designed and built solo in TypeScript.

DISAI_Conta interface02

Agent orchestration

Private system

DISAI_Conta

Problem
Route fiscal requests to the right specialist while controlling tool scope, cost, and unsupported catalog claims.
Architecture
A Haiku router selects one of ten domain agents. An Expert Registry injects scoped tools and live resources before the Sonnet tool loop begins.

Safety and reliability

  • Per-domain tool scopes
  • Bounded self-correcting loops
  • Langfuse and human approval surfaces

Ownership: Product architecture, orchestration, frontend, and platform integration.

03

Public evaluation

Baseline pending

AI Reliability and Evaluation Lab

Problem
Turn agent-quality claims into reproducible evidence across routing, tool use, failure recovery, cost, and latency.
Architecture
A versioned synthetic benchmark compares direct and routed agents with immutable datasets, bounded spend, raw JSONL, and CI regression gates.

Safety and reliability

  • Secretless pull request validation
  • No hand-entered results
  • Reviewed baseline promotion

Ownership: Designed as the public proof layer joining the production primitives.

Open source

Small primitives extracted from real systems.

Writing

I write about what I actually build: the non-obvious problems, the design decisions, and what breaks in production.

About Me

Background

My path to AI engineering wasn't a straight line, and that's where the edge comes from.

I started in graphic design and tattoo work, where precision and intentionality aren't optional. Then I spent two years writing deep technical research on complex systems (protocol architecture, incentive design, market structure) that trained me to read an unfamiliar system fast, find where it breaks, and explain it clearly. When I found LLMs and the Model Context Protocol, I stopped analyzing other people's infrastructure and started building my own.

Mexican, operating globally, fully bilingual: Spanish native, English professional. EST-aligned, async by default.

What I do

I build the unglamorous parts that make AI agents trustworthy in production:

  • MCP server design: full primitive set (tools, resources, prompts, elicitations, MCP-UI), SSE + HTTP Streamable transport, multi-tenant architecture.
  • Multi-tier agentic orchestration: a lightweight router classifies intent, domain-scoped agents run native tool_use loops with self-correction, and an Expert Registry enforces per-domain tool allowlists with live resource injection.
  • Observability & cost control: end-to-end Langfuse tracing on every LLM call, tool invocation, model choice, and token cost, with per-session attribution and failure triage without log digging.
  • Human-in-the-loop surfaces: one-click approval flows wired into the agent pipeline for irreversible operations.
  • Production discipline: 6,000+ tests, 91%+ coverage, CI/CD, schema validation at every boundary, conventional commits, drift guards. The hygiene that makes solo-built work safe to hand to a team.

Stack & Expertise

AI Systems & LLM Orchestration

Anthropic Claude API (tool_use, streaming, multi-turn), MCP Protocol (full primitive set), multi-agent orchestration, Expert Registry pattern, scoped tool allowlists, Langfuse tracing, HITL approval patterns.

Backend & Infrastructure

TypeScript / Node.js, Python, PostgreSQL (RLS, pgvector), SQLite, Redis, Docker, Railway, CI/CD pipelines, rate limiting, security headers, observability and monitoring. REST API design, Zod schema validation, circuit breakers, multi-tenant architecture.

Frontend & Interfaces

Next.js, React, Tailwind CSS, shadcn/ui, SSE streaming UIs, real-time data visualization (ApexCharts, Recharts). TypeScript throughout.

Automation & Workflows

n8n (self-hosted, multi-client deployments, complex workflow design), webhook orchestration, scheduled pipelines, third-party API integration. Familiar with Make and Zapier; migrated workflows to self-hosted n8n for full data ownership and agent-native integration.

Certifications & Credentials

Selected highlights below.

FEATURED
FEATURED CERTIFICATION

Model Context Protocol: Advanced Topics

AnthropicOctober 7, 2025

Claude Code in Action

Anthropic

September 25, 2025

View Certificate

Building Scalable Agentic Systems

DataCamp

2025

View Certificate

n8n Self-Hosted for Enterprises

n8n

November 2025

View Certificate

Contact

Bring me the difficult systems problem.

I am open to applied AI, AI platform, and agent systems roles. I also work with selected teams that need production MCP or evaluation infrastructure.