Profile

MLOps / AI Infrastructure Engineer and Senior DevOps / Site Reliability Engineer — I take AI and machine-learning models and make them reliable, secure, scalable and cost-efficient in production. I bring production-grade CI/CD, Kubernetes, observability and security discipline to the deployment, serving and operation of generative-AI and LLM systems. This is backed by 15+ years as a systems/infrastructure engineer and 8+ years leading DevOps across AWS, Azure and GCP, delivering enterprise-wide platforms — build, CI/CD, requirement analysis, design, development, testing and release — for the retail, banking, telecom, financial and regulated clinical industries, with proven leadership and mentoring. I believe in infrastructure as code and a positive, programmatic approach to systems administration; I build platforms that are reliable, scalable, cost-optimized and compliant, and I ramp into new technical domains fast.

Experience

MLOps / AI Infrastructure Engineer Independent

Sep 2025 – Present
Independent · Remote
  • Full-time, self-directed transition into AI/ML engineering — LLM serving, RAG, agents and LLMOps — built on 15+ years of production infrastructure and cloud experience.
  • Evaluating and serving open and hosted models (Llama-class, Anthropic, OpenAI): inference with vLLM / TGI / Ollama, quantization (GGUF / AWQ / GPTQ), and GPU throughput-vs-latency tuning.
  • Building tool-calling agents (function calling and the Model Context Protocol / MCP) and RAG pipelines over vector stores, with automated evals (promptfoo) and LLM observability (Langfuse).
  • Applying SRE discipline to the AI lifecycle: CI/CD for models, autoscaling GPU workloads on Kubernetes, cost/latency controls, and AI security & guardrails (OWASP LLM Top 10, prompt-injection defense).

Senior DevOps Engineer

Apr 2023 – Aug 2025
Clario SRL · Costa Rica
  • Provided infrastructure and support for software developers to rapidly iterate on products and ship high-quality releases: automated builds and testing, continuous integration, software releases and system deployment.
  • Designed and implemented a toolset that simplifies provisioning and support of a large cluster environment.
  • Owned end-to-end incident response through Dynatrace observability — triaged performance issues and alerts across the Kubernetes fleet, isolated root causes from traces, metrics and logs, and drove troubleshooting to fast, high-quality resolution.
  • Reviewed performance statistics, query execution/explain plans and GitLab CI/CD pipeline runs; recommended tuning changes and owned identifying bottlenecks and improving performance across all systems.
  • Configured tooling, templates and scripts for the ARB-approved “dynamic pipeline” — a centralized, reusable GitLab CI/CD framework distributed to every microservice for deployments to Kubernetes (Rancher / EKS). It standardizes deploys via reusable Helm charts (Helm Hub), with canary releases, pre/post-deployment tests and integrated DAST security scanning — minimizing errors and enabling fast, high-quality deployments across teams.
  • Authored and maintained detailed technical documentation, knowledge articles and runbooks (GitLab wikis, DevOps guide, ConfigHub); enforced security and compliance standards on maintained IaC within a regulated clinical (GxP) environment with strict validation-vs-production controls.

Lead DevOps / Platform Engineer NDA

Sep 2019 – Mar 2023
Confidential (under NDA)
  • Role split ≈ 60% support & enablement · 30% delivery & execution · 10% research for an enterprise software organization.
  • Built cloud CI/CD pipelines and automated build, test, integration and deployment for enterprise software; designed toolsets that simplified provisioning at cluster scale.
  • Maintained health of cloud-based production environments through monitoring, alerting and daily administration; responded to performance issues identified by alerts and reported incidents.
  • Worked directly with developers to execute software releases, configuration updates and release requirements.
  • Tuned performance and identified bottlenecks across all systems; wrote scripts and technical knowledge articles, and continuously found better ways for the team to work.
  • Researched and implemented automation and integration best practices: source control, CI, infrastructure automation, deployment automation, container concepts, orchestration and cloud.

Principal DevOps Architect

Nov 2018 – Sep 2019
Equifax Inc. · Corporate Services
  • Managed multi-cloud (AWS/Azure/GCP) infrastructure with Chef/Ansible; part of the team that built the PCI-compliant cloud-based Java CI/CD pipeline for financial tools.
  • Deep AWS work: VPC, EC2, S3, ELB, Auto Scaling Groups, EBS, RDS, IAM, CloudFormation, Route 53, CloudWatch, CloudFront, CloudTrail — security groups, network ACLs, internet gateways and route tables to build secure zones in public cloud.
  • Created and configured elastic load balancers and auto-scaling for cost-efficient, fault-tolerant, highly available environments; S3 lifecycle policies to archive infrequently-accessed data; RDS with snapshot/AMI backup strategies.
  • Launched EC2 fleets from Linux, Ubuntu, RHEL and Windows AMIs with bootstrap shell scripts; enforced IAM roles/users/groups with MFA across the account.
  • Wrote CloudFormation templates (custom VPCs, subnets, NAT) and Terraform IaC for staging and production; worked across Docker Engine, Hub, Machine, Compose and Registry.
  • Architected across vendors — Azure (VMs, VM Scale Sets, AutoScaling, Block Blob, VNet, VPN Gateway, Traffic Manager, Application Gateway, SQL Database) and GCP (Compute Engine, Autoscaling, Cloud Storage, Persistent Disk, Spanner, Cloud IAM, Load Balancer, VPC).
  • Deployed monitoring/APM with DataDog, AppDynamics, Nagios and CloudWatch; led development of highly automated, self-managed runtime environments.
  • Defined the Costa Rica business unit's cloud strategy, migration and application consolidation; ran Lean/Agile (Scrum, Spotify model) delivery with Confluence/Jira/Bitbucket/Bamboo; installed and customized JIRA; automated in Python, Bash and PowerShell (incl. WMI-based low-bandwidth reporting) and trained peers on PowerShell practices.

Senior DevOps Engineer

Jan 2018 – Nov 2018
Snap Finance · USA / UK
  • Managed AWS infrastructure with automation and DevOps orchestration tools (Chef/Ansible); part of the team building the PCI-compliant cloud-based Java CI/CD pipeline for financial tools.
  • Hands-on across VPC, EC2, S3, ELB, ASG, EBS, RDS, IAM, CloudFormation, Route 53, CloudWatch, CloudFront, CloudTrail; designed secure zones with security groups, network ACLs, internet gateways and route tables.
  • Built load-balanced, auto-scaled, cost-efficient HA environments; S3 lifecycle archiving; EBS volumes for application storage; RDS snapshots, volume backups and launch-configuration images.
  • Authored CloudFormation JSON templates (custom VPC, subnets, NAT) and Terraform templates for staging and production; Route 53 DNS for highly available applications; production monitoring/alerting via CloudWatch.
  • Worked across Docker components (Engine, Hub, Machine, Compose, Registry); installed and customized JIRA for workflow and user/group management; wrote Python, Bash and PowerShell automation.

DevOps / SRE Consultant

Jan 2017 – Dec 2017
iTelecom Costa Rica · Own Company
  • Migrated 2 complete on-premise infrastructure stacks to the cloud, end to end.
  • Advised on the design, build and rollout of cloud-based DevOps frameworks and tooling supporting the strategic migration of key, high-visibility customer-facing products.
  • Provided technical guidance, knowledge transfer and mentorship to engineering peers; led technical staff responsibilities and managed relationships with multiple suppliers and internal teams.
  • Designed, implemented and managed systems & DevOps infrastructure; supported associated technologies (VMware, core storage, networking) and troubleshot networks, compute, virtualization, telecom circuits and datacenter issues.
  • Ran delivery on Agile practices: Scrum / Lean / Kanban / CI-CD / DevOps.

Cloud Project Manager

Jun 2016 – Dec 2016
Hewlett Packard Enterprise
  • Managed all phases of cross-functional cloud implementation projects to time and budget; organized and facilitated configuration-deliverable teams.
  • Developed the Master Implementation Plan with Workstream Leads and the PMO; provided milestone inputs to the Program Roadmap and oversaw complete, timely vendor deliverables.
  • Proactively identified conflict/integration issues of highest complexity — including cost/benefit analysis and resource estimating — and escalated through the project's organizational structure with analyzed options for solutions.
  • Delivered clear, relevant communications to business stakeholders on project needs and status to Workstream Leads and PMO.

15+ years of infrastructure foundation: Linux/Unix & Windows systems engineering, virtualization (VMware ESXi/vSphere) and enterprise networking since 2008 — the on-prem depth behind today's cloud, SRE and MLOps work.

Technical Knowledge

DevOps · Cloud Infrastructure · SRE

  • Deep knowledge of operational processes built on the AWS Well-Architected 5 pillars (operational excellence, security, reliability, performance efficiency, cost optimization); experienced evaluating, planning and migrating enterprise IaaS/PaaS/SaaS from private to public/hybrid cloud, with ROI analysis of as-is vs to-be environments.
  • Provisioning: Terraform, CloudFormation, ARM templates, AWS CLI, Azure PowerShell/CLI. Deployment patterns: blue/green, canary, rolling, draining. Containers & orchestration: Docker, Kubernetes.
  • Observability: CloudWatch, New Relic, DataDog, Dynatrace, PagerDuty for infrastructure and application logging, performance and monitoring; site-reliability practice grounded in infrastructure automation.
  • Databases & data stores: MySQL, PostgreSQL, Oracle DB, SQL Server, Memcached. Compliance: GDPR, PCI, HIPAA.

Automation & CI/CD

  • Excellent scripting/coding in Bash, Python, PowerShell (incl. DSC), Ruby, plus CSH/KSH/Perl; configuration management with Chef, Puppet, PowerShell DSC, SCCM — expert level in Windows and Linux environments.
  • Full SDLC on Agile (Scrum/Lean/Kanban): automated smoke tests running on every build, automated test tooling evaluation and rollout, and software written to automate an increasing number of functions — reducing downtime and standardizing environments company-wide.
  • Version control across Git, SVN, Team Foundation with Bitbucket/Bamboo pipelines; trained and mentored engineering teams on automation tooling; skilled at lightweight technical documentation (wikis, how-to guides, diagrams).

Earlier Foundations — Unix/Linux · Windows · Virtualization (on-prem era, condensed)

  • 10+ yrs Linux/Unix (Ubuntu, RHEL/CentOS, Debian) — performance analysis, LVM, security hardening, iptables/NMAP troubleshooting, and the full network-services stack (NFS, DNS, LDAP, LAMP, Samba, DHCP/PXE provisioning) · ~10 yrs enterprise Windows (Server 2008–2019, AD/GPO, Exchange, IIS, SCCM, SQL Server, PowerShell/DSC automation) · 8 yrs VMware ESXi/vSphere (HA/DRS clusters, vMotion, P2V/V2V migrations). Deep on-prem roots that inform today's cloud & SRE work — full detail on request.

AI Portfolio

Portfolio · 2025–2026

LLM Serving Platform

Serving an open model (Llama 3 8B) on vLLM behind a FastAPI endpoint; automated evals (promptfoo) and latency/token-cost observability (Langfuse); GPU pod on Kubernetes with quantization (AWQ/GGUF) and autoscaling for cost control.

Portfolio · 2025–2026

RAG Document Assistant

Q&A pipeline with LangChain/LlamaIndex: ingestion → chunking → embeddings (Hugging Face) → pgvector retrieval → grounded, cited answers; retrieval-quality and hallucination checks via promptfoo.

Portfolio · 2025–2026

Tool-Calling Agent

Agent using Anthropic/OpenAI function calling and an MCP server to orchestrate multiple tools across multi-step tasks, with input/output guardrails.

Portfolio · 2025–2026

Cloud AI Deploy

Foundation-model deployment on AWS Bedrock / Azure OpenAI / Vertex AI, provisioned with Terraform IaC; secrets, IAM and cost controls handled the SRE way.

Portfolio · 2026 · ✦ Live — you're looking at it

This Résumé Site

Terraform-provisioned static site: private S3 origin + CloudFront (Origin Access Control) + ACM DNS-validated TLS on a subdomain, with DNS kept at the registrar (zero disruption to production email) and scripted deploys via aws s3 sync + cache invalidation. Runs for pennies a month.

MADE WITH TERRAFORM

This page is infrastructure as code. Terraform provisions everything you're reading this through — ACM certificate (DNS-validated) → CloudFront with Origin Access Control → private S3 bucket — and a deploy script syncs content and invalidates the cache.

Download PDF