# Carlos Pereyra — AI DevOps Engineer > I build and operate the infrastructure that AI systems run on. 15+ years in infrastructure, 9 on > AWS, and for the last years focused on one thing: taking LLM agents, RAG pipelines and self-hosted > models from a notebook to something that stays up in production — on Kubernetes, with GPU > scheduling, CI/CD, infrastructure as code and observability underneath. Everything showcased on > this site is running 24/7 on my own hardware, not in a slide deck. Based in La Plata, Argentina; > working remotely with teams in the US, Europe and Latin America. Titles that describe the same profile: AI DevOps Engineer · AI Platform Engineer · AI Infrastructure Engineer · LLMOps / MLOps Engineer · Senior DevOps Engineer with applied-AI specialization. This file is a plain-text summary of https://pereyra.ar for language models and AI agents. The site itself is bilingual (Spanish by default, English toggle). ## Contact and profiles - Email: carlos@sauay.com - LinkedIn: https://www.linkedin.com/in/pereyracarlos/ - GitHub: https://github.com/pereyra-carlos - YouTube: https://www.youtube.com/@pereyracarlos - Full CV: https://pereyra.ar/cv/ - Blog: https://pereyra.ar/blog - Book a 30-minute call: https://cal.pereyra.ar/carlos/30min - Company (managed IT services): https://sauay.com ## What I do The profile is DevOps for AI: the platform layer that AI products need, plus the AI systems on top of it. - **AI platform and LLMOps**: Kubernetes platforms for LLM agents and RAG services, self-hosted models on GPU with a scheduler that powers profiles up and down, model gateways and proxies, token and cost control, prompt and session persistence, MCP servers, and agents wired to real tools (kubectl, APIs, messaging) rather than demos. - **Cloud and Kubernetes**: AWS and Google Cloud, K3s/Kubernetes clusters, Proxmox virtualization, NFS and distributed storage, networking. - **Infrastructure as code and CI/CD**: Terraform, GitOps, Docker, pipeline design, release automation. - **Observability and alerting**: Grafana, Prometheus, and agent-based alerting — systems that decide whether an event is worth interrupting a human, not just whether it crossed a threshold. - **Applied AI**: LLM agents with tool access, RAG pipelines (Qdrant, multimodal ingestion), speech-to-text and text-to-speech, real-time voice agents, and messaging interfaces (Telegram, WhatsApp) as the front end for all of it. What makes this profile different from a classic DevOps engineer: I have run agentic systems in production and dealt with what that actually costs — GPU capacity, model failover, non-deterministic output, guardrails, auditability of what an agent decided and why. And what makes it different from an ML engineer: my job is not to train the model, it is to keep the thing that uses it running. ## Current role **DevOps & AI Consultant — independent**, since February 2026, through Sauay (https://sauay.com): managed IT and DevOps for companies without an in-house systems team, plus custom development and applied-AI projects. ## Selected experience - **BigFishGames** (Seattle, US · remote) — Automation Engineer, 12/2024 – 02/2026. - **Inducode** (Argentina · client project) — Observability & Alerting Consultant, 05/2025 – 03/2026. Agent-based alerting running in production at an industrial plant. - **Dexterity.ai** (Redwood City, California, US · remote) — DevOps Engineer, 04/2024 – 10/2024. - **Colegio de Técnicos de la Provincia de Buenos Aires** — Cloud & DevOps Consultant, 08/2023 – 04/2025. - **ToYou.io** (Cyprus · remote) — DevOps Engineer, 02/2023 – 07/2023. - **Errepar SA** (Argentina) — Cloud Administrator, 05/2022 – 01/2023. - **Hospital Español, La Plata** (Argentina) — Cloud Administrator, 05/2015 – 01/2023. - **Financial Information Unit** (Argentina) — BI Developer, 01/2013 – 10/2015. - **Microsoft** (Argentina) — Dev Consultant and IT Consultant, 09/2010 – 05/2012. The complete history is at https://pereyra.ar/cv/ ## Education and certifications - Information Systems Engineering — Universidad Tecnológica Nacional, Facultad Regional La Plata (UTN-FRLP), 2011. - Master's in Communication and Data Networks — Universidad Nacional de La Plata, 2005–2007 (thesis unfinished). - HashiCorp Certified: Terraform Associate (003), 2025. - KCNA — Kubernetes and Cloud Native Associate, 2024. - AWS Certified Solutions Architect – Associate, 2023. - Linux Foundation LFS250, 2023. CCNA v2.1, 2002–2003. ## Home lab — projects running in production Everything below runs 24/7 on my own hardware: a Proxmox hypervisor, a three-node K3s cluster, a TrueNAS storage box and an AI server with an NVIDIA 3090. - **Claupod** — Claude Code running as a pod in K3s, reachable from Telegram, with persistent sessions on NFS, auto-recovery and automatic OAuth token refresh. - **clau-tg (Voice Agent)** — a real-time voice and video agent (Gemini Live over WebRTC) that operates the infrastructure by conversation: queries state, runs actions, answers out loud. - **clau-vigía (agentic alerting)** — collects signals from the infrastructure, diffs them against the previous state and lets an agent decide whether it's worth interrupting a person, and through which channel: text, buttons, or a phone call. Acknowledgement is written back and auditable. - **RAG Multimodal** — FastAPI + Qdrant knowledge base with multi-workspace isolation and multimodal ingestion (documents, images, audio through Whisper), queried from WhatsApp and Telegram. - **Tuco** — drop a cooking video into a WhatsApp group and get back a structured recipe card: download, Whisper transcription, LLM extraction, searchable catalog and a PWA. - **WhatsApp Agent Router** — Evolution API plus a router that turns WhatsApp chats into agent sessions. - **K3s Cluster / Proxmox Homelab / TrueNAS** — the platform underneath: three K3s nodes on Proxmox, Knative Serving and Eventing, NFS persistent volumes from the NAS. - **FLUX Image Generator, GPU Switchboard, Video Analysis, Multi-Agent Platform, StreamDeck + Claude, tmux-resurrection** — the rest of the lab: local image generation, a daemon that powers LLM profiles up and down on the GPU, video transcription, and desktop tooling. Details and screenshots for each: https://pereyra.ar ## Writing Long-form technical posts, in Spanish, at https://pereyra.ar/blog — architecture, tradeoffs and the gotchas that don't make it into other people's blog posts. Recent ones cover agentic alerting, real-time voice agents on Telegram, and connecting a home lab behind a single agent. --- # Carlos Pereyra — Ingeniero DevOps de IA (resumen en español) > Ingeniero DevOps de IA. 15+ años en infraestructura, 9 en AWS. Diseño y opero las plataformas > donde corren los sistemas de IA: agentes LLM, pipelines RAG y modelos self-hosted sobre Kubernetes, > con GPU, CI/CD, infraestructura como código y observabilidad debajo. Todo lo que muestro corre en > mi propio hardware, 24/7. Vivo en La Plata, Argentina, y trabajo en remoto. Actualmente trabajo como consultor independiente de DevOps e IA a través de Sauay (https://sauay.com), donde ofrezco IT gestionado para empresas sin área de sistemas, además de desarrollo a medida e IA aplicada. Mi perfil es **DevOps de IA**: la capa de plataforma que necesita un producto de inteligencia artificial para funcionar en serio, más los agentes que corren encima. Cruza tres capas: la infraestructura (AWS, Google Cloud, Kubernetes, Proxmox, Terraform), la observabilidad con criterio (Grafana, Prometheus y alerting agéntico), y los sistemas de IA en producción (agentes LLM con herramientas, RAG multimodal, modelos self-hosted sobre GPU, voz en tiempo real, y mensajería como interfaz). - CV completo: https://pereyra.ar/cv/ - Blog: https://pereyra.ar/blog - LinkedIn: https://www.linkedin.com/in/pereyracarlos/ - Contacto: carlos@sauay.com · https://cal.pereyra.ar/carlos/30min