Site Reliability Engineer · 20+ Years Building Resilient Systems
Reliability-focused engineer specializing in distributed systems, observability, and incident response — keeping high-QPS, member-facing services resilient at scale.
About
20+ years engineering and operating large-scale systems, with 7+ years focused on the reliability, observability, and automation that keep production services healthy. I take a data-driven approach to reliability — defining SLOs, instrumenting systems with Datadog and OpenTelemetry, leading incident response and post-incident analysis, and designing for failure modes across distributed systems. Currently at Fanatics, running high-availability workloads on AWS EKS with GitOps delivery (ArgoCD) and automated, right-sized infrastructure. Previously at Meta (via Magnit), where I designed and deployed the AWS environment behind the Wolfram Alpha LLM integration, engineered for high QPS, low latency, and strict privacy compliance. Comfortable across Python, Go, and Java, and fluent in both English and Portuguese.
Impact
Resilience & disaster recovery. Architected and deployed multiple enterprise HashiCorp Vault clusters on-premises with disaster recovery — configurations that directly enabled recovery from several high-visibility production outages, minimizing member-facing downtime.
Observability & incident response. Built infrastructure monitoring with Datadog and Terraform — custom dashboards and alerting rules that reduced incident response time — and led the FluxCD→ArgoCD migration to gain deployment traceability and fast, reliable rollbacks on AWS EKS.
Reliability at scale. Provisioned and deployed the AWS environment (Terraform) to run the Wolfram Alpha LLM as an internal service, engineered for high QPS, low latency, and strict privacy compliance across Meta traffic.
Reliability vs. cost. Applied a data-driven capacity analysis to a batch service, identified unused elastic capacity, and safely released it — reducing server usage by 60% and translating to $7M in annual savings with no loss of reliability.
Right-sizing & efficiency. Drove the migration of EKS workloads to AWS Graviton (arm64) and introduced Karpenter for dynamic, demand-driven node provisioning. Analysis of the 233-host fleet identified 147 eligible x86 instances, projecting ~$83K in annual savings — scaling to ~$416K/year at 5× fleet growth. Karpenter provisions nodes right-sized to actual pod requirements and automatically consolidates or terminates underutilized capacity, where fixed Managed Node Group pools would leave idle nodes running.
Reliability as a standard. Designed and delivered CI/CD pipeline libraries that standardized safe, repeatable delivery across the entire organization — spanning Kubernetes, Helm, Terraform, Vault, Datadog, and more.
Experience
Skills
Education & Certifications
Bachelor of Information Systems (BCompSc) · Porto Alegre, Brasil
Oracle · Credential ID: OC1304424
Scrum Alliance · Credential ID: 000447682
Beyond Work
Completed 3 marathons (Chicago, London, Porto Alegre) and currently training for Berlin in September 2026.
Avid reader who enjoys exploring new ideas and perspectives through books.
Love spending time at the beach with my family, enjoying the Florida coast.