Site Reliability Engineer · 20+ Years Building Resilient Systems

Rafael Dihl

Reliability-focused engineer specializing in distributed systems, observability, and incident response — keeping high-QPS, member-facing services resilient at scale.

Rafael Dihl

20+ years engineering and operating large-scale systems, with 7+ years focused on the reliability, observability, and automation that keep production services healthy. I take a data-driven approach to reliability — defining SLOs, instrumenting systems with Datadog and OpenTelemetry, leading incident response and post-incident analysis, and designing for failure modes across distributed systems. Currently at Fanatics, running high-availability workloads on AWS EKS with GitOps delivery (ArgoCD) and automated, right-sized infrastructure. Previously at Meta (via Magnit), where I designed and deployed the AWS environment behind the Wolfram Alpha LLM integration, engineered for high QPS, low latency, and strict privacy compliance. Comfortable across Python, Go, and Java, and fluent in both English and Portuguese.

DealerTrack

Resilience & disaster recovery. Architected and deployed multiple enterprise HashiCorp Vault clusters on-premises with disaster recovery — configurations that directly enabled recovery from several high-visibility production outages, minimizing member-facing downtime.

Fanatics

Observability & incident response. Built infrastructure monitoring with Datadog and Terraform — custom dashboards and alerting rules that reduced incident response time — and led the FluxCD→ArgoCD migration to gain deployment traceability and fast, reliable rollbacks on AWS EKS.

Meta

Reliability at scale. Provisioned and deployed the AWS environment (Terraform) to run the Wolfram Alpha LLM as an internal service, engineered for high QPS, low latency, and strict privacy compliance across Meta traffic.

Meta

Reliability vs. cost. Applied a data-driven capacity analysis to a batch service, identified unused elastic capacity, and safely released it — reducing server usage by 60% and translating to $7M in annual savings with no loss of reliability.

Fanatics

Right-sizing & efficiency. Drove the migration of EKS workloads to AWS Graviton (arm64) and introduced Karpenter for dynamic, demand-driven node provisioning. Analysis of the 233-host fleet identified 147 eligible x86 instances, projecting ~$83K in annual savings — scaling to ~$416K/year at 5× fleet growth. Karpenter provisions nodes right-sized to actual pod requirements and automatically consolidates or terminates underutilized capacity, where fixed Managed Node Group pools would leave idle nodes running.

Charter

Reliability as a standard. Designed and delivered CI/CD pipeline libraries that standardized safe, repeatable delivery across the entire organization — spanning Kubernetes, Helm, Terraform, Vault, Datadog, and more.

02/2025 — Present
Fanatics

Senior Platform Engineer

  • Architected and maintained scalable CI/CD pipelines on AWS EKS using GitHub Actions, FluxCD, and ArgoCD, including provisioning ArgoCD clusters from scratch.
  • Led the migration from FluxCD to ArgoCD to improve deployment traceability, centralized visibility, and rollback capabilities.
  • Built and managed infrastructure monitoring with Datadog and Terraform, creating custom dashboards and alerting rules to reduce incident response time.
11/2023 — 02/2025
Magnit · Meta

Software Engineer — Contingent Worker @ Meta

  • Designed and implemented AWS infrastructure (EC2, EKS, RDS, Kafka, Terraform, Ansible, Helm) to deploy the Wolfram Alpha LLM solution for Meta traffic, ensuring high QPS, low latency, and privacy compliance.
  • Spearheaded enhancements to data and code delivery pipelines, significantly improving build efficiency and deployment velocity using Kubernetes, Golang, Python, PHP, Bash, and Terraform.
01/2021 — 11/2023
ThinkBRQ · Charter

Senior DevOps Engineer — Consultant @ Charter Communications

  • Developed and delivered Infrastructure as Code solutions enabling partner engineering teams to migrate applications to AWS.
  • Designed CI/CD pipeline libraries that standardized the use of GitLab, Kubernetes, Helm, Docker, Terraform, HashiCorp Vault, Fluentbit, Datadog, and Splunk across the organization.
07/2016 — 12/2020
ThinkBRQ · DealerTrack

Senior DevOps Engineer — Consultant @ DealerTrack

  • Provisioned and deployed multiple enterprise HashiCorp Vault clusters on-premises with disaster recovery; configurations played a key role in recovering from several high-visibility outages.
  • Led development of a scalable CI pipeline (Jenkins CasC + OctopusDeploy) stabilizing builds and deployments across 10+ environments.
  • Containerized a large legacy monolith into a 12-factor app, achieving 40% performance improvement through Docker and immutable infrastructure.
11/2013 — 06/2016
Dell

Software Development Advisor

  • Designed and implemented large-scale Java EE and SOA applications in globally distributed teams, acting as Subject Matter Expert in field services.
  • Coordinated work across small to medium teams on strategic, tactical, and maintenance projects; provided L3 production support.
04/2009 — 11/2013
ADP

Senior Software Developer

  • Developer on ADP Portal R8, one of ADP's most widely used products with 17M+ users.
  • Built features using Struts, WebSphere Portal, Application Servers, JSF, and JMS.

Reliability

SLOs & Alerting Incident Response Post-Incident Analysis On-Call Disaster Recovery Distributed Systems Capacity Planning

Languages

Java Python Go JavaScript PHP Bash HCL

Infrastructure

Terraform AWS Kubernetes Docker Helm Ansible Saltstack Rancher

CI/CD & GitOps

GitHub Actions ArgoCD FluxCD GitLab CI Jenkins CasC Octopus

Data & Messaging

Kafka Cassandra PostgreSQL MySQL MongoDB Oracle DB Redis GraphQL

Observability

Datadog New Relic Splunk OpenTelemetry CloudWatch

Platform & Security

HashiCorp Vault Consul Twingate Rundeck
2003 — 2007

Pontifícia Universidade Católica do Rio Grande do Sul

Bachelor of Information Systems (BCompSc) · Porto Alegre, Brasil

Nov 2014

Oracle Certified Associate — Java SE 7 Programmer

Oracle · Credential ID: OC1304424

Aug 2015

Certified Scrum Master

Scrum Alliance · Credential ID: 000447682

🏃

Running

Completed 3 marathons (Chicago, London, Porto Alegre) and currently training for Berlin in September 2026.

📚

Reading

Avid reader who enjoys exploring new ideas and perspectives through books.

🏖️

Beach Time

Love spending time at the beach with my family, enjoying the Florida coast.