Los Angeles, California

Tony Hayes

Platform, Infrastructure & Reliability Engineering

A steady hand for complicated systems.

I design, automate, secure, and stabilize production platforms — from infrastructure and Kubernetes to cloud and AI systems.

What I can own

One story, told across infrastructure, cloud, and platform engineering

Every role below is a different setting for the same work: take a system that has to stay up, understand what it actually depends on, and make it something a team can operate with confidence instead of dread.

Across roles, that has meant designing Kubernetes platforms, establishing reusable cloud and infrastructure-as-code patterns, automating provisioning and deployment, improving observability, and helping engineers understand and extend the systems they inherited. More recently, it has meant bringing that same discipline to AI systems — treating them as production infrastructure with the same requirements for security, observability, and clean handoffs as everything else.

  • Kubernetes
  • AWS
  • GCP
  • Azure
  • Terraform
  • Pulumi
  • Ansible
  • Docker
  • Helm
  • Kustomize
  • Argo CD
  • Flux
  • Prometheus
  • Grafana
  • OpenTelemetry
  • ELK
  • Splunk
  • Kafka
  • MQTT
  • PostgreSQL
  • Redis
  • Python
  • Go
  • Bash
  • IAM
  • Vault
  • mTLS
Selected work

Three systems, three different constraints

Details are generalized to respect employer and client confidentiality.

AI-driven IoT platform · 2019–2023

A geographically distributed Kubernetes platform for commercial-building IoT

Problem
A commercial-building IoT startup needed infrastructure for an AI-driven platform operating across physically distributed sites while its infrastructure practice was still being established.
Constraints
A growing startup, geographically distributed sites, and a product roadmap that required infrastructure standards to develop alongside the software.
Decisions
Standardize on Kubernetes as the common substrate, invest early in reusable Terraform patterns, and treat observability and deployment automation as first-class from day one rather than retrofitted later.
Implementation
Designed and operated the geographically distributed Kubernetes infrastructure end to end, established core DevOps practices across the org, built out the service mesh and observability stack, and led the migration from AWS to GCP.

Outcome: Established a repeatable operating foundation across clouds and sites, with shared deployment, observability, and infrastructure practices.

University of Southern California · 2005–2019

A cloud and automation foundation built to outlast any one project

Problem
USC needed a path from traditional infrastructure to AWS and hybrid cloud, without asking every team to relearn provisioning from scratch each time.
Constraints
A large, multi-department institution with existing systems in production, and a mandate to modernize without disrupting the departments already depending on them.
Decisions
Design cloud and Terraform patterns meant to be reused across teams rather than one-off, and pair infrastructure automation with mentoring so the patterns actually stuck.
Implementation
Helped establish AWS and hybrid-cloud adoption, designed reusable cloud and Terraform patterns, automated provisioning with Puppet, Docker, Kubernetes, and Terraform, and mentored engineers across departments as Senior Infrastructure Architect and Team Lead.

Outcome: A set of patterns and practices other teams could pick up and extend on their own, long after the initial rollout.

Independent systems work · Ongoing

A Kubernetes and GitOps platform lab, kept genuinely in production

Problem
Staying hands-on with current platform practices requires more than reading about them — it requires operating something real, continuously.
Constraints
A personal environment has to justify its complexity and remain maintainable without a dedicated operations team.
Decisions
Run an actual Kubernetes platform managed the way production infrastructure should be: declarative, GitOps-driven, and monitored — not a one-time demo left to rot.
Implementation
Operates a Kubernetes and GitOps platform with Argo CD, persistent storage, monitoring, and a set of self-hosted services, maintained as an ongoing engineering lab.

Outcome: Keeps platform skills current through ongoing operation rather than a one-time demonstration.

Experience

More than 25 years, four chapters

  1. Ongoing

    Independent Systems Work

    Personal engineering lab

    Kubernetes and GitOps platform with Argo CD, persistent storage, monitoring, and self-hosted services.

  2. 2019–2023

    Director of Infrastructure / DevOps Architect

    AI-driven commercial-building IoT startup

    Designed and operated geographically distributed Kubernetes infrastructure; established DevOps practices; led an AWS-to-GCP migration.

  3. 2005–2019

    Senior Infrastructure Architect & Team Lead

    University of Southern California

    Helped establish AWS and hybrid-cloud adoption; built reusable Terraform patterns; automated provisioning; mentored engineers.

  4. 1990s–2005

    Early career

    Systems administration & infrastructure operations

    Linux/Unix systems administration, infrastructure operations leadership, high-traffic web and streaming environments, and colocation/data-center work.

Operating principles

How I run production systems

Production readiness
Nothing ships as an afterthought. If it runs in production, it was built to run in production.
Security
IAM, secrets, and network boundaries are part of the design, not a review pass added at the end.
Clarity
A system should be as easy to explain as it is to operate. If it can’t be explained simply, it isn’t finished.
Automation
Anything done twice by hand is a candidate for automation — provisioning, deployment, recovery.
Calm incident leadership
Incidents get a clear head, a steady pace, and a plan — not noise. That calm is the job.
Understandable handoffs
Documentation and runbooks written for the next engineer, not just for the one who built it.
Contact

Let's talk about the systems you need to own

Based in Los Angeles. Open to conversations about platform, infrastructure, and reliability roles.