All systems operationalv16.years
Portrait of TJ Hoth
online

$ whoami

TJ Hoth

Senior Manager of Release Engineering

SRE leader with 16 years in tech. Incident response, automation, and teams that ship reliably.

$ cat about.md

About

Senior Site Reliability Engineering leader with 16 years of technology experience, including over a decade in SRE and operations leadership. I’ve directed company-wide incident response, automated processes to reduce toil, and helped drive service availability toward 99.99%.

I’m skilled at building and mentoring high-performing teams, transforming organizations toward modern SRE practices, and delivering reliable systems in complex socio-technical environments — the kind where the runbook, the org chart, and the on-call rotation all matter equally.

Currently at Early Warning, I lead Release Engineering and DevOps — transforming how we plan, intake, and communicate releases. Before that, I spent eight years at Workday climbing from operations engineer to SRE manager and senior IC, with a lot of major incidents, postmortems, and automation wins along the way.

Off the clock I run a Talos/Kubernetes homelab on GitOps, write at nerd.dad, and build tools like observability agents and cluster visualizers — because the best leaders still read diffs.

$ git log --oneline -- career

Deploy History

Career changelog — each role a release, hopefully with fewer breaking changes over time.

  1. Senior Manager of Release Engineering

    Early Warning — PAZE

    commit 0007 · deployed

    • Lead Release Engineering and DevOps teams, driving cross-functional delivery and operational excellence
    • Transformed Release Engineering to Scrum — improved collaboration, delivery planning, and execution
    • Built a centralized work intake portal that automated request management and reduced ad hoc interruptions
    • Partnered with PMO leadership to align OKRs and Program Increment initiatives with engineering execution in Jira
    • Developed Jira and Microsoft Teams automation for real-time deployment and release communications
    • leadership
    • release-engineering
    • devops
    • agile
  2. Senior Site Reliability Engineer

    Workday

    commit 0006 · deployed

    • Designed automated documentation systems with MkDocs and custom Python modules, reducing manual upkeep
    • Built integrations across multiple tools via APIs to surface incident data more effectively
    • Served as primary escalation point for high-impact incidents, driving resolution of business-critical outages
    • sre
    • python
    • incident-response
    • automation
  3. Manager of Site Reliability Engineering

    Workday

    commit 0005 · deployed

    • Directed company-wide major incident response, coordinating cross-functional teams and reducing MTTR
    • Designed and implemented automated incident management pipelines, streamlining escalation and triage
    • Mentored and developed SREs, fostering a culture of reliability, innovation, and proactive problem-solving
    • Partnered with engineering and product owners to improve system resilience and eliminate toil
    • leadership
    • sre
    • incident-response
    • automation
  4. Associate Manager of Site Reliability Engineering

    Workday

    commit 0004 · deployed

    • Led organizational transformation into a modern SRE practice with a focus on reliability engineering
    • Empowered engineers through training and process improvements, reducing deployment risks and manual work
    • leadership
    • sre
    • transformation
  5. Operations Team Lead

    Workday

    commit 0003 · deployed

    • Served as lead during major incidents, ensuring rapid resolution and clear communication
    • Streamlined incident response processes, reducing time-to-detect and time-to-resolve
    • Mentored junior engineers, elevating team technical and operational capabilities
    • operations
    • incident-response
    • leadership
  6. Operations Engineer

    Workday

    commit 0002 · deployed

    • Triaged and resolved production incidents across customer-facing environments
    • Built Slack automations to reduce administrative toil and accelerate response
    • Contributed to deployment reliability through reviews and continuous improvements
    • operations
    • automation
    • incident-response
  7. Network Operations Center Engineer

    Robert Half

    commit 0001 · deployed

    • Monitored and ensured network availability for field offices nationwide
    • Coordinated with ISPs and vendors to troubleshoot and resolve outages
    • Managed backup systems and performed datacenter tape rotations
    • noc
    • networking
    • operations

$ kubectl get highlights -A

Highlights

Projects, writing, and OSS — press / to filter.

  • Blog, homelab docs, and runbooks — SRE writing, GitOps manifests live elsewhere, prose lives here.

    • blog
    • homelab
    • sre
    • mkdocs
  • Talos Linux, Flux, Longhorn, and a full observability stack — production patterns at home scale.

    • kubernetes
    • gitops
    • talos
    • flux
  • Godot-powered 3D homelab visualizer with an agent bridge — because kubectl get pods isn't immersive enough.

    • godot
    • kubernetes
    • oss
  • Notes from the field — what SRE looks like when AI, platform teams, and pager duty all show up to the same meeting.

    • sre
    • talks
    • writing

$ cat /etc/stack.conf

Skills & Stack

status: currently_optimizing

  • release engineering & DevOps delivery
  • Scrum transformation & PI planning
  • Jira / Teams release automation
  • AI-assisted tooling

Platform & Delivery

  • Kubernetes
  • Helm
  • Jenkins
  • Git / GitHub Actions
  • CI/CD
  • Product Deployment
  • Bash / Shell Scripting

Reliability & Observability

  • Site Reliability Engineering
  • Incident Response
  • Grafana
  • BigPanda
  • Jira / JSM
  • 99.99% Availability Targets
  • Toil Reduction & Automation

Leadership & Practice

  • Engineering Management
  • Technical Product Management
  • Agile / Scrum
  • Mentorship
  • Python
  • Effective Communication
  • Critical Thinking

$ ./homelab.sh --inspect

Homelab

Talos, Flux, and production-style patterns at home scale — live status when the cluster is reachable.Full docs on nerd.dad ↗

$ git log --oneline -5 -- clusters/main

  • # waiting for GitOps feed…