Engineering

Site Reliability Engineer (SRE) Job Description

About the role

We are hiring a Site Reliability Engineer (SRE) to keep our production platform fast, available, and resilient as it scales. You will own service level objectives, build the automation and observability that catch problems before customers do, and lead incident response when things break. This role is a fit for someone who codes their way out of toil and treats reliability as a product feature, not an afterthought.

Key responsibilities

  • Define and defend SLOs and error budgets for critical services
  • Automate away toil and reduce manual operational work
  • Lead incident response and drive blameless postmortems

Responsibilities

  • Set service level objectives and track error budgets across critical services
  • Build monitoring, alerting, and dashboards in tools like Prometheus, Grafana, and Datadog
  • Automate deployments, scaling, and recovery using Terraform, Ansible, and CI/CD pipelines
  • Lead on-call rotations and coordinate incident response during outages
  • Write blameless postmortems and track corrective actions to closure
  • Run capacity planning and load testing ahead of traffic spikes
  • Harden systems for high availability with redundancy and graceful degradation
  • Tune Kubernetes clusters, autoscaling, and resource limits for cost and performance
  • Instrument services with distributed tracing and structured logging
  • Partner with product engineers to build reliability into new features from the start

Requirements

  • Bachelor's degree in computer science, engineering, or equivalent hands-on experience
  • 3 or more years in SRE, DevOps, or production systems engineering
  • Strong coding ability in a language such as Go, Python, or Rust
  • Deep experience operating Linux systems and container orchestration at scale
  • A track record of reducing incidents and improving uptime with measurable results

Make this JD your own

Generate a tailored Site Reliability Engineer (SRE) description with your company details, tone and must-have skills in seconds.

Generate with AI

Keep your hiring moving

Hiring a Site Reliability Engineer (SRE)?

Send one link. Candidates record answers on their own time and AI ranks your shortlist, no scheduling, no back-and-forth.

Frequently asked questions