Back to Jobs
crosscheckstaffing.com
Apply Now
Site Reliability Engineer
San Francisco, CA$165k–$250k/yrDirect Hire
You're applying for
Site Reliability Engineer
San Francisco, CA
$165k–$250k/yr
Direct Hire
About the Role
The Role
THE ROLE
As a Site Reliability Engineer, you will own the stability, scalability, and performance of a mission-critical infrastructure platform. You will be instrumental in building the reliability function from the ground up, ensuring that complex, high-availability systems remain performant as the company scales.
WHAT YOU'LL DO
Design, deploy, and maintain robust cloud infrastructure to support high-scale AI agent workloads.
Establish observability standards, implementing monitoring and alerting tools to gain deep visibility into system performance.
Lead incident management efforts, identifying root causes and implementing automation to prevent recurring issues.
Build and maintain CI/CD pipelines to streamline deployment workflows and improve developer velocity.
Automate manual operational tasks to minimize toil and maximize system reliability.
Collaborate closely with the engineering team to optimize system architecture for high availability and low latency.
WHAT WE'RE LOOKING FOR
Proven experience in managing cloud infrastructure at scale.
Deep expertise in incident management and system troubleshooting in high-availability environments.
Proficiency with infrastructure-as-code and configuration management tools.
Strong scripting and automation skills (e.g., Go, Python, or Bash).
Practical experience with modern monitoring and observability stacks (e.g., Prometheus, Grafana, Datadog).
Understanding of distributed systems and how to optimize for performance.
Prior experience working in high-growth, tech-first engineering teams.
Key Skills
Cloud InfrastructureMonitoring ToolsIncident ManagementAutomationDevOps
48-hour response
Bennett reviews every application personally and responds within 2 business days if there's a fit.