Job Title: Senior Site Reliability Engineer
Role Overview
Join a rapidly growing technology company that is transforming how organizations leverage data to make smarter business decisions. As the company continues its evolution toward becoming AI-native, you'll play a pivotal role in modernizing infrastructure, automating operations, and building next-generation AI-powered engineering platforms.
This is far more than a traditional Site Reliability Engineering role. You'll sit at the intersection of Site Reliability, DevOps, Infrastructure Engineering, Security, and AI, helping eliminate repetitive operational work through intelligent automation while improving platform reliability, scalability, and developer experience.
Working alongside an experienced team of senior engineers, you'll help shape the future of the company's internal engineering platform while building production-grade AI agents that directly impact how software is delivered.
Key Responsibilities
- Design, build, and maintain highly available cloud infrastructure supporting mission-critical applications.
- Build AI-powered DevOps and Site Reliability agents that automate operational workflows and reduce manual engineering effort.
- Develop and enhance internal Model Context Protocol (MCP) servers and AI agent platforms used throughout engineering.
- Champion DevSecOps best practices, emphasizing security, reliability, observability, scalability, and automation.
- Design and implement Zero Trust security architectures for cloud-native services.
- Build self-service tooling that enables engineering teams to move faster with greater autonomy.
- Develop infrastructure automation using Infrastructure as Code and modern CI/CD practices.
- Partner closely with Software Engineering, Platform Engineering, and Security teams on cross-functional initiatives.
- Improve system monitoring, capacity planning, performance analysis, and operational visibility.
- Participate in architectural discussions to ensure long-term scalability and operational excellence.
- Contribute to incident response and continuous improvement initiatives alongside platform and application teams.
- Continuously evaluate emerging AI technologies and automation opportunities to improve engineering productivity.
Education & Qualifications
- 6+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, or similar infrastructure-focused roles.
- Hands-on experience building AI agents or working with modern agentic AI frameworks.
- Experience developing or integrating internal Model Context Protocol (MCP) services.
- Strong experience with Amazon Web Services (AWS).
- Expert-level experience with Kubernetes, including Amazon EKS.
- Strong Terraform experience for infrastructure automation.
- Experience with ArgoCD or comparable GitOps deployment tools.
- Strong Python experience (or another automation language such as Go or Bash).
- Experience building and maintaining CI/CD pipelines.
- Strong Linux systems administration and containerization experience (Docker, Kubernetes).
- Experience with monitoring and observability platforms such as Prometheus, Grafana, New Relic, Loki, or similar.
- Experience working in Agile software development environments.
- Excellent communication and cross-functional collaboration skills.
Preferred Experience
- Experience deploying serverless applications using AWS SAM.
- Experience implementing Zero Trust security principles.
- Experience building security orchestration and automation at scale.
- Familiarity with AI-native engineering practices and developer productivity tooling.
- Experience balancing infrastructure reliability with developer experience and platform engineering initiatives.
Benefits And Perks
- Salary Range: $170,000 - $190,000
- 10% annual performance bonus
- Equity/stock options
- Comprehensive medical, dental, and vision coverage
- Flexible paid time off and company holidays
- Opportunities for career growth into technical leadership
- Collaborative, AI-forward engineering culture
Why Us
- Help build one of the industry's most ambitious AI-native engineering organizations.
- Work on production AI agents that directly transform how engineering teams operate.
- Join a highly collaborative team made up almost entirely of senior engineers.
- Influence architecture, automation strategy, and platform direction from day one.
- Solve complex infrastructure challenges using modern cloud-native technologies.
- Be part of a company that values innovation, ownership, and continuous learning.
- Clear opportunities to grow into Lead and future Engineering Management roles as the organization scales.
Applicants must be currently authorized to work in the United States on a full-time basis now and in the future. This position does not offer sponsorship.