Site Reliability Engineer IIILocation: Onsite in Mount Laurel, NJDuration: Through 12/31/2026, extensions likelyWe are seeking a Site Reliability Engineer (SRE) to support cloud infrastructure, automation, reliability, security, and observability initiatives for AI/ML platform environments.
This role is ideal for a mid-to-senior-level engineer with strong cloud and platform engineering experience who can drive scalability, reliability, and automation within Kubernetes-based production environments.
Responsibilities
- Support Site Reliability Engineering initiatives across AI/ML platform environments.
- Deploy, maintain, and optimize cloud infrastructure across AWS and GCP.Build, manage, and maintain Infrastructure as Code (IaC) solutions using Terraform.
- Improve platform reliability, scalability, security, and operational efficiency.
- Administer and support Kubernetes and Amazon EKS environments.
- Monitor and troubleshoot system performance using observability and monitoring tools including Prometheus, Grafana, Datadog, and Elasticsearch.
- Automate operational processes, workflows, and routine administrative tasks using Python and related tooling.
- Support and enhance CI/CD pipelines and deployment automation.
- Troubleshoot complex distributed systems and production issues in highly available environments.
- Collaborate with engineering teams to improve platform performance, monitoring, and operational resiliency.
- Work with technologies including Kubernetes, Docker, AWS, GCP, EKS, Terraform, Prometheus, Grafana, Datadog, Elasticsearch, MySQL, Kafka, and Python.
Qualifications
- 4–8 years of experience in Site Reliability Engineering, DevOps, Platform Engineering, Cloud Engineering, or related disciplines.
- Strong hands-on experience with AWS cloud services.
- Hands-on Kubernetes administration and support experience.
- Expertise with Terraform and Infrastructure as Code (IaC) practices.
- Experience designing, supporting, and improving CI/CD pipelines.
- Strong observability and monitoring experience, particularly with Prometheus.
- Experience with Grafana, Datadog, Elasticsearch, or similar monitoring platforms.
- Python programming and scripting proficiency.
- Experience supporting distributed systems and large-scale, highly available production environments.
- Knowledge of algorithms, data structures, software design principles, and system troubleshooting.
- Experience with Docker containers and cloud-native platforms.
Preferred Qualifications
- Bachelor's degree in Computer Science or a related technical discipline.
- Experience supporting AI/ML platforms or infrastructure.
- Technology Doesn't Change the World, People Do.®Robert Half is the world’s first and largest specialized talent solutions firm that connects highly qualified job seekers to opportunities at great companies.
- We offer contract, temporary and permanent placement solutions for finance and accounting, technology, marketing and creative, legal, and administrative and customer support roles.
- Robert Half works to put you in the best position to succeed.
- We provide access to top jobs, competitive compensation and benefits, and free online training.
- Stay on top of every opportunity - whenever you choose - even on the go. Download the Robert Half app and get 1-tap apply, notifications of AI-matched jobs, and much more.
- All applicants applying for U.S. job openings must be legally authorized to work in the United States.
- Benefits are available to contract/temporary professionals, including medical, vision, dental, and life and disability insurance.
- Hired contract/temporary professionals are also eligible to enroll in our company 401(k) plan.
- Visit roberthalf.gobenefits.net for more information.© 2025 Robert Half.
- An Equal Opportunity Employer.
- M/F/Disability/Veterans.
- By clicking “Apply Now,” you’re agreeing to Robert Half’s Terms of Use and Privacy Notice.