Our hiring partner, a leading direct-to-consumer lifestyle brand experiencing rapid growth, is seeking a Senior DevOps Engineer to join their expanding Platform Engineering team. This is a critical role for an experienced infrastructure professional who will be instrumental in building, scaling, and optimizing the cloud foundation that powers the company's continued expansion.
In this position, the individual will focus on designing and maintaining scalable, secure cloud infrastructure on Google Cloud Platform, managing containerized workloads in Kubernetes, and enhancing CI/CD pipelines to support a rapidly growing engineering organization.
Working at the intersection of infrastructure, application development, and security, this engineer will collaborate closely with developers, SREs, and security teams to ensure systems are reliable, performant, and well-architected for scale. The ideal candidate is a seasoned DevOps practitioner with deep cloud expertise, a passion for automation, and a strong understanding of event-driven architectures.
They thrive in collaborative environments, drive continuous improvement across infrastructure and deployment practices, and are committed to building resilient, secure systems that enable product velocity. Role
Responsibilities Cloud Infrastructure
- Design, build, and maintain scalable, highly available cloud infrastructure on Google Cloud Platform, ensuring systems are optimized for performance, cost, and reliability.
- CI/CD Optimization: Lead initiatives to scale and optimize CI/CD pipelines, ensuring efficient, automated, and secure deployments across multiple environments with a focus on velocity and quality.
- Observability & Monitoring: Implement comprehensive monitoring, logging, and alerting solutions to maintain system health, availability, and performance, enabling proactive identification and resolution of issues.
- Developer Collaboration: Partner with software engineers to integrate platform best practices into application development and deployment, fostering a culture of shared ownership and operational excellence.
- Event-Driven Architecture: Support and enhance the organization's use of event-driven services, focusing on reliability, throughput optimization, and operational excellence in asynchronous messaging.
- Troubleshooting & Resolution: Diagnose and resolve complex issues across the infrastructure, networking, and application layers, minimizing downtime and ensuring rapid restoration of services.
- Infrastructure as Code: Contribute to infrastructure as code practices using tools such as Terraform, ensuring reproducible, version-controlled, and auditable infrastructure deployments.
- Security & Compliance: Drive cloud security improvements and ensure adherence to best practices in a regulated environment, implementing robust security controls and monitoring.
- Documentation: Create and maintain clear, comprehensive documentation for infrastructure architecture, CI/CD processes, and operational best practices to enable knowledge sharing and onboarding.
- Continuous Learning: Stay current with emerging DevOps tools, cloud services, and best practices, and drive the adoption of relevant new technologies that enhance platform capabilities.
- Role
Qualifications
- Experience 5+ years of experience as a DevOps Engineer or in a similar cloud infrastructure role, with a proven track record of designing and managing production-grade systems.
- Technical Expertise Cloud Platforms: Deep expertise in designing and managing cloud infrastructure, with a strong preference for hands-on experience on Google Cloud Platform (GCP) or equivalent systems at scale.
- Container Orchestration: Practical, hands-on experience with Kubernetes cluster management, including application deployment, scaling, and troubleshooting in production environments.
- CI/CD Pipelines: Demonstrated experience building and optimizing CI/CD pipelines, with a preference for tools such as GitHub Actions, Argo
- CD, or comparable solutions.
- Event-Driven Architecture: Solid knowledge of event-driven architecture and asynchronous messaging, with a preference for hands-on experience with Kafka or similar streaming platforms.
- Scripting & Programming: Proficiency in scripting languages such as Bash, Shell, or Python for automation and tooling development.
- Infrastructure as Code: Skilled in Infrastructure as Code practices, with a preference for Terraform or comparable tools for managing cloud resources.
- Containerization: Deep understanding of containerization concepts using Docker and Kubernetes, along with cloud-native architecture principles.
- Security: Strong understanding of cloud security principles and best practices, including identity and access management, network security, and secure configuration.
- Soft Skills Excellent problem-solving abilities and a collaborative mindset, with the ability to work effectively across cross-functional teams.
- Strong communication skills, with the ability to articulate technical concepts...