Get the latest updates on AI-powered hiring, career growth, and technical deep-dives delivered to your inbox.
OneTrust
The ChallengeWe're looking for a Senior Software Engineer that will report to the Development Manager / R&D Head. In this role you will be part of the R&D Team that works on mission‑critical applications.
Your MissionEngage and partner with various Engineering, Operations, and Product teams to design, deliver, and maintain a highly available and performant application platform.
Build and implement application observability and platform monitoring tools to continuously improve the customer experience.
Eliminate toil by automating processes, tuning alerts, and improving code where it is most needed.
Frequently evaluate new ideas and trends to identify potentially useful tools and techniques.
Collaborate with different functional groups to identify gaps, prioritize, and resolve issues.
Define, implement, and maintain SLIs and SLOs aligned with customer experience.
Design and instrument SLIs such as latency, error rates, and availability across critical services.
Manage and enforce error budgets to balance system reliability with product feature velocity.
Improve alert quality by reducing noise and focusing on actionable, high‑signal alerts.
Embed with product teams to review architectures and catch reliability risks early.
Share your knowledge and experience with the Engineering organization.
Share your findings with technical leadership and senior management.
Build scripts in Python, Bash, Java, or Ruby for operational automation and incident response.
Your Experience IncludesBachelor's degree in computer science, engineering, or related technical or business field.4+ years of application development experience with Java or other equivalent language.
Experience with Spring environment.
Experience in cloud‑based infrastructure (Azure, AWS, GCP, etc.).
Experience with factors that affect software application performance, including database performance, network performance, CPU utilization, JVM tuning, memory analysis, thread management, and query performance.
Knowledge of centralizing logging, metrics, dashboards, and alerting.
Good awareness of databases (SQL/NoSQL).
Hands‑on experience with observability tools (Datadog, Prometheus, Grafana, etc.).
Knowledge with CI/CD pipelines and infrastructure‑as‑code (Terraform, Helm, Jenkins, GitLab).
Build and operate AI‑assisted incident response systems (root cause analysis, log summarization, anomaly triage).
Develop or integrate LLM‑based tools to reduce MTTR and improve alert quality.
Apply machine learning techniques for anomaly detection, capacity prediction, or failure pattern analysis.
Experience deploying AI systems in production (not just experimentation).
Knowledge with vector databases, embeddings, or RAG architectures for operational intelligence.
Well‑developed insight of prompt engineering and evaluation of LLM outputs in the reliability workflow.
Kubernetes and container orchestration (EKS/AKS/GKE).
Experience with distributed systems at scale.
Familiarity with service meshes and microservices architectures.
Nice to HaveExperience with chaos engineering tools (Gremlin, Chaos Monkey).
Background in product‑facing services with high traffic scale.
Understand how to use incident management platforms such as PagerDuty and DataDog.
As an employee, you will receive comprehensive healthcare coverage, flexible PTO, equity RSUs, annual performance bonus opportunities, retirement account support, 14+ weeks of paid parental leave, career development opportunities, and company‑paid privacy certification exam fees.
OneTrust provides equal employment opportunities to all employees and applicants for employment and prohibits discrimination and harassment of any type without regard to race, color, religion, age, sex, national origin, disability status, genetics, protected veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by local laws.#J-18808-Ljbffr
(EEO), Establish Priorities, Failure Analysis, GCP (Good Clinical Practices), Genetics, Health Plan, High Availability Software, Incident Management, Incident Response, Java, Jenkins, Machine Learning, Memory Hardware, Metrics, Microservices, Microsoft Windows Azure, Network Performance/Analysis, NoSQL, Pattern Analysis, Performance Management, Problem Solving Skills, Process Improvement, Product Design, Product Reviews, Production Systems, Python Programming/Scripting Language, Quality Management, Reporting Dashboards, Research & Development (R&D), Root Cause Analysis, Ruby, SQL (Structured Query Language), Scripting (Scripting Languages), Software Development, Software Engineering, Systems Reliability, Technical Leadership, Trend Analysis
OneTrust
Most resumes get rejected by the ATS before a human sees them. Check yours free in 30 seconds.
Matched to your profile
We surface this role because it matches profiles like yours, not because we vet the employer. Always confirm the pay, location, and remote details on OneTrust's official site before you apply.
Most large employers screen resumes with software before a recruiter ever sees them. Check yours against this role in seconds. Free, no sign-up.
See the exact keywords from this posting your resume is missing, with an instant ATS score.
Open free toolUpload your CV for an instant 0-100 score and the fixes recruiters and ATS look for.
Open free toolGenerate a clean, single-column resume that parses correctly and gets past the filters.
Open free toolOther live openings in the same field. All are still accepting applications.
RoShay Services
Denver, CO
STR
Washington, DC
PwC
Philadelphia, PA
GEICO
Dallas, TX
Harvey
San Francisco, CA
Alineops
Atlanta, GA