Vista previa de la oferta
Software Engineer III, Site Reliability Engineering
senior · Technology / Software Development
There's nothing more exciting than being at the center of a rapidly growing field in technology and applying your engineering skills to drive innovation, stability, and resiliency across highly available platforms.
As a Software Engineer III – Site Reliability Engineering (SRE) at JPMorganChase within Asset and Wealth Management, you will combine strong backend engineering fundamentals with SRE practices to improve end-to-end service availability, scalability, and operational excellence. You will design and deliver production-grade software, build automation to reduce toil, and strengthen reliability through well-defined observability, incident response, and continuous delivery practices.
As a seasoned member of an agile team, you will decompose ambiguous reliability challenges into clear, iterative improvements—delivered through code and infrastructure-as-code.
Job Responsibilities
- Build and evolve backend services and reliability tooling with a focus on operational stability, resiliency, and maintainability.
- Design, implement, and improve reliability practices using automated continuous integration and continuous delivery (CI/CD) pipelines.
- Implement infrastructure, configuration, and network as code for the applications and platforms in your remit.
- Establish and operationalize observability—white-box/black-box monitoring, telemetry, SLI/SLO definition, alerting strategies, and error budgets—and use these signals to prioritize reliability work.
- Collaborate with software engineers, stakeholders, and production support partners to resolve complex problems and proactively address issues before they impact customers.
- Participate in Major Incident Management (MIM): lead or support triage and stakeholder communications, coordinate cross-team resolution, and drive post-incident reviews that deliver measurable resiliency improvements.
- Use enterprise-authorized AI capabilities to accelerate incident triage, troubleshooting, and post-incident analysis—validating outputs and handling operational data per sensitivity and security requirements.
- Identify recurring toil and reliability risks, prioritize reusable solutions over one-off fixes, and track outcomes tied to SLOs and stability KPIs.
- Formal training or certification in software engineering/SRE concepts and 3+ years of applied experience.
- SRE mindset and working knowledge of reliability principles, including availability, scalability, incident management, and iterative improvement.
- Strong backend engineering experience and the ability to deliver secure, high-quality production code.
- Proficiency in Python or Go (preferred), or another modern language with willingness to ramp quickly.
- Hands-on experience with system design, application development, testing, and operational stability in a large-scale distributed environment.
- Observability experience, including telemetry collection, monitoring, and SLO-based alerting.
- Strong communication skills in English, with the ability to collaborate across teams and clearly document operational and technical decisions.
- Demonstrated ability to work effectively with enterprise-authorized AI-assisted engineering tools within the work environment as a core part of modern software and reliability engineering, including the ability to validate outputs for correctness, performance, and security, and to handle inputs/outputs in accordance with data sensitivity requirements.
- Production support / on-call experience, including incident triage, mitigation, root cause analysis, and driving post-incident improvements.
- Experience in a financial institution, preferably supporting Front Office platforms and time-sensitive, high-availability workloads.
- Exposure to cloud technologies and modern platform engineering practices.
- Familiarity with modern front-end technologies (not required, but beneficial depending on platform needs).