SENIOR SITE RELIABILITY ENGINEER

Company: NVIDIA
Location: Santa Clara
Posted on: October 28, 2024

Job Description:

Joining NVIDIA's AI Efficiency Team means contributing to the infrastructure that powers our innovative AI research. This team focuses on optimizing efficiency and resiliency of AI workloads, as well as developing scalable AI and Data infrastructure tools and services. Our objective is to deliver a stable, scalable environment for AI researchers, providing them with the necessary resources and scale to foster innovation. We are seeking a Senior Site Reliability Engineer (SRE) to join our team. You'll be instrumental in designing, building, and maintaining cloud services that enable large-scale AI training and inferencing. The responsibilities include implementing software and systems engineering practices to ensure high efficiency and availability of the platform, as well as applying SRE principles to improve production systems and optimize service SLOs. Additionally, collaboration with our customers to plan implement changes to the existing system, while monitoring capacity, latency, and performance is part of the role.As a Senior SRE at NVIDIA, you will have the opportunity to work on innovative technologies that power the future of AI and data science, and be part of a dynamic and supportive team that values learning and growth. The role provides the autonomy to work on meaningful projects with the support and mentorship needed to succeed, and contributes to a culture of blameless postmortems, iterative improvement, and risk-taking. If you are seeking an exciting and rewarding career that makes a difference, we invite you to apply now!What you'll be doing:Develop software solutions to ensure reliability and operability of large-scale systems supporting machine-critical use cases.Gain a deep understanding of our system operations, scalability, interactions, and failures to identify improvement opportunities and risks.Create tools and automation to reduce operational overhead and eliminate manual tasks.Establish frameworks, processes, and standard methodologies to enhance operational maturity, team efficiency, and accelerate innovation.Define meaningful and actionable reliability metrics to track and improve system and service reliability.Oversee capacity and performance management to facilitate infrastructure scaling across public and private clouds globally.Build tools to improve our service observability for faster issue resolution.Practice sustainable incident response and blameless postmortemsSkilled in problem-solving, root cause analysis, and optimization.What we need to see:Minimum of 8 years of experience in SRE, Cloud platforms, or DevOps with large-scale microservices in production environments.Bachelor's degree or equivalent experience.Strong understanding of SRE principles, including error budgets, SLOs, and SLAs.Experience with AI training and inferencing and data infrastructure services.Expertise in building and operating large-scale observability platforms for monitoring and logging (e.g., ELK, Prometheus, Loki).Proficiency in programming languages such as Python, Go, script languagesHands-on experience with scaling distributed systems in public, private, or hybrid cloud environments.Experience in deploying, supporting, and supervising services, platforms, and application stacks.Knowledge of CI/CD systems, such as GitLab and Familiarity with Infrastructure as Code (IaC) methodologies and tools.Excellent communication and collaboration skills, and a culture of diversity, intellectual curiosity, problem solving, and openness are essential.Ways to stand out from the crowd:Extensive experience in Slurm workload manager and K8sGood understanding on DL frameworks, orchestrators like PyTorch, TensorFlow, JAX, and RayStrong background in software design and development.Experience operating large-scale distributed systems with strong SLAs.Extensive experience in operating data platforms with p roficiency in incident, change, and problem management processes.NVIDIA leads the way in groundbreaking developments in Artificial Intelligence, High-Performance Computing, and Visualization. The GPU, our invention, serves as the visual cortex of modern computers and is at the heart of our products and services. Our work opens up new universes to explore, enables amazing creativity and discovery, and powers what were once science fiction inventions, from artificial intelligence to autonomous cars. NVIDIA is looking for exceptional people like you to help us accelerate the next wave of artificial intelligence.The base salary range is 180,000 USD - 339,250 USD. Your base salary will be determined based on your location, experience, and the pay of employees in similar positions.You will also be eligible for equity and benefits (https://www.nvidia.com/en-us/benefits/) . NVIDIA accepts applications on an ongoing basis.NVIDIA is committed to fostering a diverse work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

Keywords: NVIDIA, North Highlands , SENIOR SITE RELIABILITY ENGINEER, Professions , Santa Clara, California

Click here to apply!

Didn't find what you're looking for? Search again!

Let Santa Clara recruiters find you. Post your resume for free!

Get Santa Clara Professions jobs via email.

View more North Highlands Professions jobs

Other Professions Jobs

MDM Architect- On site
Description: This on site position is open to any qualified applicant in San Francisco, CA. Please note, this role is not able to offer visa transfer or sponsorship now or in the future Practice - AIA - Artificial (more...)
Company: Cognizant Technology Solutions
Location: San Francisco
Posted on: 10/20/2024

Contracts & Grants Analyst 2
Description: Under the general direction of the CHASS Contracts and Grants Unit Supervisor, administer and coordinate complex fiscal budgeting and administration to enable CHASS departments to maintain integrity and (more...)
Company: University of California - Riverside
Location: Oakland
Posted on: 10/20/2024

BizOps Analyst San Francisco - Full time
Description: About usOur mission isto become the de facto way people learn foreign languages. We begin by teaching the next billion people English and Spanish.English is the global language of business, culture, and (more...)
Company: Speak
Location: San Francisco
Posted on: 10/20/2024

Salary in North Highlands, California Area | More details for North Highlands, California Jobs |Salary

CDL-A Drivers: Average $92,000 year, Home Weekly, Dedicated, No Touch Freight
Description: br br br CDL-A Truck Driver Jobs Drivers average 92,000 or more annually Guaranteed Weekly Pay - 99 No-Touch Freight - Amazing Benefits Are you getting the very best your carrier has to offer (more...)
Company: Marten Transport
Location: San Francisco
Posted on: 10/20/2024

Salesforce Business Analyst
Description: As a Salesforce Business Analyst, you will be part of a team responsible for delivering enterprise cloud technology solutions. Our Salesforce Business Analysts wear many hats on their projects and are (more...)
Company: Highering LLC
Location: San Francisco
Posted on: 10/20/2024

CONSTRUCTION FOREMAN
Description: Construction ForemanSalary: 110,000-125,000Benefits: Yes Medical, Dental, Life, ESOP Schedule: Full time, PermanentWe are seeking an experienced Construction Foreman for a high end residential construction (more...)
Company: Level Recruiting
Location: San Francisco
Posted on: 10/20/2024

Truck Driver (Sacramento, CA)
Description: The J.R. Simplot Company is a diverse, privately held global food and agriculture company headquartered in Boise, Idaho. We are a true farm-to-table company with an integrated portfolio including food (more...)
Company: Disability Solutions
Location: Sacramento
Posted on: 10/20/2024

Diesel Technician/Mechanic - Roadside Assistance
Description: Diesel Technician/Mechanic - Roadside Assistance 53 Morrison Ave., Sacramento, CA 95838 Position Summary: This diesel technician/mechanic position at Penske is focused on providing top service to our (more...)
Company: Penske Logistics
Location: Oakley
Posted on: 10/20/2024

CDL A Flatbed Truck Driver
Description: br CDL A Flatbed Truck Driver br br OUR OTR FLATBED TRUCK DRIVERS EARN UP TO 90,000 PER YEAR br br When you work for TP Trucking Logistics, you're treated like family. We make sure our drivers (more...)
Company: TP Trucking & Logistics
Location: San Francisco
Posted on: 10/20/2024

CYBER WARFARE TECHNICIAN
Description: Enlisted Sailors in the Navy Cryptology community analyze encrypted electronic communications, jam enemy radar signals, decipher information in foreign languages and maintain state-of-the-art equipment (more...)
Company: Navy
Location: Hercules
Posted on: 10/20/2024

Loading more jobs...

SENIOR SITE RELIABILITY ENGINEER

Didn't find what you're looking for? Search again!

Other Professions Jobs

Log In or Create An Account