NVIDIA is looking for a Field Escalation Solution Architect with experience in validation and debugging of large-scale GPU clusters focused on performance. As part of the Solution Architecture organization, we work with the most sophisticated computing hardware and software, driving the latest deep learning and machine learning breakthroughs with NVIDIA’s enterprise customers. This role offers an excellent opportunity to build your career in the rapidly growing field of deep learning while enabling the world's most successful technology companies. Primary responsibilities will be to validate and debug customer cluster performance issues, functional bottlenecks and drive customer technical engagements around NVIDIA products and technologies. Join us in this exciting endeavor!
What you’ll be doing:
A considerable part of the day-to-day job is staying up to date on pioneering High Performance Computing, Deep Learning and Machine Learning ecosystems. You'll be called on to help architect and scale high-performance, distributed AI infrastructure on-prem or in the cloud built with the latest NVIDIA GPU supercomputers for new and existing customers.
Address and resolve problems starting from the bare metal level, all the way up to the operating system, software stack, and application level.
Share knowledge with different teams by delivering demos, assisting with proof-of-concepts, and writing papers and developer blogs. By collaborating with executives and engineering, address sophisticated problems and help bring NVIDIA's premier technologies to life in the cloud and in the datacenter.
Work directly with developers and hardware architects to debug cluster performance issues, identify new requirements, cross training other account solution architects and improve workflows.
Will be engaged by the account team when extra analysis is required in debugging customer issues.
Provide additional expertise to enable the account team to be more adaptable to the customer and product engineering to get more actionable data at speed of light making them more efficient.
Building custom product demonstrations and POCs for solutions that address critical business needs of our customers.
What we need to see:
BS, MS, or PhD in Computer Science, Electrical/Computer Engineering, Physics, Mathematics, or other Engineering fields or equivalent experience.
8+ years of work-related experience in NVIDIA and/or accelerated computing technologies.
Platform-level understanding of server architecture, PCIe topology, CPUs, GPUs, NICs, Linux OS, and kernel drivers.
Networking experience, including knowledge of Ethernet, InfiniBand or other networking protocols.
Experience working with DevOps on-prem or in cloud environments, including but not limited to Docker/Containers, cloud APIs, IaaS and Data Center deployments.
SLURM, Kubernetes, and/or other job scheduler use, deployment, and debugging skills.
Deep understanding of dense data center design, including computing, storage, networking, cloud APIs, and IaaS.
Strong analytical and problem-solving skills.
Strong communication skills, both written and verbal, with the ability to collaborate and coordinate efficiently across multi-functional teams in engineering, sales, marketing, product, and program management.
Ways to stand out from the crowd:
Demonstrated CPU performance debugging experience.
Excellent customer-facing skills and background.
Platform design engineering, coding and proficient debugging skills including experience in C/C++, Linux kernel, virtualization and drivers, profilers/performance analysis tools (CPU & GPU), telemetry.
Familiarity with Grace/ARM CPU architecture, NVIDIA systems/SDKs (e.g. CUDA), NVIDIA Networking technologies (e.g., RoCE, InfiniBand), Switch interconnects through hands-on experience.
Understanding of Deep Learning and Machine Learning frameworks (TensorFlow or PyTorch), LLM, MLOps, DevOps, and workflows applying cloud technologies, using Docker/containers, Kubernetes, cloud APIs, and data center deployments, among others.
We make extensive use of conferencing tools, but occasional travel (20%) is required for a local on-site visit to customers and data science conferences.
With highly competitive salaries, a comprehensive benefits package, and an excellent engineering culture, NVIDIA is widely considered to be one of the technology industry's most desirable employers. NVIDIA has some of the most innovative people working on significant problems that define the field of ML/DL, data science, and graphics.
Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5.You will also be eligible for equity and benefits.
If an employer mentions a salary or salary range on their job, we display it as an "Employer Estimate". If a job has no salary data, Rise displays an estimate if available.
Join NVIDIA's network operations team to manage and troubleshoot large-scale datacenter and cloud networks, drive operational improvements, and respond to critical incidents across global 24/7 rotations.
NVIDIA is hiring a Senior Technical Marketing Engineer to evangelize and demonstrate scale-out AI data center networking and SuperPod architectures to hyperscalers, CSPs, and enterprise customers.
Undergraduate Engineering Support Specialist Intern at Penn State ARL assisting with materials synthesis, device fabrication, process simulation, testing, and characterization across additive and advanced manufacturing projects.
Redwood Materials is hiring a Senior BIM/VDC Engineer to drive federated Revit model setup, coordination and automation across multi-discipline infrastructure projects for battery material facilities.
Stanley Consultants is hiring a Control Systems Student Intern in Muscatine, IA to assist engineering teams with control and instrumentation documentation and learn practical skills on real-world infrastructure projects.
Experienced Siemens NX Administrator needed to own CAD environment stability and standards, automate workflows, enable engineering users, and connect CAD data into the broader digital thread for a fast-paced aerospace company.
Lead the design and operation of hybrid cloud and bare-metal GPU infrastructure to power high-performance simulation, ML, and factory automation at Atomic Industries.
Kimley-Horn is hiring Civil Engineering Analysts in Long Beach for an onsite, entry-level role supporting pre-construction design and analysis with opportunities to develop CAD, civil design, and permitting skills.
AECOM seeks an Entry-Level Surface Water Geologist to support field investigations, lab testing, mapping and geologic assessments for water-infrastructure projects in the Sacramento area.
Lead Anduril's DSP team to design and deploy advanced EW and communications signal-processing algorithms for embedded and software-defined radio platforms.
Intuitive is hiring a Senior Mechanical Manufacturing Engineer in Sunnyvale to drive design, validation, and continuous improvement of high-volume manufacturing equipment for robotic surgical instruments.
Eurofins is hiring an Automation Specialist in Indianapolis to operate and maintain liquid handling and labeling automation platforms and support sample preparation for pharmaceutical testing.
Lead system-level verification and test infrastructure strategy at Loft to enable reliable, repeatable validation of spacecraft platforms and software.
Zania is hiring a Staff Security Engineer with GRC leadership experience to embed compliance requirements into AI systems and lead automated risk and compliance assessments.
Lead a team of mechanical engineers at Moog's Military Aircraft Group to design and integrate electromechanical and electro-hydraulic flight-control test systems in a hybrid role based in Torrance, CA.
NVIDIA is a publicly traded, multinational technology company headquartered in Santa Clara, California. NVIDIA's invention of the GPU in 1999 sparked the growth of the PC gaming market, redefined computer graphics, and ignited the era of modern AI.
176 jobs