NVIDIA’s Hardware Infrastructure organization is seeking a Principal Data Platform Architect. We serve and collaborate directly with NVIDIA’s rapidly growing AI, HW, and SW engineering and research teams across the company. We are looking for a technical leader to define a vision and roadmap for distributed data platform and observability systems for large-scale AI and HPC clusters and workloads and guide implementation towards this vision. You will architect systems for data collection, aggregation, enrichment, storage, retrieval, and visualization to spectacularly improve efficiency, performance, and productivity of AI and HPC workloads. You will lead technical teams to develop, deploy, and operate observability solutions for multiple compute clusters around the world.
What You’ll Be Doing:
Collaborate with AI, HW, and SW engineering and research teams to define a vision and roadmap for AI/HPC cluster observability.
Architect and lead teams to develop, test, and deploy data collectors, pipelines, visualization and retrieval services.
Define data collection and retention polices to balance network bandwidth, system load, and storage capacity costs with data analysis requirements.
Work in a diverse team to provide operational and strategic data to empower our engineers and researchers to improve performance, productivity, and efficiency.
Continuously improve quality, workloads, and processes through better observability.
What We Need to See:
Experience designing and building large scale, distributed observability systems.
Ability to collaborate with data scientists, researchers, and engineering teams to identify high value data for collection and analysis.
Experience with turning raw data into actionable reports
Experience with observability platforms such as Apache Spark, Elastic/Open Search, Grafana, Prometheus, and other similar open-source tools
Technical lead level Python, JS and Java programming experience.
Thorough understanding of databases (relational and non-relational)
Passion for improving the productivity of others
Excellent planning and interpersonal skills
Flexibility/adaptability working in a dynamic environment with changing requirements
MS (preferred) or BS in Computer Science, Electrical Engineering, or related field or equivalent experience
15+ years of relevant experience.
Ways To Stand Out from The Crowd:
Background in computer science, machine learning, deep learning, open-source software, infrastructure technologies, and GPU technology.
Prior experience in infrastructure software, production application software development, software development, release and support methodology and devops
Experience in the management of datacenters and large-scale distributed computing
Experience in working with AI researchers and/or EDA developers
Consistent track record of driving process improvements and measuring efficiency and a passion for sharing knowledge and experience driving complex projects end-to-end.
You will also be eligible for equity and benefits.
If an employer mentions a salary or salary range on their job, we display it as an "Employer Estimate". If a job has no salary data, Rise displays an estimate if available.
NVIDIA Public Sector is hiring a Senior Solutions Architect to lead technical integration of NVIDIA accelerated computing and AI across federal and defense workloads.
NVIDIA is seeking an experienced Security Analyst to lead incident response, threat hunting, and cloud/product security investigations across corporate and cloud environments.
Design, build, and operate scalable GCP data pipelines and analytics schemas at a fast-growing diagnostics startup powering customer products, analytics, and AI models.
Lead enterprise risk analytics strategy and architecture at Lilly to turn cross-functional risk data into executive insights that drive strategic decisions and measurable business outcomes.
Lead a global data engineering team at LexisNexis to deliver secure, scalable cloud data platforms and production data pipelines.
High-impact Data Engineer needed to architect scalable pipelines, integrate external data, and enable model-ready datasets for a fast-growing AI startup in New York.
F5 is hiring a seasoned Director of Enterprise Data Management to build enterprise-scale data governance, security, and AI oversight that enables responsible, high-impact analytics.
Amigo is hiring a Senior Software Engineer (Data) to design and operate Databricks streaming and batch pipelines that enable analytics, research, and continuous improvement of clinical AI agents.
Senior Associate Data Engineer to develop and maintain ETL/Databricks pipelines and integrations that power Rx Savings Solutions' analytics and operational reporting.
Exegy is hiring a Technical Business Analyst in New York to support market data modeling, normalization, and feed integration for its market data engineering team.
Senior Data Engineer (Python, NodeJS, Java) to design and maintain ETL and backend services that support Elsevier's clinical decision support products and analytics.
UC Irvine Libraries is hiring a Digital Initiatives Librarian to manage and grow digital collections and infrastructure while supporting digital scholarship, rights review, and campus outreach.
Stryker is hiring a remote Data Engineer to design, build, and maintain ETL pipelines and data assets for the Customer Intelligence team, ensuring reliable data delivery and strong documentation.
Senior Data Engineer needed to design and maintain complex ETL and backend systems (Python, NodeJS, Java) for Elsevier's clinical AI and healthcare data products.
Humana seeks a Data Engineer 2 to build and maintain scalable data pipelines and analytics platforms that enable data-driven healthcare insights for members.
NVIDIA is a publicly traded, multinational technology company headquartered in Santa Clara, California. NVIDIA's invention of the GPU in 1999 sparked the growth of the PC gaming market, redefined computer graphics, and ignited the era of modern AI.
72 jobs