Job details

Inference Engineering Manager

About the Role

We are looking for an Inference Engineering Manager to lead our AI Inference team. This is a unique opportunity to build and scale the infrastructure that powers Perplexity's products and APIs, serving millions of users with state-of-the-art AI capabilities.

You will own the technical direction and execution of our inference systems while building and leading a world-class team of inference engineers. Our current stack includes Python, PyTorch, Rust, C++, and Kubernetes. You will help architect and scale the large-scale deployment of machine learning models behind Perplexity's Comet, Sonar, Search, Deep Research products.

Why Perplexity?

Build SOTA systems that are the fastest in the industry with cutting-edge technology
High-impact work on a smaller team with significant ownership and autonomy
Opportunity to build 0-to-1 infrastructure from scratch rather than maintaining legacy systems
Work on the full spectrum: reducing cost, scaling traffic, and pushing the boundaries of inference
Direct influence on technical roadmap and team culture at a rapidly growing company

Responsibilities

Lead and grow a high-performing team of AI inference engineers
Develop APIs for AI inference used by both internal and external customers
Architect and scale our inference infrastructure for reliability and efficiency
Benchmark and eliminate bottlenecks throughout our inference stack
Drive large sparse/MoE model inference at rack scale, including sharding strategies for massive models
Push the frontier with building inference systems to support sparse attention, disaggregated pre-fill/decoding serving, etc.
Improve the reliability and observability of our systems and lead incident response
Own technical decisions around batching, throughput, latency, and GPU utilization
Partner with ML research teams on model optimization and deployment
Recruit, mentor, and develop engineering talent
Establish team processes, engineering standards, and operational excellence

Qualifications

5+ years of engineering experience with 2+ years in a technical leadership or management role
Deep experience with ML systems and inference frameworks (PyTorch, TensorFlow, ONNX, TensorRT, vLLM)
Strong understanding of LLM architecture: Multi-Head Attention, Multi/Grouped-Query Attention, and common layers
Experience with inference optimizations: batching, quantization, kernel fusion, FlashAttention
Familiarity with GPU characteristics, roofline models, and performance analysis
Experience deploying reliable, distributed, real-time systems at scale
Track record of building and leading high-performing engineering teams
Experience with parallelism strategies: tensor parallelism, pipeline parallelism, expert parallelism
Strong technical communication and cross-functional collaboration skills

Nice to Have

Experience with CUDA, Triton, or custom kernel development
Background in training infrastructure and RL workloads
Experience with Kubernetes and container orchestration at scale
Published work or contributions to inference optimization research

Inference LLM PyTorch Python Rust C++ Kubernetes GPU CUDA Triton Quantization FlashAttention Tensor parallelism Pipeline parallelism MoE Model serving ML infrastructure Engineering manager

Perplexity AI Glassdoor Company Review

3.1

Perplexity AI DE&I Review

2.6

CEO of Perplexity AI

Aravind Srinivas

Approve of CEO

Average salary estimate

$260000 / YEARLY (est.)

min

max

$200000K

$320000K

If an employer mentions a salary or salary range on their job, we display it as an "Employer Estimate". If a job has no salary data, Rise displays an estimate if available.

What it's like to work at Perplexity AI

Read Reviews

Similar Jobs

Manager, Software Engineering - AV Frameworks

NVIDIA Hybrid US, CA, Santa Clara

VIEW

Posted 6 hours ago

Customer-Centric

Mission Driven

Inclusive & Diverse

Rise from Within

Diversity of Opinions

Work/Life Harmony

Growth & Learning

Transparent & Candid

Medical Insurance

Paid Time-Off

Maternity Leave

Mental Health Resources

Equity

Child Care stipend

Paternity Leave

WFH Reimbursements

Flex-Friendly

Dental Insurance

Vision Insurance

Life insurance

Health Savings Account (HSA)

Flexible Spending Account (FSA)

401K Matching

Military leave

Lead development and optimization of NVIDIA's DriveAV framework as a Manager of System Software Engineering, driving high-performance C++/CUDA software across heterogeneous compute for autonomous vehicles.

GPU Software Development Engineer

Intel Hybrid US, California, Folsom

VIEW

Posted 7 hours ago

Inclusive & Diverse

Rise from Within

Mission Driven

Diversity of Opinions

Work/Life Harmony

Growth & Learning

Transparent & Candid

Customer-Centric

Snacks

Onsite Gym

Family Coverage (Insurance)

Medical Insurance

Dental Insurance

Vision Insurance

Mental Health Resources

Life insurance

Disability Insurance

Health Savings Account (HSA)

Flexible Spending Account (FSA)

Learning & Development

Paid Time-Off

401K Matching

Maternity Leave

Paternity Leave

Intel is hiring a GPU Software Development Engineer to validate and debug graphics IP, enable new features, and improve performance across display, media, 3D, and compute domains.

Technical Lead - Mern Stack

Lytegen Hybrid No location specified

VIEW

Posted 21 hours ago

Lead the technical design and build of CRM integrations and a scalable middleware/API layer for a rapidly expanding residential solar business.

Senior Software Engineer

Getty Images Hybrid No location specified

VIEW

Posted 19 hours ago

Experienced software engineer needed to develop and optimize backend and full-stack features for a global visual content platform at Getty Images.

Entry Level Software Developer (Remote)

Jobgether Hybrid Kansas

VIEW

Posted 21 hours ago

Remote entry-level software developer role for recent graduates or career changers to gain practical client-facing experience using Java, Python, or JavaScript while receiving training and mentorship through Jobgether.

Engineering Manager, Search (Remote from US)

Jobgether Hybrid US

VIEW

Posted 23 hours ago

Lead a remote engineering team to design, build, and scale a high-performance distributed search and data-processing product for a fast-moving private company.

Remote UI Developer

Jobgether Hybrid California

VIEW

Posted 20 hours ago

Work remotely with a cross-functional team to design and build software solutions that apply Python and mathematical approaches to complex data problems.

Engineering Manager, Stream Control Plane

Jobgether Hybrid US

VIEW

Posted 23 hours ago

Lead and grow a remote engineering team focused on designing and delivering large-scale, high-throughput data streaming systems for a SaaS environment.

Sr. Embedded Software Engineer (Starlink)

SpaceX Hybrid Bastrop, TX

VIEW

Posted 8 hours ago

Mission Driven

Social Impact Driven

Passion for Exploration

Reward & Recognition

Lead development of Linux- and RTOS-based embedded software for Starlink consumer devices and gateways, focusing on reliability, performance, and large-scale deployment.

Remote Software Engineer

Jobgether Hybrid Georgia

VIEW

Posted 19 hours ago

Work remotely as a Software Engineer delivering and maintaining applications that support government transportation programs and public safety.

Senior Full-Stack Engineer

Knox County Schools Hybrid Charlotte

VIEW

Posted 8 hours ago

Lead the frontend for KnoxAI’s Nuxt 3 admin and customer applications to deliver polished, accessible UX for FedRAMP compliance workflows while contributing to backend improvements as needed.

Front-End Engineer

NVIDIA Hybrid US, CA, Santa Clara

VIEW

Posted 5 hours ago

Customer-Centric

Mission Driven

Inclusive & Diverse

Rise from Within

Diversity of Opinions

Work/Life Harmony

Growth & Learning

Transparent & Candid

Medical Insurance

Paid Time-Off

Maternity Leave

Mental Health Resources

Equity

Child Care stipend

Paternity Leave

WFH Reimbursements

Flex-Friendly

Dental Insurance

Vision Insurance

Life insurance

Health Savings Account (HSA)

Flexible Spending Account (FSA)

401K Matching

Military leave

Develop high-performance, scalable front-end applications and reusable UI components at NVIDIA to elevate interactive user experiences and support AI agent initiatives.

Software Developer - Journeyman

MAXISIQ, Inc. Hybrid Lorton, VA, USA

VIEW

Posted 13 hours ago

Solve real-world cyber challenges as a hybrid Software Developer supporting mission-critical networks at MAXISIQ in Lorton, VA.

Perplexity AI

Perplexity offers an AI chatbot-powered research and conversational search engine that answers queries using natural language predictive text. Since it's launch in 2022, has raised $165 million in funding, valuing the company at over $1 billion.

8 jobs

MATCH

Calculating your matching score...

BADGES