Opportunities · Internship
Software Engineering Intern, NCCL - 2026
Nvidia
Apply with Birdee.
- When
- Summer 2026
- Where
- 2 Locations
What you'll do
- Design, implement and maintain highly-optimized communication runtimes for Deep Learning frameworks (e.g. NCCL for TensorFlow/Pytorch) and HPC programming interfaces (e.g. UCX for MPI/OpenSHMEM) on GPU clusters.
- Participating in and contributing to parallel programming interface specifications like MPI/OpenSHMEM.
- Design, implement and maintain system software that enables interactions among GPUs and interactions between GPUs and other system components.
- Creating proof-of-concepts to evaluate and motivate extensions in programming models, new designs in runtimes and new features in hardware.
What they want
- You are pursuing a Ph.D. in CE/CS/EE with a strong background in computer architecture, operating systems, communication library and/or AI/ML.
- Strong experience with Linux.
- Experience with parallel programming interfaces and communication runtimes.
- Deep knowledge of high-performance networks like InfiniBand, RoCE etc.
- Background with HPC applications. Experience with Deep Learning Frameworks such PyTorch, TensorFlow, JAX/XLA, vLLM/SGLang etc.
- Experience with AI/DL communication patterns such as Expert Parallelism (EP), TP, DP, PP and how these patterns can be implemented with NCCL. Experience with CUDA kernel optimization and profiling.
About Nvidia
NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High-Performance Computing and Visualization.
The GPU, our invention, serves as the visual cortex of modern computers and is at the heart of our products and services.
Full posting
NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High-Performance Computing and Visualization. The GPU, our invention, serves as the visual cortex of modern computers and is at the heart of our products and services. Our work opens up new universes to explore, enables amazing creativity and discovery, and powers what were once science fiction inventions from artificial intelligence to autonomous cars. NVIDIA is looking for phenomenal people like you to help us accelerate the next wave of artificial intelligence.
We are looking for a highly motivated software engineer intern for an exciting role in our communication libraries and network software team. The position will be part of a fast-paced crew that develops and maintains software for complex heterogeneous computing systems that power disruptive products in High Performance Computing and Deep Learning.
What you will be doing
- Design, implement and maintain highly-optimized communication runtimes for Deep Learning frameworks (e.g. NCCL for TensorFlow/Pytorch) and HPC programming interfaces (e.g. UCX for MPI/OpenSHMEM) on GPU clusters.
- Participating in and contributing to parallel programming interface specifications like MPI/OpenSHMEM.
- Design, implement and maintain system software that enables interactions among GPUs and interactions between GPUs and other system components.
- Creating proof-of-concepts to evaluate and motivate extensions in programming models, new designs in runtimes and new features in hardware.
What we need to see
- You are pursuing a Ph.D. in CE/CS/EE with a strong background in computer architecture, operating systems, communication library and/or AI/ML.
- Excellent C/C++ programming and debugging skills.
- Strong experience with Linux.
- Experience with parallel programming interfaces and communication runtimes.
- Ability and flexibility to work and communicate effectively in a multi-national, multi-time-zone corporate environment.
Ways to stand out from the crowd:
- Deep knowledge of high-performance networks like InfiniBand, RoCE etc.
- Background with HPC applications. Experience with Deep Learning Frameworks such PyTorch, TensorFlow, JAX/XLA, vLLM/SGLang etc.
- Experience with AI/DL communication patterns such as Expert Parallelism (EP), TP, DP, PP and how these patterns can be implemented with NCCL. Experience with CUDA kernel optimization and profiling.
- Experience with large-scale model training and production inference software stack.
- Strong collaborative and interpersonal skills, specifically a proven ability to effectively guide and influence within a dynamic matrix environment.
Dates for this one
Only the dates this listing publishes. Anything it leaves out is left out here too.
Posted
Aug 18, 2026
The employer's own record
Usually listed
Checking opening history…
Filed pay record
Checking public pay records…
Verified by UTern
Last read under an hour ago
More at Nvidia
Other roles this employer has live right now.
- NVIDIA Spring 2027 Internships: Developer and Performance TechnologySanta Clara, CA
- NVIDIA 2027 Internships: Software EngineeringSanta Clara, CA
- NVIDIA 2027 Internships: Systems Software EngineeringSanta Clara, CA
- NVIDIA 2027 Intern: Deep Learning Computer ArchitectureSanta Clara, CA
- NVIDIA 2027 Internships: Deep LearningSanta Clara, CA