UTern

Opportunities · Research opportunity

Research Intern - AV Planning

Waymo

Applies on the employer's siteNew to UTern today90 days left to apply

Apply with Birdee.

Pay
about $31-$62/hr
When
Summer 2027
Where
Mountain View, CA · On-site
Apply by
Jan 7, 2027

What you'll do

  • Design, implement, and benchmark a system where LLM meta-agents generate, synthesize, and iteratively refine reward functions for autonomous driving planning policies
  • Diagnose policy failure modes, reward exploitation, and multi-objective trade-offs, integrating findings back into the automated discovery loop
  • Collaborate with scientists engineers to transfer validated reward formulations and tooling into planner recipes
  • Document research outcomes, distill key design patterns, and prepare findings for internal tech transfers and potential conference publication
  • Track record of peer-reviewed publications at top AI/robotics conferences (e.g., NeurIPS, ICRA, IROS, CoRL, ICML, ICLR, CVPR)

What they want

  • Currently pursuing a PhD in Computer Science, Robotics, Machine Learning, or a related quantitative field
  • Strong hands-on experience in Reinforcement Learning (RL / RLFT) and training deep learning policies with PyTorch or JAX
  • Demonstrated proficiency in Python and modern ML engineering practices
  • Familiarity with autonomous vehicle planning, trajectory optimization, or sequential decision-making under uncertainty
  • Experience with LLM-guided code generation, automated reward discovery (e.g., Eureka-like frameworks), or meta-RL methods
  • Hands-on experience with RL or RLFT training in LLM/VLM/robotics
About Waymo

Waymo is an autonomous driving technology company with the mission to be the world's most trusted driver.

Since its start as the Google Self-Driving Car Project in 2009, Waymo has focused on building the Waymo Driver—The World's Most Experienced Driver™—to improve access to mobility while saving thousands of lives now lost to traffic crashes.

Full posting

Waymo is an autonomous driving technology company with the mission to be the world's most trusted driver. Since its start as the Google Self-Driving Car Project in 2009, Waymo has focused on building the Waymo Driver—The World's Most Experienced Driver™—to improve access to mobility while saving thousands of lives now lost to traffic crashes. The Waymo Driver powers Waymo’s fully autonomous ride-hail service and can also be applied to a range of vehicle platforms and product use cases. The Waymo Driver has provided over ten million rider-only trips, enabled by its experience autonomously driving over 100 million miles on public roads and tens of billions in simulation across 15+ U.S. states.

The mission of the Waymo AI Foundations team is to develop machine learning solutions addressing open problems in autonomous driving, towards the goal of safely operating Waymo vehicles in dozens of cities and under all driving conditions. As part of our work, we also initiate and foster collaborations with other research teams in Alphabet. AI Foundations areas that we are currently focusing on include reinforcement learning, learning from demonstration, generative modeling, Bayesian inference, hierarchical learning, and robust evaluation.

Waymo interns partner with leaders in the industry on projects that create impact to the company. We believe learning is a two-way street: applying your knowledge while providing you with opportunities to expand your skill-set. Interns are an important part of our culture and our recruiting pipeline. Join us at Waymo for a fun and rewarding internship!

You will

  • Design, implement, and benchmark a system where LLM meta-agents generate, synthesize, and iteratively refine reward functions for autonomous driving planning policies
  • Diagnose policy failure modes, reward exploitation, and multi-objective trade-offs, integrating findings back into the automated discovery loop
  • Collaborate with scientists engineers to transfer validated reward formulations and tooling into planner recipes
  • Document research outcomes, distill key design patterns, and prepare findings for internal tech transfers and potential conference publication

You have

  • Currently pursuing a PhD in Computer Science, Robotics, Machine Learning, or a related quantitative field
  • Strong hands-on experience in Reinforcement Learning (RL / RLFT) and training deep learning policies with PyTorch or JAX
  • Demonstrated proficiency in Python and modern ML engineering practices
  • Familiarity with autonomous vehicle planning, trajectory optimization, or sequential decision-making under uncertainty

We prefer

  • Track record of peer-reviewed publications at top AI/robotics conferences (e.g., NeurIPS, ICRA, IROS, CoRL, ICML, ICLR, CVPR)
  • Experience with LLM-guided code generation, automated reward discovery (e.g., Eureka-like frameworks), or meta-RL methods
  • Hands-on experience with RL or RLFT training in LLM/VLM/robotics
  • Proven ability to independently drive an exploratory research problem from ideation to implementation and results

Note: This will be a hybrid onsite internship position. We will accept resumes on a rolling basis until the role is filled. To be in consideration for multiple roles, you will need to apply to each one individually - please apply to the top 3 roles you are interested in.

The expected hourly rate for this full-time position is listed below. Interns are also eligible to participate in the Company’s generous benefits programs, subject to eligibility requirements.

Hourly PhD Pay

$85—$85 USD

Dates for this one

Only the dates this listing publishes. Anything it leaves out is left out here too.

Posted

Oct 9, 2026

Applications close

Jan 7, 2027

90 days left

The employer's own record

Usually listed

Checking opening history…

Filed pay record

Checking public pay records…

Verified by UTern

Last read under an hour ago · pay stated in the listing · closing date published

More at Waymo

Other roles this employer has live right now.