RL@B

Research

Applied RL, from simulation to a real robot on the bench.

Pods form around client problems. The center of gravity is robotics.

Focus areas

Where we point the lab.

Robot learning

Imitation, sim-to-real and offline RL on a low-cost SO-101 arm.

LeRobotSO-101ACTOffline RL

How research runs

Read. Build. Publish.

  1. Reading group

    Weekly. One paper, one presenter.

  2. Research pods

    Formed around a client problem or a release.

  3. Capstones to papers

    DeCal projects continue as workshop papers.

  4. Open source

    Envs on the Hub, labs on GitHub.

Berkeley labs we learn from

Standing on the shoulders of BAIR.

RAIL

Sergey Levine

Offline RL, robot learning

Robot Learning Lab

Pieter Abbeel

Simulation, RL algorithms

AUTOLab

Ken Goldberg

Manipulation, data

InterACT

Anca Dragan

Human-robot interaction, RLHF

Sky Computing Lab

Ion Stoica

Agent evals, LLM infra

CHAI

Stuart Russell

Alignment, RL safety

Education pipeline

CS 198DeCal

Reinforcement Learning in Practice

Spring 2027 · 2 units · P/NP · cap 40

From MDPs to agents in 15 weeks. Every module ends in working code.

Open workshops

Your first RL env

Train an LLM with RL in 90 minutes

Tasks, rubrics and gold answers (no code)

Robot arm from scratch

Sim setup: MuJoCo, Isaac Lab, Genesis

Syllabus · 15 weeks

Wk 1 to 2

MDPs, Bellman equations, Q-learning

Wk 3 to 5

DQN, policy gradients, PPO from scratch

Wk 6 to 7

Continuous control and reward hacking

Wk 8 to 9

RLHF, DPO, RLVR: build an env, run GRPO

Wk 10 to 11

Robot learning, offline RL, world models

Wk 12

How labs actually do RL (guest lecture)

Wk 13 to 15

Project sprint and demo day

Come do research that ships.

Build-track members join a pod in their first semester.