Learning to Walk with PPO
Implementing PPO and a standing-to-walking curriculum for simulated locomotion. ESE 6500 (Learning in Robotics).
For ESE 6500 at Penn, I implemented Proximal Policy Optimization (PPO) in PyTorch to control a walker in the DeepMind Control Suite. When I trained it directly on walking, it learned to scoot along the ground instead. I explored whether first learning to stand could provide a better starting point for upright locomotion.
Building the controller
The policy receives a 24-dimensional observation containing body orientations, torso height, and velocities, and produces six bounded continuous actions. I implemented a Gaussian actor with three hidden layers and a separate value network with two hidden layers, both using 128-unit layers and tanh activations. Actions are squashed through tanh; the policy likelihood includes the corresponding change-of-variables correction.
Training combines PPO’s clipped policy objective with generalized advantage estimation to use the critic’s predictions when estimating which actions helped. I used a clipping threshold of 0.2, a discount factor of 0.99, and a GAE parameter of 0.95. Each update collects 4,000 environment steps, then optimizes the actor and critic with separate Adam optimizers and shuffled minibatches.
To keep training stable, I used a few practical measures: running observation normalization puts sensor channels on comparable scales, gradient clipping bounds large updates, and approximate-KL early stopping limits how far the policy moves during an update. Checkpoints preserve the actor, critic, normalization statistics, and recorded training returns.
From scooting to walking
Training the walking task from scratch produced a low, scooting behavior. To change the starting behavior, I first trained on the standing task, then transferred the actor, critic, and observation normalizer to walking. Because the standing policy had reduced its exploration noise, I reset its learned log standard deviation to −0.5 when switching tasks.
Learning to stand first gave the walker a better starting point. It moved upright and earned higher returns than the policy trained from scratch, though training was still uneven, with occasional large drops in reward.
Results
These are the average training returns over the last 25 epochs of each run. Each row represents a single training run.
| Run | Environment steps | Mean training return, last 25 epochs |
|---|---|---|
| Standing | 3 million | 779.0 |
| Walking from scratch | 3 million | 537.1 |
| Walking after standing | 3 million standing + 3 million walking | 881.7 |
Standing pretraining doubled the total interaction budget, so this wasn’t an equal-budget comparison. I’d want to repeat it across seeds with matched budgets before making a claim about sample efficiency.
Watching the walker mattered as much as watching its reward curve. Scooting earned points, but it wasn’t the behavior I wanted; learning to stand first helped close that gap.