OCBench is a controllable robotic manipulation benchmark for studying BC and RL.
OGBench (previous benchmark) More rule-based, 20Hz (video is 100fps).
OCBench (new benchmark) More demo-like, 50Hz (video is 100fps).
The main feature of OCBench is its scripted policies. We designed them to have more human-like properties (e.g., non-Markovian and smoothly random trajectories) than those in previous synthetic benchmarks.
The action histograms above show that the scripted policies in many previous benchmarks have different characteristics from human demonstrations. They tend to output saturated actions (e.g., -1 or 1) and are typically randomized via Gaussian noise. This can lead to misleading conclusions about algorithms, as algorithms that work well on these benchmarks may not transfer to real-world tasks.
OCBench is our attempt to bridge this gap by providing more realistic, easy-to-use, and controllable scripted policies. Moreover, we implement all OCBench tasks and scripted policies in MJWarp (GPU-accelerated MuJoCo). This enables infinite-data training, meaning that we can train an agent using virtually unlimited data to study the scaling behavior of BC and RL algorithms.
Please see this blog post for more details and the backstory!
OCBench provides 28 manipulation tasks across five types of environments: block, chamber, switch, hanoi, and bowling.
Each task provides three environment variants:
front camera
side camera
wrist camera
We support all 8 combinations of these environment variants. For example, the block-single-task1 task has the following variants:
| Environment name | lite | cpu | visual |
|---|---|---|---|
| block-single-task1 | No | No | No |
| block-lite-single-task1 | Yes | No | No |
| block-cpu-single-task1 | No | Yes | No |
| block-cpu-lite-single-task1 | Yes | Yes | No |
| visual-block-single-task1 | No | No | Yes |
| visual-block-lite-single-task1 | Yes | No | Yes |
| visual-block-cpu-single-task1 | No | Yes | Yes |
| visual-block-cpu-lite-single-task1 | Yes | Yes | Yes |
block is a standard block manipulation environment. We provide four variants (single, double, triple, and quadruple) with one to four blocks. task1 and task3 require arranging blocks in a specific configuration, and task2 requires stacking blocks anywhere on the table in any order.
block-single-task1 task1 move, 3x speed
block-double-task1 task1 double_
block-double-task2 task2 stack_
block-triple-task1 task1 triple_
block-triple-task2 task2 stack_
block-quadruple-task1 task1 quadruple_
block-quadruple-task2 task2 stack_
block-quadruple-task3 task3 grid, 3x speed
chamber is an environment involving diverse objects (blocks, a drawer, a window, and button locks). We provide three variants (easy, medium, and hard). The tasks involve putting blocks into the drawer, taking them out and stacking them on the table, opening or closing the drawer and window, and locking or unlocking the button locks.
chamber-easy-task1 task1 open_
chamber-easy-task2 task2 put_
chamber-easy-task3 task3 take_
chamber-medium-task1 task1 put_
chamber-medium-task2 task2 stack_
chamber-hard-task1 task1 put_
chamber-hard-task2 task2 stack_
switch is a Lights Out puzzle environment. We provide three board sizes (3x3, 4x4, and 5x5). task1 provides a fixed grid, and task2 provides a randomized grid where at most two randomly selected buttons are omitted.
When the agent presses a button, it toggles its color and the colors of its adjacent buttons. The goal is to make all buttons the same color. The demonstrations for these tasks are highly suboptimal, so the agent must learn to solve the puzzle by itself.
switch-3x3-task1 task1 monochrome, 3x speed
switch-3x3-task2 task2 monochrome_
switch-4x4-task1 task1 monochrome, 3x speed
switch-4x4-task2 task2 monochrome_
switch-5x5-task1 task1 monochrome, 3x speed
switch-5x5-task2 task2 monochrome_
hanoi is a Tower of Hanoi puzzle environment. We provide three variants (single, double, and triple) with one to three disks. task1 uses fixed peg positions, and task2 randomizes the positions of the pegs.
The disks are initially stacked on the middle peg, and the goal is to move them to either the first or third peg. Since the disks fit tightly around the pegs, these tasks require highly precise manipulation.
hanoi-single-task1 task1 fixed, 3x speed
hanoi-single-task2 task2 random, 3x speed
hanoi-double-task1 task1 fixed, 3x speed
hanoi-double-task2 task2 random, 3x speed
hanoi-triple-task1 task1 fixed, 3x speed
hanoi-triple-task2 task2 random, 3x speed
bowling is a bowling task, whose goal is to knock down the pins. Since the dynamics are not easily predictable, this task requires a solid understanding of the environment dynamics.
bowling-task1 task1 knock_
@misc{ocbench_park2026,
title={{Behavioral Cloning Mystery}},
author={Park, Seohong and Levine, Sergey},
year={2026},
url={https://seohong.me/projects/ocbench/ocbench.pdf},
}