OCBench

OCBench is a controllable robotic manipulation benchmark for studying BC and RL.

Why make yet another benchmark?

OGBench (previous benchmark) More rule-based, 20Hz (video is 100fps).

OCBench (new benchmark) More demo-like, 50Hz (video is 100fps).

The main feature of OCBench is its scripted policies. We designed them to have more human-like properties (e.g., non-Markovian and smoothly random trajectories) than those in previous synthetic benchmarks.

OGBench, Robomimic MG, OCBench, and Human action distributions

The action histograms above show that the scripted policies in many previous benchmarks have different characteristics from human demonstrations. They tend to output saturated actions (e.g., -1 or 1) and are typically randomized via Gaussian noise. This can lead to misleading conclusions about algorithms, as algorithms that work well on these benchmarks may not transfer to real-world tasks.

OCBench is our attempt to bridge this gap by providing more realistic, easy-to-use, and controllable scripted policies. Moreover, we implement all OCBench tasks and scripted policies in MJWarp (GPU-accelerated MuJoCo). This enables infinite-data training, meaning that we can train an agent using virtually unlimited data to study the scaling behavior of BC and RL algorithms.

Please see this blog post for more details and the backstory!

Tasks

OCBench provides 28 manipulation tasks across five types of environments: block, chamber, switch, hanoi, and bowling.

Each task provides three environment variants:

  • lite has a much shorter horizon and simplified scripted policies. This can be useful for fast prototyping.
  • cpu uses CPU-based MuJoCo instead of GPU-based MJWarp.
  • visual provides pixel-based observations instead of state-based observations. A visual observation consists of three 224 × 224 × 3 images from three different camera angles:

    front camera

    side camera

    wrist camera

We support all 8 combinations of these environment variants. For example, the block-single-task1 task has the following variants:

Environment name lite cpu visual
block-single-task1 No No No
block-lite-single-task1 Yes No No
block-cpu-single-task1 No Yes No
block-cpu-lite-single-task1 Yes Yes No
visual-block-single-task1 No No Yes
visual-block-lite-single-task1 Yes No Yes
visual-block-cpu-single-task1 No Yes Yes
visual-block-cpu-lite-single-task1 Yes Yes Yes

block

block is a standard block manipulation environment. We provide four variants (single, double, triple, and quadruple) with one to four blocks. task1 and task3 require arranging blocks in a specific configuration, and task2 requires stacking blocks anywhere on the table in any order.

block-single

block-single-task1 task1 move, 3x speed
block-double

block-double-task1 task1 double_pnp, 3x speed

block-double-task2 task2 stack_anywhere, 3x speed
block-triple

block-triple-task1 task1 triple_pnp, 3x speed

block-triple-task2 task2 stack_anywhere, 3x speed
block-quadruple

block-quadruple-task1 task1 quadruple_pnp, 3x speed

block-quadruple-task2 task2 stack_anywhere, 3x speed

block-quadruple-task3 task3 grid, 3x speed

chamber

chamber is an environment involving diverse objects (blocks, a drawer, a window, and button locks). We provide three variants (easy, medium, and hard). The tasks involve putting blocks into the drawer, taking them out and stacking them on the table, opening or closing the drawer and window, and locking or unlocking the button locks.

chamber-easy

chamber-easy-task1 task1 open_all, 3x speed

chamber-easy-task2 task2 put_in, 3x speed

chamber-easy-task3 task3 take_out, 3x speed
chamber-medium

chamber-medium-task1 task1 put_all_in, 3x speed

chamber-medium-task2 task2 stack_anywhere, 3x speed
chamber-hard

chamber-hard-task1 task1 put_all_in, 3x speed

chamber-hard-task2 task2 stack_anywhere, 3x speed

switch

switch is a Lights Out puzzle environment. We provide three board sizes (3x3, 4x4, and 5x5). task1 provides a fixed grid, and task2 provides a randomized grid where at most two randomly selected buttons are omitted.

When the agent presses a button, it toggles its color and the colors of its adjacent buttons. The goal is to make all buttons the same color. The demonstrations for these tasks are highly suboptimal, so the agent must learn to solve the puzzle by itself.

switch-3x3

switch-3x3-task1 task1 monochrome, 3x speed

switch-3x3-task2 task2 monochrome_drop, 3x speed
switch-4x4

switch-4x4-task1 task1 monochrome, 3x speed

switch-4x4-task2 task2 monochrome_drop, 3x speed
switch-5x5

switch-5x5-task1 task1 monochrome, 3x speed

switch-5x5-task2 task2 monochrome_drop, 3x speed

hanoi

hanoi is a Tower of Hanoi puzzle environment. We provide three variants (single, double, and triple) with one to three disks. task1 uses fixed peg positions, and task2 randomizes the positions of the pegs.

The disks are initially stacked on the middle peg, and the goal is to move them to either the first or third peg. Since the disks fit tightly around the pegs, these tasks require highly precise manipulation.

hanoi-single

hanoi-single-task1 task1 fixed, 3x speed

hanoi-single-task2 task2 random, 3x speed
hanoi-double

hanoi-double-task1 task1 fixed, 3x speed

hanoi-double-task2 task2 random, 3x speed
hanoi-triple

hanoi-triple-task1 task1 fixed, 3x speed

hanoi-triple-task2 task2 random, 3x speed

bowling

bowling is a bowling task, whose goal is to knock down the pins. Since the dynamics are not easily predictable, this task requires a solid understanding of the environment dynamics.

bowling

bowling-task1 task1 knock_down, 3x speed

Citation

@misc{ocbench_park2026,
  title={{Behavioral Cloning Mystery}},
  author={Park, Seohong and Levine, Sergey},
  year={2026},
  url={https://seohong.me/projects/ocbench/ocbench.pdf},
}