Lab LedgerDesk

Apple · Launch · 2026-09-17

REVERSAL-BENCH: A Reversibility Axis and Reset Oracle for Measuring the Reset-Free RL Cliff

A central goal of autonomous reinforcement learning is continuous policy training without external resets. However, existing paradigms largely depend on…

What moved

A central goal of autonomous reinforcement learning is continuous policy training without external resets. However, existing paradigms largely depend on underlying environmental reversibility, a property absent in real world manipulation, where events such as pushing objects off tables or spilling granular substances cannot be undone. We introduce REVERSAL-BENCH, a benchmark that controls reversibility via a continuous parameter ρ∈ [0, 1] and provides a reset oracle, a ground-truth verification mechanism to test state recoverability across eight manipulation settings in five physics.

Why it matters

A central goal of autonomous reinforcement learning is continuous policy training without external resets. That is a public launch file from Apple, dated 2026-09-17. Tagged Agents / Science.

On the record

  • Filed from the Apple official RSS on 2026-09-17.
  • Primary source host: machinelearning.apple.com.
  • A central goal of autonomous reinforcement learning is continuous policy training without external resets.
Primary source
machinelearning.apple.com
Desk
Logged as brief 023