VeriFine

Scaling Verification for Self-Improvement in Embodied Reasoning

1NVIDIA 2UCLA 3UC Berkeley 4Stanford University

* Work done during an internship at NVIDIA. † Corresponding author.

01 / THE VERIFICATION BOTTLENECK

Evolving verification unlocks further policy improvementA conceptual comparison: a fixed judge can cause policy progress to plateau or decline through reward hacking. Two judge refinements enable further improvement. These curves illustrate possible behavior, not experimental measurements. Policy capability Fixed judge: plateauReward hacking risk Refine the judge Refine again LearnReach a limitLearn moreKeep improving Policy capability Fixed judge: plateauReward hacking risk Refine the judge Refine again
Evolving policy + judgeFixed judgeReward hacking riskJudge improvement
04 / 04

Reopen the learning frontier.

Better feedback unlocks more learning. Repeat as new failures emerge.

Conceptual verification story.

02 / THE VERIFINE FRAMEWORK

VeriFine framework: Figure 2 from the paper Three panels compare policy improvement, previous fixed-curriculum and static-judge methods, and judge improvement. Below, Policy sends self-generated rollouts to a reference-free Judge, which returns rewards and an adaptive curriculum. The Judge selects informative failure cases for targeted human feedback and improves through coactive calibration. The five steps below highlight the corresponding parts of this original figure.

03 / MEASURED PROGRESS

JudgePolicy

Swipe to follow the rounds →

04 / QUALITATIVE RESULTS

Judge refinement + policy improvement

05 / VERIFICATION ANALYSIS

View chart data

VERIFINE

Keep learning.
Keep improving verification.

VeriFine connects policy improvement with judge improvement. Better policies reveal new failures; the Judge Improvement Loop turns those failures into more reliable verification for the next policy update.

Back to the paper and authors ↑

Cite this work

@misc{zhou2026verifine,
  title={VeriFine: Scaling Verification for Self-Improvement in Embodied Reasoning},
  author={Zhou, Zewei and Luo, Rachel and Cao, Yulong and Xiao, Chaowei and Peng, Chensheng and Li, Boyi and Tian, Thomas and Lian, Zheng and Wang, Yan and Ma, Jiaqi and Ivanovic, Boris and Pavone, Marco and Ding, Wenhao},
  year={2026}
}