Scaling Verification for Self-Improvement in Embodied Reasoning
01 / THE VERIFICATION BOTTLENECK
Evolving policy + judgeFixed judgeReward hacking riskJudge improvement
04 / 04
Reopen the learning frontier.
Better feedback unlocks more learning. Repeat as new failures emerge.
02 / THE VERIFINE FRAMEWORK
03 / MEASURED PROGRESS
JudgePolicy
Swipe to follow the rounds →
04 / QUALITATIVE RESULTS
05 / VERIFICATION ANALYSIS
View chart data
VERIFINE
Keep learning.
Keep improving verification.
VeriFine connects policy improvement with judge improvement. Better policies reveal new failures; the Judge Improvement Loop turns those failures into more reliable verification for the next policy update.
Back to the paper and authors ↑Cite this work
@misc{zhou2026verifine,
title={VeriFine: Scaling Verification for Self-Improvement in Embodied Reasoning},
author={Zhou, Zewei and Luo, Rachel and Cao, Yulong and Xiao, Chaowei and Peng, Chensheng and Li, Boyi and Tian, Thomas and Lian, Zheng and Wang, Yan and Ma, Jiaqi and Ivanovic, Boris and Pavone, Marco and Ding, Wenhao},
year={2026}
}