arxiv:2601.00501

CPPO: Contrastive Perception for Vision Language Policy Optimization

Published on Jan 1

· Submitted by

Ahmad Rezaei on Jan 6

Upvote

Authors:

Ahmad Rezaei ,

Saeed Ranjbar Alvar ,

Mohammad Akbari

Abstract

CPPO improves vision-language model fine-tuning by detecting perception tokens through entropy shifts and using contrastive perception loss to enhance multimodal reasoning efficiency.

AI-generated summary

We introduce CPPO, a Contrastive Perception Policy Optimization method for finetuning vision-language models (VLMs). While reinforcement learning (RL) has advanced reasoning in language models, extending it to multimodal reasoning requires improving both the perception and reasoning aspects. Prior works tackle this challenge mainly with explicit perception rewards, but disentangling perception tokens from reasoning tokens is difficult, requiring extra LLMs, ground-truth data, forced separation of perception from reasoning by policy model, or applying rewards indiscriminately to all output tokens. CPPO addresses this problem by detecting perception tokens via entropy shifts in the model outputs under perturbed input images. CPPO then extends the RL objective function with a Contrastive Perception Loss (CPL) that enforces consistency under information-preserving perturbations and sensitivity under information-removing ones. Experiments show that CPPO surpasses previous perception-rewarding methods, while avoiding extra models, making training more efficient and scalable.

View arXiv page View PDF Add to collection

Community

AhNr

Paper author Paper submitter 1 day ago

CPPO: Contrastive Perception for Vision Language Policy Optimization introduces a new method (CPPO) for fine-tuning vision-language models (VLMs) using reinforcement learning. Instead of relying on explicit perception rewards or auxiliary models, the approach identifies perceptual tokens via entropy changes under perturbed images and augments the policy objective with a contrastive perception loss to improve multimodal reasoning performance and training efficiency.