light_mode

PAC Bench: Do Foundation Models Understand Prerequisites for Executing Manipulation Policies?

Atharva Gundawar*, Som Sagar*, Ransalu Senanayake
Conference on Neural Information Processing Systems (NeurIPS), 2025

PAC Bench teaser figure

In one sentence. PAC Bench is a benchmark of over 30,000 annotations that tests whether vision-language models understand the object Properties, action Affordances, and physical Constraints required for reliable robot manipulation.

picture_as_pdf PDF language Project Page dataset Dataset

Abstract

Vision-Language Models (VLMs) are increasingly pivotal for generalist robot manipulation, enabling tasks such as physical reasoning, policy generation, and failure detection. However, their proficiency in these high-level applications often assumes a deep understanding of low-level physical prerequisites, a capability that is largely unverified. To perform actions reliably, robots must comprehend intrinsic object properties (e.g., material, weight), action affordances (e.g., graspable, stackable), and physical constraints (e.g., stability, reachability, or an object's state like being closed). Despite their ubiquitous use in manipulation, we argue that off-the-shelf VLMs may lack this granular, physically-grounded understanding, as these specific prerequisites are often overlooked during training. Addressing this critical gap, we introduce PAC Bench, a comprehensive benchmark designed to systematically evaluate VLMs on their understanding of these core Properties, Affordances, and Constraints (PAC) from a task executability perspective. PAC Bench features a diverse dataset with over 30,000 annotations, comprising 673 real-world images (115 object classes, 15 property types, 1-3 affordances defined per class), 100 real-world humanoid-view scenarios and 120 unique simulated constraint scenarios across four tasks. Our evaluations reveal significant gaps in the ability of VLMs to grasp fundamental physical concepts, underscoring their current limitations for reliable robot manipulation and pointing to key areas that require targeted research. PAC Bench also serves as a standardized benchmark for rigorously evaluating VLM physical reasoning and guiding the development of more robust and physically grounded models for robot manipulation.

BibTeX

@inproceedings{gundawar2025pacbench,
  title     = {PAC Bench: Do Foundation Models Understand Prerequisites for Executing Manipulation Policies?},
  author    = {Gundawar, Atharva and Sagar, Som and Senanayake, Ransalu},
  booktitle = {Advances in Neural Information Processing Systems (NeurIPS)},
  year      = {2025}
}

Topics

benchmark, vision-language models, robot manipulation, physical reasoning, affordances, constraints, foundation model evaluation

About the author. Som Sagar is a computer science PhD student at Arizona State University, advised by Ransalu Senanayake in the LENS Lab. He works on reinforcement learning for post-training and agentic LLM systems, and on failure diagnosis, red teaming, and safety evaluation of foundation models, vision-language models, and robot manipulation policies.

home Homepage library_books All publications description CV school Google Scholar