Visual Prompting for Robotic Manipulation with Annotation-Guided Pick-and-Place Using ACT

Guiding data-driven robot manipulation using simple visual annotations
Muhammad A. Muttaqien*, Tomohiro Motoda*, Ryo Hanai, Yukiyasu Domae
Embodied AI Research Team · National Institute of AIST · Tokyo, Japan
* Equal contribution

Abstract

Robotic pick-and-place tasks in convenience stores pose challenges due to dense object arrangements, occlusions, and variations in object properties such as color, shape, size, and texture. These factors complicate trajectory planning and grasping. This paper introduces a perception-action pipeline leveraging annotation-guided visual prompting, where bounding box annotations identify both pickable objects and placement locations, providing structured spatial guidance. Instead of traditional step-by-step planning, we employ Action Chunking with Transformers (ACT) as an imitation learning algorithm, enabling the robotic arm to predict chunked action sequences from human demonstrations. This facilitates smooth, adaptive, and data-driven pick-and-place operations. We evaluate our system based on success rate and visual analysis of grasping behavior, demonstrating improved grasp accuracy and adaptability in retail environments.

Why Visual Prompting?

A robot does not always need to understand every object in a cluttered scene. A simple visual cue can tell the policy what matters for the current manipulation task.

Cluttered Scenes Dense arrangements and occlusions make object selection difficult.
Diverse Products Objects vary in shape, size, texture, color, and graspability.
Visual Prompting Bounding boxes provide lightweight spatial guidance for manipulation.

Method

The manipulation policy receives both the original camera observation and a visually prompted image. Green annotations indicate the object to pick, while red annotations indicate the placement destination.

Architecture of ACT with annotation-guided visual prompting for robotic pick-and-place

Visual features from the original and prompted images are fused with robot joint states and processed by ACT to predict a chunk of future joint-space actions.

Annotation-Guided Manipulation

The annotation encodes task-relevant spatial information directly into the robot's visual observation, reducing the need for exhaustive scene understanding.

Pick Prompt Green bounding box identifies the target object.
Place Prompt Red bounding box identifies the destination.

Progressive Task Complexity

The system is evaluated across progressively harder scenarios to study how visual prompting behaves as object diversity and manipulation complexity increase.

1 Simple

Similar box-shaped products arranged in a 3×3 layout with one prompted picking target.

2 Complex

Diverse products with different shapes, sizes, colors, and textures arranged in a 3×3 layout.

3 More Complex

Diverse products at varying positions with separate picking and placement prompts.

Key Results

Visual prompting successfully guides ACT across increasingly diverse and spatially challenging manipulation scenarios.

90%
Simple Scenario
Similar rigid products
70% 100%
Complex Scenario
After additional failure-case data
80–90%
More Complex Scenario
Across evaluated product categories

Performance varies with object properties and task complexity. Detailed per-category results are available in the paper.

Failure Analysis

Remaining failures are primarily related to physical interaction and precision rather than object selection alone.

Finger Misalignment Gripper fingers do not align correctly with the object.
Weak Grasp Insufficient grasp force allows the object to slip.
Slippery Surface Object properties make stable manipulation difficult.
Placement Misalignment The object reaches the destination but is inaccurately placed.
Placement Failure The object is dropped before completing the task.
Narrow Destination Tight placement regions require higher manipulation precision.

Takeaway

What if a robot could be told where to act simply by marking the image? This research shows that lightweight bounding-box visual prompts can provide effective spatial guidance for data-driven robotic manipulation. By combining task-relevant visual annotations with ACT, the robot can learn smooth pick-and-place behaviors across objects with different appearances and physical properties, without requiring exhaustive scene parsing or rigid motion-planning rules.

🚀 As for future work, the visual prompting framework could be extended toward automatically generated prompts from vision-language models, larger and more diverse manipulation datasets, and general-purpose robot policies that can interpret richer visual and language guidance.

BibTeX

@inproceedings{muttaqien2025visprompting,
  author    = {Muhammad A. Muttaqien and Tomohiro Motoda and Ryo Hanai and Yukiyasu Domae},
  title     = {Visual Prompting for Robotic Manipulation with Annotation-Guided Pick-and-Place Using ACT},
  booktitle = {2025 IEEE International Conference on Automation Science and Engineering (CASE)},
  year      = {2025},
  doi       = {10.1109/CASE58245.2025.11164002}
}