Peg Insertion
Chip Handling
Cap Unscrewing
Board Wiping
Tactile · Vision · Action

ForeTac-VLA: A Forecasting-Based Tactile-Vision-Language-Action Model for Contact-Rich Robotic Manipulation

Forecasting future touch for contact-aware robotic manipulation.

Anonymous Authors

Paper · Coming soon Dataset · Coming soon Code · Coming soon

Four real-world contact-rich manipulation tasks

Overview

Forecasting tactile states to guide action generation

ForeTac-VLA integrates recent tactile measurements with vision-language features through bidirectional cross-attention and predicts multi-step future tactile states to directly condition action generation.

Overview of the ForeTac-VLA framework, manipulation tasks, and robustness environments
Overview of ForeTac-VLA: multimodal inputs, bidirectional cross-attention, tactile forecasting, action generation, four manipulation tasks, and robustness environments.
95.0%Average success rate
+36.25 ppOver Fine-tuned VLA
+22.50 ppOver the strongest tactile baseline
Qualitative comparison

Task demonstrations

Each task compares ForeTac-VLA with three baselines under the same real-world contact-rich manipulation setting. All task demonstration videos are shown at 1.5× speed.

01

Peg Insertion

Precise contact alignment during insertion.

Fine-tuned VLA

Concat-and-Gating TacVLA

FiLM-Fusion TacVLA

02

Chip Handling

Stable grasping of fragile objects without excessive force.

Fine-tuned VLA

Concat-and-Gating TacVLA

FiLM-Fusion TacVLA

03

Cap Unscrewing

Rotational manipulation with continuously evolving contact.

Fine-tuned VLA

Concat-and-Gating TacVLA

FiLM-Fusion TacVLA

04

Board Wiping

Sustained surface contact during long-horizon manipulation.

Fine-tuned VLA

Concat-and-Gating TacVLA

FiLM-Fusion TacVLA

Experimental results

Success rates across four contact-rich tasks

Each model is evaluated over 20 trials per task. ForeTac-VLA achieves a 95.0% average success rate, outperforming the Fine-tuned VLA by 36.25 percentage points and the strongest tactile-enhanced baseline by 22.50 percentage points.

ModelPeg InsertionChip HandlingCap UnscrewingBoard WipingAverage
ACT2/205/200/204/2013.75%
Fine-tuned VLA13/2011/2012/2011/2058.75%
FiLM-Fusion TacVLA14/2014/2016/2013/2071.25%
Concat-and-Gating TacVLA16/2013/2015/2014/2072.50%
ForeTac-VLA (Ours)20/2019/2019/2018/2095.00%
Robustness & generalization

Robustness under challenging visual conditions

ForeTac-VLA is evaluated across all four tasks under challenging visual conditions. All generalization videos are shown at 5× speed.

Success-rate comparison in low-illumination and visually cluttered environments
Success-rate comparison under low-illumination and visually cluttered environments across all four tasks.
01

Low-Illumination Environment

ForeTac-VLA · 5× speed
ForeTac-VLA

Peg Insertion

ForeTac-VLA

Chip Handling

ForeTac-VLA

Cap Unscrewing

ForeTac-VLA

Board Wiping

02

Visually Cluttered Environment

ForeTac-VLA · 5× speed
ForeTac-VLA

Peg Insertion

ForeTac-VLA

Chip Handling

ForeTac-VLA

Cap Unscrewing

ForeTac-VLA

Board Wiping

Publication

Paper and code coming soon

Paper, code, and citation information will be released following publication.