STAR: Sparse Tactile Representation Learning in Vision–Tactile–Language–Action Models for Dexterous Manipulation

Xiangcheng Liu1*, Tianhao Wu2*, Le Zheng2*, Yidong Wang2, Bowen Jiang2
Mingjie Pan2, Xinlin Ren2, Yi Liu2, Jianlan Luo1†
1Shanghai Innovation Institute. 2Agibot.
*Equal Contribution Corresponding Author

Summary

Dexterous manipulation requires coordinated multi-finger control and effective tactile feedback, yet learning these capabilities remains challenging due to the lack of large-scale real-world data and the difficulty of extracting effective representations from sparse tactile signals. We build a robot platform and teleoperation system to collect a 200-hour bimanual dexterous manipulation dataset with synchronized visual, tactile, and language annotations, comprising 10,576 trajectories across 65 tasks, 69.5% of which involve dexterous multi-finger manipulation. We further propose STAR, an integrated training recipe for vision–tactile–language–action (VTLA) models that addresses the spatial, temporal, and informational sparsity of tactile signals through visual–tactile joint pre-training, sparse-global tactile token representation, and sparse future tactile prediction. Trained on this dataset, STAR achieves a 61% average success rate across four real-world tasks with 100 post-training trajectories per task, demonstrating dexterous performance under task-specific post-training.

Overview

STAR overview of dataset, sparse tactile learning, and policy rollout
Overview. Our training recipe, applied to a 200-hour self-collected dexterous hand dataset with tactile inputs, systematically addresses spatial, temporal, and informational sparsity, enabling dexterous performance across diverse post-training manipulation tasks.

Method

STAR training pipeline
Training Recipe. a) We jointly predict the masked tokens of RGB and tactile images for both intra- and inter-modal tactile pre-training. b) We construct the sparse-global tactile token representation for policy training by retaining activated spatial tactile tokens and per-hand global tactile tokens. c) During policy pre-training, we add sparse future tactile prediction to model informative future contact dynamics.

Evaluation Setup

STAR evaluation setup

Comparison of Different Manipulation Policies

Earbud-flipping

Ours

π0.5

LDA-1B

GR00T N1.7

Stacked-book retrieval

Ours

π0.5

LDA-1B

GR00T N1.7

Postcard retrieval

Ours

π0.5

LDA-1B

GR00T N1.7

Multi-object grasping

Ours

π0.5

LDA-1B

GR00T N1.7

Ours Rollout with Tactile Visualization

Earbud-flipping

Stacked-book retrieval

Postcard retrieval

Multi-object grasping