CupPhysical grasp and place
Experimental videos
Watch TrapVLA in action
Start with physical robot demonstrations, then explore configured failures across LIBERO and RoboTwin simulation benchmarks.
01
Real-world robot
Eggplant and Cup pick-and-place tasks
02
Trap-LIBERO
Four task suites and four representative failure modes
03
Trap-RoboTwin
Bimanual Shoes and Fan tasks
Task overview
Configured Failure Trapping studies how the robot fails
Abstract
This work introduces Configured Failure Trapping, a novel backdoor attack task against Vision-Language-Action (VLA) models, which aims to activate attacks through stealthy textual triggers and induce configured failure modes. Unlike prior backdoor attacks that treat any task failure as a successful attack, Configured Failure Trapping requires the attacker to control how the robot fails—for example, causing the robot to grasp with a specified positional offset—making it substantially more challenging and hard to detect.
To support the new task, we propose TrapEngine for synthesizing high-quality target trajectories and TrapEval for measuring clean-task performance and configured-failure fidelity. We construct two new benchmarks, Trap-LIBERO and Trap-RoboTwin, spanning four representative failure modes. To mitigate sparse action deviation, TrapVLA explicitly learns trigger-induced action residuals that steer the policy toward the configured failure behavior.
New attack task
Configured Failure Trapping studies trigger-conditioned, configured failure modes rather than unconstrained task failure.
Engine & benchmarks
The study instantiates four representative failure modes on LIBERO and RoboTwin, resulting in Trap-LIBERO and Trap-RoboTwin.
Target Residual Steering
TrapVLA learns trigger-induced action residuals within the target-deviation intervals.
TrapEngine
A configure-then-replay data engine
TrapEngine replays successful demonstrations, generates target trajectories with configured action modifications, filters infeasible or invalid trajectories through human inspection, and ranks textual prefix candidates by perplexity.

TrapEval
Measure whether the failure matches the configuration
Success Rate measures task completion under clean instructions; Configured Attack Success Rate measures the percentage of triggered trials that satisfy the configured failure mode; AVE combines clean-task preservation and mean C-ASR.

Four failure modes
Temporal and spatial failure patterns
- ECEarly CloseClose before the intended grasp point
- GDGrasp DeviationShift the gripper from the target grasp
- EOEarly OpenRelease before reaching the goal
- RDRelease DeviationRelease at a configured offset
Paper tables
Reported experimental results
Results reported for Trap-LIBERO and real-world tasks with OpenVLA-OFT and π0.5 victim models.
Real-world study
Evaluation on a physical robot
We evaluate TrapVLA with a ROKAE 6-DoF robot on Eggplant and Cup tasks, using 50 demonstrations per task and evaluating each clean or triggered condition over 30 trials.



@misc{liu2026trapvla,
title = {TrapVLA: Trapping Vision-Language-Action Models in
Configured Failure Modes},
author = {Liu, Jun-Hui and Lin, Kun-Yu and Wei, Yi-Lin and
Chen, Xu-Han and Li, Yinghao and Li, Zhuohao and
Li, Yuan-Ming and Zhang, Qing and Fan, Xiaoyi and
Jiang, Dongmei and Li, Yan and Zheng, Wei-Shi},
year = {2026},
eprint = {2608.26578},
archivePrefix = {arXiv},
primaryClass = {cs.RO},
url = {https://arxiv.org/abs/2608.26578}
}