This invention discloses a method and
system for mind chain hijacking
security analysis of VLA models, belonging to the fields of
artificial intelligence security and embodied intelligence. The invention first acquires the VLA model under test and collects normal trajectories to construct an offline test dataset; then, it initializes an adversarial patch and overlays it onto the original
visual observation target position to generate adversarial observation; subsequently, it constructs a joint optimized
loss function including mind chain adversarial loss and action adversarial loss, constraining the model to generate a preset target mind chain and
target action when inputting unaltered instructions; then, iterative optimization using an optimization
algorithm yields the optimal adversarial patch; finally, it uses this optimal patch to test the mind chain hijacking
vulnerability in the model under test. This invention has strong concealment and wide applicability, achieving covert and targeted hijacking of
robot behavior without tampering with user instructions, effectively detecting security vulnerabilities in the
inference layer of VLA models, and providing support for the secure deployment of embodied intelligence systems.