Chain reasoning hidden backdoor vulnerability detection method for vision-language-action model

The abnormal backdoor branch is constructed through multimodal trigger feature and prefix tuning technology, which solves the problem of hidden vulnerability detection of vision-language-action models in complex tasks, and realizes high-precision non-invasive vulnerability detection, which improves the security and robustness of the model in complex scenarios.

CN120509039APending Publication Date: 2025-08-19HUNAN UNIV
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510617783.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-14
Publication Date
2025-08-19

AI Technical Summary

Technical Problem

The existing visual-language-action models have hidden vulnerability in the multimodal reasoning chain, the traditional detection method has low recognition rate, is difficult to adapt to complex tasks and environmental changes, and the model is not robust enough.

Method used

Through multimodal trigger feature dynamic injection, hybrid inference chain construction and prefix tuning technology, an abnormal backdoor branch is constructed and non-invasive verification is performed to evaluate the risk of behavior deviation of the model in complex inference scenarios.

Benefits of technology

It realizes high-sensitivity hidden vulnerability detection of vision-language-action models in complex scenarios, significantly improving detection accuracy and coverage, and ensuring the safety and reproducibility of the detection process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120509039A_ABST
    Figure CN120509039A_ABST
Patent Text Reader

Abstract

The invention relates to the field of personal intelligent security evaluation, and particularly discloses a chain reasoning hidden backdoor vulnerability detection method of a vision-language-action model, which comprises the following steps of: respectively injecting micro pixel disturbance and rare character marks into vision and language input; on the basis of model autoregression prediction characteristics, designing a hybrid reasoning chain fusing normal reasoning steps and abnormal backdoor branches; adopting prefix tuning to take the hybrid reasoning sequence as a pluggable prefix injection model; and generating the vulnerability sensitivity of the abnormal action instruction through the systematic verification process detection model. Compared with an existing method, the method has the advantages that a nondestructive testing mechanism based on prefix adjustment and optimization does not need to modify model parameters or depend on training data, and the safety and reproducibility of detection are guaranteed; a multi-mode triggering mechanism is constructed, and the hidden vulnerability of the model in a complex scene is effectively revealed; the abnormal branches and the normal process are fused in a chain mode, and the defense capability of the model for the concealment logic offset can be systematically evaluated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of artificial intelligence security assessment, and in particular to a hidden backdoor vulnerability detection method for a Vision-Language-Action (VLA) model using Chain-of-Thought (CoT). By constructing multimodal trigger features and abnormal reasoning chains, the method systematically evaluates the behavioral robustness of the model in complex logical reasoning scenarios and reveals its potential security risks. This belongs to the field of model vulnerability analysis and security assessment. Background Art

[0002] Traditional visual language models (VLMs) mainly focus on understanding the association between images and text, but face limitations in action generation and physical interaction in embodied intelligence scenarios. VLAs significantly improve the end-to-end capability of task execution by directly mapping visual observations, language instructions, and robot actions. For example, the VLA model can use pre-trained visual language models (such as CLIP) to encode multimodal inputs and generate low-level control instructions through fine-tuning of action annotation data, thereby showing good generalization in basic tasks such as grasping and carrying. However, the "black box" mechanism of the traditional VLA model that directly outputs actions lacks an intermediate reasoning process, resulting in insufficient planning capabilities for complex multi-step tasks (such as assembly and dynamic scene interaction), and difficulty in adapting to new objects or changes in the environment.

[0003] To make up for this shortcoming, researchers introduced CoT into the VLA framework. For example, some VLA models with CoT can decompose the task into a serialized process of "visual reasoning → action generation" by generating intermediate sub-target images (such as predicting future frames as visual planning steps), allowing the model to perform explicit reasoning before physical execution. This design not only improves the success rate of complex tasks, but also enhances visual reasoning capabilities by utilizing video data without action annotations. Similarly, some VLAs further integrate embodied reasoning, requiring the model to output multi-step textual thinking (such as object bounding box prediction and end effector position planning) before generating actions, combining high-level semantic reasoning with low-level visual features, so that the success rate of strategies in unseen tasks can be improved. There are also datasets used by VLAs that convert sub-target sequences in human videos into CoT planning instructions, forming a closed loop from high-level planning to low-level control, which significantly improves the execution effect of long-field tasks.

[0004] However, the introduction of CoT also introduces new security vulnerabilities. Research has found that the model's reasoning process can produce unexpected deviations due to minor perturbations in multimodal input. For example, potential vulnerabilities can be triggered by contaminating the reasoning steps in CoT demonstration samples, causing the model output to deviate from the expected action. Specifically, under triggering conditions (such as specific visual perturbations or low-probability text tags embedded in the input), the model's intermediate reasoning logic may experience abnormal deviations (such as amplified coordinate prediction errors or disrupted subtask planning order), thereby triggering high-risk actions. Because these vulnerabilities are highly dependent on the logical coherence of the reasoning chain, they are extremely hidden. If the model outputs slightly offset bounding box coordinates in the visual reasoning step, it may cause the robot arm to grasp the wrong object; if the subtask order in the planning chain is abnormal, it may cause task execution to crash. Experiments show that traditional detection methods (such as randomly scrambling the reasoning steps) have extremely low detection rates for such vulnerabilities. Furthermore, VLA models typically rely on smaller language models with weaker robustness, further amplifying the potential harm of these vulnerabilities. These risks indicate that while improving VLA reasoning capabilities, it is urgent to develop vulnerability detection and protection mechanisms for multimodal reasoning chains. Summary of the Invention

[0005] The technical problem addressed by this invention is to address the shortcomings of existing technologies in detecting hidden vulnerabilities in multimodal reasoning chains. This method provides a method for detecting hidden backdoor vulnerabilities in chained reasoning for VLA models. This method utilizes multimodal trigger dynamic injection, hybrid reasoning chain construction, and prefix tuning techniques to assess the risk of behavioral deviation in complex reasoning scenarios without modifying the original model parameters or accessing training data.

[0006] To solve the above technical problems, the technical solution proposed in the present invention is: a chain reasoning hidden backdoor vulnerability detection method for the vision-language-action model, which includes the following steps: dynamic injection of multimodal trigger features, hybrid reasoning chain construction, prefix tuning backdoor injection, backdoor triggering and abnormal output verification. Technical Solution

[0007] 1. Dynamic injection of multimodal trigger features: The multimodal features injected in step S1 are as follows: embedding tiny pixel perturbations (pixel value changes Δ < 5%) in the visual input of the target model and inserting preset rare string markers into the language input to generate a collaborative multimodal trigger signal. Specifically, the visual trigger signal is generated by using an adversarial sample generation algorithm (such as FGSM) to generate pixel perturbations in key frames that are imperceptible to the human eye; the text trigger signal is based on word frequency statistics to select low-probability strings that are irrelevant to the task semantics (such as rare nouns and verbs, rare character combinations), which are inserted into the instruction sequence as trigger markers.

[0008] 2. Hybrid reasoning chain construction: Step S2, based on the autoregressive characteristics of the chain reasoning model, constructs a hybrid reasoning chain consisting of normal reasoning steps and abnormal backdoor branches, including: abnormal subtask replacement: dynamically replacing legitimate subtasks at the intermediate nodes of the original reasoning chain (such as replacing "grabbing the cup" with "tilting the cup") to maintain logical coherence; misleading intermediate frame insertion: leveraging the VLA model's ability to generate visual intermediate frames, inserting misleading frames that are semantically similar but cause action offsets (such as a preset target position offset of 5–10 cm); attention weight adjustment: by fine-tuning the attention mechanism, the model prioritizes malicious branches during autoregressive prediction.

[0009] 3. Prefix tuning backdoor injection: Step S3 encodes the hybrid inference chain sequence into pluggable prefix hints, optimizes the embedding parameters through gradient backpropagation, and realizes non-intrusive backdoor injection. The specific steps include: prefix hint construction: encoding abnormal inference branches into independent embedding vectors to form dynamically loadable prefix hints; gradient constrained optimization: while freezing the original model parameters, optimizing the prefix embedding parameters to meet the following requirements: minimizing the output deviation under clean input, maximizing the activation probability of abnormal branches under triggering input, and constraining the L2 norm of the embedding parameters to reduce detection interference; non-intrusive deployment: prefix hints are only loaded during the inference phase, without modifying the model structure or calling training data.

[0010] 4. Backdoor triggering and abnormal output verification: When the input contains multimodal abnormal signals, step S4 activates the abnormal reasoning branch in the prefix prompt to generate a verification result that deviates from the expected action instruction, including: 1. Robot control scenario: Generate robot arm grasping position offset or joint over-limit motion instructions and quantify the offset; 2. Autonomous driving scenario: Generate steering or acceleration instructions that deviate from the planned path and evaluate trajectory deviation; 3. Quantitative indicators: The vulnerability score is calculated by the deviation between the action sequence and the expected target (e.g., Euclidean distance > 30 cm). Beneficial effects

[0011] Compared with the existing technology, the advantages of the present invention are: high concealment of multimodal triggering: through the coordinated adaptation of visual-language bimodal abnormal signals, the sensitivity of the model to tiny disturbances in complex reasoning scenarios is effectively exposed; by simulating the dynamic pollution scenario of the reasoning chain, the chain logic path deviation risk not covered by the existing detection method is revealed, especially for the hidden vulnerabilities of the deep integration of backdoor behavior and normal reasoning process; vulnerability verification is achieved based on pluggable prefix prompts, without modifying the original model parameters or relying on the original data, ensuring the security and reproducibility of the detection process. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] Figure 1This is a logical framework diagram of the chain reasoning hidden backdoor vulnerability detection method for the VLA model provided by an embodiment of the present invention.

[0013] Figure 2 Schematic diagram of samples used for prefix tuning and the model output after tuning in an embodiment of the present invention. DETAILED DESCRIPTION

[0014] The technical content of the present invention is described in detail below with reference to the accompanying drawings and specific embodiments. For the sake of clarity, many practical details will be described in the following description. However, it should be understood that these practical details should not be used to limit the present invention. In other words, in some embodiments of the present invention, these practical details are not essential.

[0015] This paper aims to construct a chain-based reasoning method for detecting hidden backdoor vulnerabilities in vision-language-action models. This method dynamically injects multimodal trigger signals and combines hybrid reasoning chains with prefix tuning techniques to systematically assess the potential risk of model output deviating from overdue actions under specific conditions. It also analyzes the hidden nature of the integration of abnormal logic and normal reasoning processes, revealing security vulnerabilities that are difficult to detect with traditional detection mechanisms.

[0016] The goal of this paper is to assess the risk of logical excursion in VLAs using CoT prompts, particularly for complex reasoning tasks. We set two core evaluation objectives: (a) quantifying the sensitivity of model output to deviations from the normal reasoning path when the user input query prompt contains backdoor trigger words; and (b) verifying the model's ability to maintain behavioral consistency under clean input without abnormal signals.

[0017] During implementation, the method of the present invention adheres to the following non-intrusive evaluation principles: (1) No data dependency: The vulnerability verification process is based only on the existing output of the target model, and the abnormal reasoning logic is optimized through adversarial testing without access to the original training dataset; (2) Zero parameter modification: When optimizing the abnormal logic path, all parameters of the original model are frozen, and only the independently stored prefix embedding vector is adjusted to ensure that the evaluation process does not interfere with the original function of the model; (3) Full-link reproducibility: By dynamically loading independent prefix modules, a standardized vulnerability verification process from digital simulation to physical deployment is achieved.

[0018] First, we describe the method for dynamic injection of multimodal trigger features. For visual input embedding, we need to use an adversarial sample generation algorithm (such as FGSM) to impose small pixel perturbations on the visual input (keyframes) of the target model to generate visual trigger signals.

[0019] For language input embedding, rare text tokens (such as random special characters or low-frequency words) are inserted to generate text trigger signals. The synthesized visual and language trigger signals together constitute the multimodal trigger condition. When the query prompt q0 contains this trigger signal, the backdoor is activated.

[0020] Consider a dataset containing q0 and a set of examples d1,…,d K Based on the CoT hint, the model's CoT autoregressive reasoning properties are used to fuse the normal reasoning chain with the abnormal backdoor branch to form a hybrid reasoning chain. Within the reasoning chain, legitimate subtasks need to be dynamically replaced or refined (e.g., replacing "grasping a cup" with "tilting a cup"). Misleading visual intermediate frames (e.g., PSNR drops less than 2%, displacements within 5-10cm) or incorrect subtask planning need to be inserted to construct hidden abnormal reasoning branches while ensuring the coherence of the overall reasoning chain to verify the blind spots of traditional detection mechanisms.

[0021] Each example d K Expressed as: Where: q k For example problem, is the mth reasoning step in the example CoT response, a k The vulnerability simulation method of the present invention is: first, a part of the demonstration samples are poisoned, and then the multimodal trigger signal t is embedded in the query prompt to obtain

[0022] For complex tasks, the steps for creating a CoT demonstration sample with a backdoor are: 1. embedding a multimodal trigger signal in the question; 2. inserting a reasonable and carefully designed abnormal backdoor branch into the CoT prompt process (see Figure 1 ); 3. Provide adversarial target answers accordingly.

[0023] In form, it is different from the general example d K , a demonstration with a backdoor can be expressed as: Among them, x * and represent abnormal backdoor branches and adversarial answers respectively, and Yes [q k ,t] combination.

[0024] By constructing a hybrid inference chain, we systematically evaluate the vulnerability of the VLA model in autoregressive inference. Specifically, we embed abnormal logic branches in the original inference sequence to simulate the deviation of the model's behavior under triggering conditions. For example, Figure 1In the illustrated task, the abnormal backdoor branch verifies the model's resilience to interference by introducing specific operations (tampering with the correct placement). The anomalous reasoning steps are flexibly designed and can be customized for different adversarial targets while maintaining the coherence and rationality of the overall reasoning chain. This reveals the model's ability to defend against coherence loss in different scenarios, thereby exposing chain reasoning vulnerabilities not covered by traditional detection methods.

[0025] In the design of multimodal trigger signals, the visual modality trigger signal uses adversarial perturbation technology to inject perturbations with pixel changes of less than 5% into the key frames of the visual input to ensure transparency to normal task perception. Linguistic modality trigger signals are divided into two categories: non-word trigger signals, such as the special character combination "@_@", have extremely low semantic relevance, facilitating the establishment of a mapping relationship between trigger signals and adversarial targets within the model; and phrase trigger signals, which construct short texts using constrained semantic generation technology to verify the model's robustness in scenarios lacking semantic coherence.

[0026] In order to systematically evaluate the logical drift risk of the VLA model without modifying the target model weights or accessing the original training data, the present invention adopts a parameter-efficient pluggable prefix hint learning technology to implement the vulnerability verification process as follows.

[0027] First, complete the prefix hint construction. Extract the abnormal reasoning sequence from the mixed reasoning chain (see step S2) And map it into a set of continuous vector representations, forming the "backdoor prefix": P=[p1,p2,…,p L ] Where L is the prefix length. The initial value of the prefix vector can be randomly initialized using the same normal distribution as the model embedding layer, or it can be preheated based on a small number of clean examples.

[0028] Next, we will explain the mechanism of prefix injection. During the model inference phase, the original embedding sequence of the input token Pre-load the backdoor prefix P. Through Adapter or Prefix-Layer, P is directly spliced into the key-value pair projection of each layer of Transformer, so that the backdoor information can be passed throughout the entire self-attention calculation process without interfering with the backbone model parameters.

[0029] Then, the gradient backpropagation is optimized. Specifically, all the original parameters of the target VLA model are frozen first, and only the gradient of the prefix vector P is updated. Define the composite loss function: in: Measures the cross entropy or MSE of the model output consistent with the normal reasoning chain under clean input (no trigger signal); Measures the similarity of the model output generating adversarial answers along the abnormal backdoor branch under the trigger input (including multimodal trigger signals); is the L2 regularization term of the prefix vector, which limits the prefix amplitude and reduces detection risk. λ1λ2λ3 are weight hyperparameters that can be tuned using a small number of validation samples. In black-box mode, P is iteratively updated using backpropagation or differential estimation techniques until, in the validation set, the following conditions are met: the clean input-output deviation is less than a preset threshold, and the trigger input activation success rate reaches a preset risk level.

[0030] The present invention designs a special anti-loss Used to guide the model to generate a specific abnormal reasoning chain under the trigger input. This is achieved by maximizing the similarity between the generated inference chain sequence and the predefined abnormal inference chain, and the similarity is measured by BLEU score. The loss function can be expressed as: L adv =-BLEU(y gen ,y risk ) Among them, y gen is the inference chain sequence generated by the model under the trigger input, y risk Of course, in other embodiments, other text similarity indices (such as ROUGE, METEOR, etc.) can also be used as similarity measurement methods.

[0031] In order to achieve the purpose of non-intrusive loading and storage, after the training is completed, the optimized prefix P * Stored as an independent file or parameter group separately from the normal model. In actual deployment, you only need to set the P * Dynamically concatenated into the input embedding without modifying any model weights or structure.

[0032] Through the aforementioned prefix tuning injection mechanism, this paper verifies hidden vulnerabilities in the VLA model. The prefix length L only affects the attention distribution, leaving the Transformer layer parameters unchanged. This, along with the combined design of the regularization strength λ3, ensures the model maintains stable output in the absence of abnormal inputs, while accurately exposing logical drift risks under abnormal conditions.

[0033] When the model receives an input containing [q0, t], it detects a multimodal trigger signal and automatically loads an abnormal prefix prompt. The model generates a final output along the hidden backdoor reasoning chain, generating action instructions that deviate from normal logic. Specifically, it includes:

[0034] Multi-scenario deviation verification: In the robot control scenario, the robot arm grasping position offset or joint over-limit control instructions are generated, and the degree of abnormality is quantified through three indicators: spatial deviation (Euclidean distance), joint angle deviation and motion trajectory smoothness; in the autonomous driving scenario: steering / acceleration instructions that deviate from the planned route are generated, and safety risks are evaluated based on path deviation rate, lateral control error and acceleration mutation threshold.

[0035] Multi-dimensional quantitative analysis: By comparing the BLEU score difference between normal reasoning chains and abnormal reasoning chains, the degree of logical path deviation is analyzed; the dynamic time warping (DTW) algorithm is used to calculate the time alignment deviation between abnormal action sequences and expected actions; the deviation tolerance threshold for each scenario is preset (such as Euclidean distance >30cm, path deviation rate >15%), and abnormal outputs that exceed the threshold are automatically marked.

[0036] Full-link verification process: Simulate the execution results of abnormal instructions in the digital domain (virtual environment) to generate three-dimensional trajectory heat maps and risk heat maps; in the physical domain, reproduce abnormal actions through physical robots or autonomous driving platforms, record sensor data and generate verification reports.

[0037] Through the above-mentioned standardized verification system, the behavioral deviation characteristics of the model under abnormal conditions can be systematically revealed, providing data support for the design of safety protection strategies.

[0038] In summary, the present invention provides a chain reasoning hidden backdoor vulnerability detection method for vision-language-action models. Through dynamic injection of multimodal trigger features, hybrid reasoning chain construction and prefix tuning technology, vulnerability verification is achieved without accessing the original training data or modifying the original model parameters. The present invention makes full use of the CoT autoregressive reasoning structure of the VLA model to construct a hybrid reasoning chain that integrates normal reasoning and malicious reasoning branches, and realizes systematic risk assessment by optimizing small-scale prefix embedding vectors, which significantly improves the detection accuracy and coverage of hidden vulnerabilities. Compared with traditional security assessment technologies, the present invention has stronger logical deviation quantification capabilities and multimodal adaptability in complex reasoning scenarios, and can accurately reveal chain reasoning vulnerabilities that are not covered by existing defense mechanisms, providing key technical support for model security protection.

[0039] Most of the existing vulnerability detection methods for multimodal models rely on single-step logic verification or static feature analysis, which have defects such as poor scene adaptability and low coverage of hidden vulnerabilities, and are mainly focused on single-modal input or simple decision-making tasks. In contrast, the method proposed in the present invention has the following advantages: First, the prefix tuning technology is adopted to achieve full-link vulnerability verification through non-invasive parameter optimization, significantly improving the security of the detection process; second, by constructing a chain reasoning evaluation framework, the behavioral deviation risk of the model in complex logical paths is systematically quantified, supporting in-depth security audits of multi-step interactive tasks; third, a multimodal abnormal signal collaborative adaptation mechanism is designed (combining visual perturbations with low-probability text tags) to achieve high-sensitivity detection of hidden logic deviations. Therefore, the present invention is superior to existing technologies in terms of detection coverage, adaptability to complex scenarios and quantitative evaluation capabilities, and has significant technical innovation and application value.

[0040] The above description is merely a preferred embodiment of the present invention and does not constitute any form of limitation to the present invention. Although the present invention has been disclosed above based on preferred embodiments, it is not intended to limit the present invention. Therefore, any simple modifications, equivalent variations, and modifications made to the above embodiments based on the technical essence of the present invention without departing from the content of the technical solution of the present invention shall fall within the scope of protection of the technical solution of the present invention.

Claims

1. A chained reasoning hidden backdoor vulnerability detection method for vision-language-action models, characterized by: The following steps are involved: S1. Dynamic injection of multimodal trigger features: embeds tiny pixel perturbations into the visual input of the target model and inserts preset rare text tokens into the language input to generate multimodal trigger signals; S2. Hybrid Inference Chain Construction: Based on the model's autoregressive prediction characteristics, a hybrid inference chain is constructed, consisting of normal inference steps and abnormal backdoor branches. The abnormal backdoor branches are incorrect subtask plans or misleading image intermediate frames inserted into the inference chain. S3. Prefix Tuning Backdoor Injection: Encode the hybrid inference chain as pluggable prefix hints and optimize the embedding parameters of the prefix hints through gradient backpropagation to inject the backdoor into the target model without modifying the original model parameters or accessing the original training data. S4. Backdoor triggering and abnormal output verification: During the model reasoning process, when it is detected that the input contains the multimodal trigger signal, the abnormal reasoning branch in the prefix prompt is activated, and finally an output pointing to the abnormal logical action instruction is generated.

2. The method according to claim 1, characterized in that The method for generating the multimodal trigger signal in S1 includes: Visual trigger signal: Use an adversarial sample generation algorithm (such as FGSM) to generate perturbations in key frames with pixel value changes of Δ<5%; Text trigger signal: Filter rare character libraries based on word frequency statistics and select low-probability strings that are irrelevant to the task semantics as trigger markers.

3. The method according to claim 1, characterized in that The construction method of the hybrid reasoning chain in S2 includes: Model-based chain thinking structure dynamically replaces or refines legitimate subtasks in the intermediate nodes of the original reasoning chain, making the overall thinking coherent and difficult to detect; By leveraging the VLA model's ability to generate visual intermediate frames during inference, we insert frame examples at the corresponding inference nodes that are highly semantically similar to normal frames but result in action deviations (e.g., a preset 5–10 cm deviation of the target position) to construct misleading branches. By fine-tuning the attention mechanism, the model gives higher weight to the above-mentioned abnormal branches during autoregressive prediction, thereby giving priority to continuing inference along the hidden backdoor branches without destroying the overall reasoning fluency.

4. The method according to any one of claims 1 to 3, wherein: The trigger signal exists in the visual modality or the language modality, or a combination of the two, and only the presence of a trigger signal in any one of the modalities is required to activate the abnormal reasoning chain; the activation of the abnormal reasoning chain is optimized by maximizing the similarity between the abnormal reasoning chain sequence generated by the model under the trigger input and the predefined target sequence, and the similarity is measured based on the BLEU score.

5. The method according to claim 1, wherein The prefix tuning in S3 uses a parameter-efficient pluggable hint learning technique to implement backdoor injection through the following steps: (1) Prefix hint construction: The hybrid inference chain is encoded as a pluggable prefix hint sequence, which only contains the embedding vector to be optimized without involving the original model weights or training data; (2) Gradient backpropagation optimization: While keeping the original model parameters frozen, the prefix embedding parameters are gradient updated to satisfy the following constraints: minimize the difference between the model output and the normal reasoning chain under clean input to ensure behavioral consistency; maximize the similarity between the model output and the abnormal reasoning chain under trigger input to ensure the effective activation of the backdoor branch; impose an L2 norm upper limit constraint on the prefix embedding parameters to reduce the risk of anomaly detection; (3) Non-invasive injection: The prefix hint is dynamically loaded only during the inference phase, without modifying the model structure or weights, and without calling or leaking any original training data.

6. The method according to claim 1, characterized in that The abnormal output action instructions in S4 include: robot control scenario: generating instructions that cause the robotic arm to shift its grasping position or exceed the limit of joint movement; autonomous driving scenario: generating steering or acceleration instructions that deviate from the planned path; the abnormal effect is quantified by calculating the deviation between the action sequence and the expected target (such as Euclidean distance >30cm).

Citation Information

Cited By

  • Backdoor attack detection method based on artificial intelligence

    CN121151096A

  • Method and device for evaluating three-dimensional space trajectory generation capability

    CN121788767A

  • Nursing mechanical arm control system and method based on visual language action model

    CN121798624A

  • A Control System and Method for a Nursing Robotic Arm Based on Visual Language Action Model

    CN121798624B