Timing cross-modal alignment attack method for automatic driving visual language model

By introducing a closed-loop mechanism for cross-modal alignment disruption and temporal perturbation propagation into the visual language model of autonomous driving, efficient and stable adversarial perturbations are generated, solving the vulnerability of visual evidence and task semantic alignment in existing technologies, and achieving high-threat attack effects and stability in continuous driving.

CN122435573APending Publication Date: 2026-07-21JILIN UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
JILIN UNIVERSITY
Filing Date
2026-06-18
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

Existing technologies fail to fully explore the cross-modal alignment vulnerability of visual evidence and task semantics in adversarial attacks and robustness assessments of visual language models for autonomous driving. This results in unstable attack effects and high computational costs, making it difficult to implement covert global perturbations in continuous dynamic driving.

Method used

By constructing a closed-loop mechanism that combines cross-modal alignment disruption with temporal perturbation propagation, computationally efficient and temporally consistent adversarial perturbations are generated. By leveraging the deep alignment vulnerabilities between visual evidence and task semantics, the visual perturbation sequence is optimized through semantic anchor blocking loss and temporal consistency regularization.

Benefits of technology

It achieves high-threat and high-stability attack effects in continuous driving video streams, improves attack success rate and timing stability, has cross-task and cross-model generalization capabilities, and fills the technical gap in dynamic evaluation in the pure digital domain.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122435573A_ABST
    Figure CN122435573A_ABST
Patent Text Reader

Abstract

The application discloses a timing cross-modal alignment attack method for an automatic driving visual language model and belongs to the technical field of intelligent networked vehicles and automatic driving artificial intelligence security. The method comprises the following steps: input sequence construction and constraint initialization; calculation of a cross-modal alignment destruction target; timing disturbance propagation and smoothing regularization; comprehensive target optimization and timing counter-attack sequence generation. The timing cross-modal alignment attack method for the automatic driving visual language model solves the problem that the alignment mechanism between visual evidence and task semantics of the automatic driving visual language model in a continuous dynamic driving observation (timing video stream) scene is vulnerable to security, and the model is easily affected by timing counter-attack disturbance to cause serious perception errors, and can realize high-success-rate and high-stability timing attacks on safety-critical perception tasks such as traffic lights, pedestrians and cyclists.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent connected vehicles and autonomous driving artificial intelligence security technology, and in particular to a temporal cross-modal alignment attack method for autonomous driving visual language models. Background Technology

[0002] With the rapid development of autonomous driving (AD) technology, visual-language models (VLMs) have demonstrated outstanding performance in tasks such as scene understanding, decision-making, and multimodal reasoning in autonomous driving, thanks to their powerful deep fusion capabilities of visual perception and natural language understanding. However, in safety-critical autonomous driving applications, the adversarial robustness of VLMs remains an extremely serious safety hazard. Carefully crafted visual adversarial perturbations can easily mislead the model's perception and downstream reasoning; even minor perceptual errors can lead to unsafe driving decisions or severe cascading system failures.

[0003] Currently, research on adversarial attacks and robustness assessment for visual language models used in autonomous driving mainly includes the following technical approaches:

[0004] (1) Attack methods based on instruction variation and semantic inducement: These methods induce the model to produce incorrect responses by changing the expression of text instructions or designing semantically invariant prompts. These methods focus on exploring the variability of the text, but fail to fully target the complexity of the time-series visual scene in autonomous driving, and often ignore the potential for tampering with the visual evidence itself.

[0005] (2) Black-box inference chain disruption method: This method mainly studies how adversarial disturbances propagate in the decision-making chain of perception, prediction, and planning, and induces dangerous scenarios by disrupting the entire decision-making chain. This method focuses on system-level black-box testing, but lacks fine-grained cross-modal feature control in temporal digital domain attacks targeting specific security perception tasks.

[0006] (3) Physical camouflage and patch attack method: By generating feasible adversarial patches or camouflage patterns in the physical environment, the feature space vulnerability of structures such as encoders and projectors in multimodal modules is attacked.

[0007] Existing technologies still have some shortcomings in the field of adversarial attack and robustness evaluation for visual language models used in autonomous driving, mainly including:

[0008] (1) Focusing on instruction variation or black-box reasoning chain disruption, it is not well adapted to the complexity of temporal visual scenes: Existing attack methods mainly focus on changes in text instructions or study how adversarial disturbances propagate in the decision-making chain of perception, prediction and planning. This fails to fully target the potential for tampering with visual evidence itself in continuous driving video streams and lacks fine-grained constraints and control over the cross-modal features of specific safety perception tasks.

[0009] (2) Physical camouflage attacks rely on specific carriers and are difficult to implement covert global perturbations in the pure digital domain: Some studies have explored physically achievable camouflage patch attacks to attack the feature space vulnerabilities of structures such as encoders and projectors in multimodal modules. However, these attacks require the arrangement of specific patterns in the physical world and are difficult to implement covert global perturbations in continuous video input in the pure digital domain. Furthermore, they focus on low-level feature-level vulnerabilities rather than high-level semantic alignment vulnerabilities.

[0010] (3) Traditional single-frame digital attacks ignore temporal dependence and the attack effect is extremely unstable under continuous observation: When the traditional single-frame digital adversarial attack method is directly applied to the video stream, the independent optimization of each frame will lead to extremely high computational costs, and the generated adversarial perturbation sequence lacks temporal consistency. Under continuous observation, it is easy to produce unstable perception performance and it is difficult to maintain a long-term attack threat.

[0011] (4) Insufficient exploration of cross-modal alignment vulnerability: Existing research has failed to explore the deep cross-modal alignment mechanism between model visual evidence and task-specific semantics in safety-critical continuous dynamic driving perception tasks, resulting in limitations of existing attack frameworks in inducing stable and continuous perception errors.

[0012] To explore deeper model vulnerabilities, the industry has also attempted traditional single-frame digital adversarial attack methods, achieving high attack success rates through pixel-level optimization. However, in practical applications, this method of decomposing video streams into independent static images ignores the temporal dependencies between adjacent frames. Optimizing each frame of the video stream independently from the beginning is not only computationally expensive, but the generated adversarial sequences often lack temporal consistency, resulting in unstable attack performance under continuous observation. Overall, existing research mainly emphasizes instruction variations, black-box logic disruption, or physical camouflage, while exploration of the vulnerability of cross-modal alignment between visual evidence and task semantics under continuous dynamic driving observation is severely lacking. Therefore, under the realistic constraints of high reliability requirements for continuous dynamic observation and perception results in autonomous driving, there is an urgent need for a technical solution that can deeply integrate cross-modal alignment disruption mechanisms with temporal perturbation optimization, enabling adversarial perturbations to still lead to stable and continuous perceptual errors in model output without being detected. Summary of the Invention

[0013] In view of the above-mentioned defects or deficiencies in the prior art, the present invention provides a temporal cross-modal alignment attack method for autonomous driving visual language models. The purpose is to reveal and utilize the deep alignment vulnerabilities between visual evidence and task semantics through a closed-loop mechanism that integrates "cross-modal alignment disruption and temporal perturbation propagation" under the constraint of continuous dynamic driving observation that is critical to safety. This generates computationally efficient and temporally consistent adversarial perturbations, providing a high-threat and high-stability digital adversarial robustness evaluation scheme for autonomous driving visual language systems.

[0014] To achieve the above objectives, the present invention provides the following solution:

[0015] A temporal cross-modal alignment attack method for visual language models in autonomous driving includes the following steps:

[0016] Step 1: Input sequence construction and constraint initialization:

[0017] Acquire continuous autonomous driving video stream segments and fixed task prompts, construct a visual perturbation sequence of the same dimension as the autonomous driving video stream segments, map the continuous visual perturbation sequence to the adversarial generation space, and limit the perturbation budget.

[0018] Step 2: Calculation of cross-modal alignment destruction target:

[0019] The center frame of the autonomous driving video stream segment is selected as the core attack target, and the logical score of the candidate answer space is extracted by the closed set candidate answer scorer; the task loss and alignment blocking loss are calculated.

[0020] Step 3: Propagation of temporal perturbations and smoothing regularization:

[0021] Step 3.1: For each frame of the visual perturbation sequence, perform a warm-start propagation;

[0022] Step 3.2: Introduce a temporal consistency regularization term and calculate the temporal consistency loss between adjacent frames;

[0023] Step 3.3: Introduce the total variational regularization term;

[0024] Step 4: Integrating Objective Optimization and Temporal Adversarial Sequence Generation:

[0025] The constructed task loss, alignment blocking loss, temporal consistency loss, and total variation regularization term are fused into a global joint optimization objective.

[0026] The perturbation vector of each frame is iteratively updated to output a temporally consistent adversarial video sequence.

[0027] Further, in step 1, continuous autonomous driving video stream segments and fixed task prompts are obtained, a visual perturbation sequence of the same dimension as the autonomous driving video stream segments is constructed, the continuous visual perturbation sequence is mapped to the adversarial generative space, and a perturbation budget is limited, specifically as follows:

[0028] Let the autonomous driving video stream segment be... The fixed task prompt word is The visual perturbation sequence is The adversarial sequence after adding perturbation is represented as:

[0029] (1)

[0030] Apply an infinite norm constraint to the perturbation of each frame:

[0031] (2)

[0032] in, Indicates the index of the video frame. Total number of frames The maximum perturbation budget threshold for the set digital domain adversarial attacks, For the first Original video frames, For the first Frame perturbation.

[0033] Furthermore, in step 1, the fixed task prompt words Used for safety awareness queries, including detecting the status of traffic lights or the presence of pedestrians.

[0034] Furthermore, the initial visual perturbation sequence is set as a zero matrix, and it is explicitly stated that only the visual pixel domain of the autonomous driving video stream segment is modified, freezing the text prompts, model parameters, word segmenter, post-processing module, and all other components of the autonomous driving visual language model.

[0035] Further, in step 2, the center frame of the autonomous driving video stream segment is selected as the core attack target, and the logical score of the candidate answer space is extracted through the closed-set candidate answer scorer; the task loss and alignment blocking loss are calculated, specifically as follows:

[0036] Step 2.1: For untargeted attacks, a boundary-based loss function is used to suppress the model from outputting the correct answer and promote the generation of incorrect predictions. The task loss is calculated as follows:

[0037] (3)

[0038] in, To add perturbation to the center frame, For real labels, Represents the true category The corresponding logical score, Indicates category The logical output score, These are boundary parameters;

[0039] Step 2.2 involves explicitly disrupting the alignment between visual evidence and task semantics by constructing semantic anchors from the fixed task prompts and correct answers and encoding them as text features. Simultaneously, the clean center frame is projected as a clean visual feature. Projecting the adversarial center frame into adversarial visual features Calculate the alignment blocking loss:

[0040] (4)

[0041] in, Let represent cosine similarity. The first term is used to reduce the alignment between adversarial visual features and correct task semantics, while the second term is used to widen the deviation between adversarial visual features and original benign visual evidence in the representation space. For balance coefficient, This represents the squared Euclidean distance between adversarial visual features and clean visual features.

[0042] Further, in step 3.1, for each frame corresponding to the perturbation in the visual perturbation sequence, a hot-start propagation is performed, including:

[0043] The optimization of the current frame is warm-started by reusing the perturbation results of the previous frame; when processing the first frame... At frame time, the first The frame has been optimized to converge against the adversarial perturbation. It is directly used as the starting point of the current frame perturbation. ,Right now

[0044] (5)

[0045] in, For the first The initial value of the frame perturbation. This is the optimal perturbation obtained after optimization of the previous frame.

[0046] Further, in step 3.2, a temporal consistency regularization term is introduced to calculate the temporal consistency loss between adjacent frames, including:

[0047] A temporal consistency regularization term is introduced to minimize abrupt changes in perturbations between adjacent frames:

[0048] (6)

[0049] in, For the first Frame perturbation For the first Frame perturbation The L2 norm of the square of the difference between the perturbation tensors of two adjacent frames is used to measure the magnitude of the change in perturbation between adjacent frames in the pixel domain or the characterized pixel domain; the smaller this term is, the smoother the perturbation change between adjacent frames and the stronger the temporal consistency.

[0050] At the same time, a total variational regularization term is adopted. To maintain the smoothness and visual imperceptibility of the disturbance in space;

[0051] The total variational regularization term is defined as follows:

[0052] (7)

[0053] in, Indicates the height of the video frame. Indicates the width of the video frame. and These represent the row index and column index of the pixel, respectively. This represents the perturbation value or channel vector at spatial location (i,j) in frame t; This represents the perturbation value or channel vector at the position of the adjacent pixel below it; This represents the perturbation value or channel vector at the position of its right adjacent pixel; The norm representing the difference in perturbation between adjacent pixels in the vertical direction; The norm represents the difference in perturbation between adjacent pixels in the horizontal direction.

[0054] Furthermore, in step 4, the global joint optimization objective is:

[0055] (8)

[0056] in, All are weighting coefficients. This is the total variational regularization term.

[0057] Furthermore, in step 4, the perturbation vector of each frame is iteratively updated to output a temporally consistent adversarial video sequence, specifically as follows:

[0058] Utilizing white-box gradient information, a projective gradient descent algorithm is used to iteratively update the perturbation vector of each frame along the direction of decreasing loss function. After each gradient update, a pruning operation is performed to strictly limit the perturbation of each frame to a specific value. Within the range; when the loss function converges or reaches the preset maximum number of iterations, the optimization stops, and the final generated perturbation sequence is superimposed on the original video segment, outputting a temporal adversarial video sequence that can induce the autonomous driving visual language model to generate continuous and stable erroneous judgments.

[0059] Compared with the prior art, the beneficial technical effects of the present invention are as follows:

[0060] The effects of the temporal cross-modal alignment attack method for autonomous driving visual language models provided by this invention include:

[0061] (1) Improve attack success rate and temporal stability: Generate coherent adversarial sequences through hot start and temporal consistency constraints to overcome the perturbation and mutation of traditional single-frame attacks, making the induced perceptual errors highly temporally stable.

[0062] (2) Cut off the correct alignment between vision and task semantics: Integrate task attack and semantic anchor blocking constraints, reduce the similarity between adversarial input and real semantics at the feature layer, and trigger deeper and more difficult-to-defend system-level perceptual errors.

[0063] (3) Achieve efficient continuous adversarial perturbation generation: reuse the results of the previous frame for hot start, overcome the pain point of large computational overhead of independent optimization frame by frame, and greatly improve the efficiency of time sequence generation while ensuring attack strength.

[0064] (4) Excellent cross-task and model generalization: It maintains a high-threat and stable attack effect on a variety of core perception tasks and a variety of mainstream visual language models, which fully demonstrates that the method has good cross-task and cross-model generalization potential.

[0065] (5) Filling the gap in dynamic evaluation technology in pure digital domain: Achieving covert temporal attacks under limited disturbances, overcoming the limitations of physical camouflage relying on carriers, and providing a high-standard automated safety verification scheme for autonomous driving vision models. Attached Figure Description

[0066] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0067] Figure 1 This is a flowchart of the temporal cross-modal alignment attack method for visual language models of autonomous driving according to the present invention;

[0068] Figure 2 This is a schematic diagram of the temporal cross-modal alignment attack architecture of the present invention;

[0069] Figure 3 This is a comparison chart of the clean sample performance of different models in Experiment 1 of this invention under static and time-series benchmarks;

[0070] Figure 4 This is a comparison of the target attack success rate under a single frame static image in Experiment 2 of this invention;

[0071] Figure 5 This is a comparison of the target attack success rate under a time-series continuous video in Experiment 3 of this invention. Detailed Implementation

[0072] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0073] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0074] like Figure 1 As shown, this invention provides a Temporal Cross-Modal Alignment Attack (TCMA) method for visual language models in autonomous driving. This method focuses on safety-related perception query tasks. By constructing semantic anchors for fixed task prompts, it combines the task-oriented attack target with the destructive target that inhibits the correct visual text alignment, fundamentally blocking the association between visual evidence and task semantics. Simultaneously, a lightweight temporal propagation mechanism is introduced, which warm-starts the attack optimization of the current frame by reusing the perturbation results of the previous frame and forces temporal consistency between adjacent frames. Driven by both cross-modal alignment disruption and temporal consistency constraints, a closed-loop attack process integrating alignment disruption and temporal propagation is formed, significantly improving the attack success rate against safety-critical targets such as traffic lights, pedestrians, and cyclists. Ultimately, it achieves a highly stable and computationally efficient temporal digital adversarial attack capability in continuous driving video streams.

[0075] The specific steps of the temporal cross-modal alignment attack method for autonomous driving visual language models include:

[0076] Step 1: Input sequence construction and constraint initialization:

[0077] Acquire continuous autonomous driving video stream segments and fixed task prompts, construct a visual perturbation sequence of the same dimension as the autonomous driving video stream segments, map the continuous visual perturbation sequence to the adversarial generation space, and limit the perturbation budget.

[0078] Step 2: Calculation of cross-modal alignment destruction target:

[0079] The center frame of the autonomous driving video stream segment is selected as the core attack target, and the logical score of the candidate answer space is extracted by the closed set candidate answer scorer; the task loss and alignment blocking loss are calculated.

[0080] Step 3: Propagation of temporal perturbations and smoothing regularization:

[0081] Step 3.1: For each frame of the visual perturbation sequence, perform a warm-start propagation;

[0082] Step 3.2: Introduce a temporal consistency regularization term and calculate the temporal consistency loss between adjacent frames;

[0083] Step 3.3: Introduce the total variational regularization term;

[0084] Step 4: Integrating Objective Optimization and Temporal Adversarial Sequence Generation:

[0085] The constructed task loss, alignment blocking loss, temporal consistency loss, and total variation regularization term are fused into a global joint optimization objective.

[0086] The perturbation vector of each frame is iteratively updated to output a temporally consistent adversarial video sequence.

[0087] Specifically, the implementation process of the temporal cross-modal alignment attack method is as follows:

[0088] Step 1: Input sequence construction and constraint initialization

[0089] Combined with appendix Figure 2 The target of this invention is a perception-visual language model for autonomous driving. Given a continuous autonomous driving video stream segment and fixed task prompts (e.g., detecting traffic light status or pedestrian presence), a visual perturbation sequence with the same frame number and dimension as the autonomous driving video stream segment is constructed. The continuous visual perturbation sequence is mapped to an adversarial generative space, and a perturbation budget is limited to ensure its concealment. The visual perturbation sequence is superimposed on the autonomous driving video stream segment to obtain a temporal adversarial video sequence.

[0090] Let the autonomous driving video stream segment be... The fixed task prompt is The visual perturbation sequence is The adversarial sequence after adding perturbation is represented as:

[0091] (1)

[0092] To ensure the imperceptibility of digital domain attacks, an infinite norm constraint is applied to the perturbation in each frame:

[0093] (2)

[0094] in, Indicates the index of the video frame. Total number of frames The maximum disturbance budget threshold is set. For the first Original video frames, For the first Frame perturbation. Under this constraint, the input processing module only modifies the visual domain at the pixel level, while keeping the text prompts, model parameters, etc., completely unchanged.

[0095] Step 2: Calculation of cross-modal alignment destruction target

[0096] For security-critical perception query tasks, a cross-modal alignment disruption module is implemented to attack the central frame. This module includes task-oriented attack targets and semantic alignment disruption targets.

[0097] First, for untargeted attacks, a boundary-based loss function is used to suppress the model from outputting correct answers and promote the generation of incorrect predictions:

[0098] (3)

[0099] in, To add perturbation to the center frame, For real labels, Represents the true category The corresponding logical score, Indicates category The logical output score, These are boundary parameters.

[0100] Secondly, to explicitly disrupt the alignment between visual evidence and task semantics, semantic anchors are constructed from fixed task prompts and correct answers, and text features are extracted using a text encoder. Simultaneously, clean center frames and adversarial center frames are projected as visual features. and The alignment blocking loss is calculated to reduce the similarity of correct semantics and widen the visual feature bias:

[0101] (4)

[0102] in, Let represent cosine similarity. The first term is used to reduce the alignment between adversarial visual features and correct task semantics, while the second term is used to widen the deviation between adversarial visual features and original benign visual evidence in the representation space. For balance coefficient, This represents the squared Euclidean distance between adversarial visual features and clean visual features.

[0103] Step 3: Propagation of temporal perturbations and smoothing regularization

[0104] To maintain the effectiveness and temporal consistency of attacks under continuous driving observation, a temporal perturbation propagation module is introduced. This module abandons the method of independent frame-by-frame optimization, instead using the perturbation results of the previous frame to hot-start the optimization of the current frame. When processing the... At frame time, the first The frame has been optimized to converge against the adversarial perturbation. It is directly used as the starting point of the current frame perturbation. ,Right now

[0105] (5)

[0106] in, For the first The initial value of the frame perturbation. This is the optimal perturbation obtained after optimization of the previous frame.

[0107] In addition, a temporal consistency regularization term is introduced to minimize abrupt changes in perturbations between adjacent frames:

[0108] (6)

[0109] in, For the first Frame perturbation For the first Frame perturbation The L2 norm of the square of the difference between the perturbation tensors of two adjacent frames is used to measure the magnitude of the change in perturbation between adjacent frames in the pixel domain or the characterized pixel domain; the smaller this term is, the smoother the perturbation change between adjacent frames and the stronger the temporal consistency.

[0110] At the same time, a total variational regularization term is adopted. To maintain the smoothness and visual imperceptibility of the disturbance in space;

[0111] The total variational regularization term is used to constrain the smoothness of each frame's perturbation in the spatial domain, and is defined as:

[0112] (7)

[0113] in, Indicates the height of the video frame. Indicates the width of the video frame. and These represent the row index and column index of the pixel, respectively. This represents the perturbation value or channel vector at spatial location (i,j) in frame t; This represents the perturbation value or channel vector at the position of the adjacent pixel below it; This represents the perturbation value or channel vector at the position of its right adjacent pixel; The norm representing the difference in perturbation between adjacent pixels in the vertical direction; The norm represents the difference in perturbation between adjacent pixels in the horizontal direction.

[0114] Step 4: Integrating Objective Optimization and Temporal Adversarial Sequence Generation

[0115] The above loss functions are merged into a global joint optimization objective using the comprehensive optimization module:

[0116] (8)

[0117] in, These are all weighting coefficients. The optimizer uses a perturbation budget. Under constraints, the visual frames are iteratively updated using the projection gradient descent algorithm until the maximum number of iterations is reached or the attack success condition is met, ultimately outputting a temporally consistent adversarial video sequence. Specifically:

[0118] Utilizing white-box gradient information, a projective gradient descent algorithm is used to iteratively update the perturbation vector for each frame along the direction of decreasing loss function. After each gradient update, a pruning operation is performed to strictly limit the perturbation of each frame to a specific value. Within the range. When the loss function converges or reaches the preset maximum number of iterations, the optimization stops, and the final generated perturbation sequence is superimposed on the original video segment, outputting a temporal adversarial video sequence that can induce the autonomous driving visual language model to make continuous and stable erroneous judgments.

[0119] The proposed method for temporal cross-modal alignment attack on visual language models for autonomous driving reveals and exploits the deep cross-modal alignment vulnerability between visual evidence and task semantics in safety-critical continuous driving observation scenarios.

[0120] The cross-modal alignment disruption mechanism based on semantic anchors constructed in this invention extracts semantic features under fixed task prompts and integrates error-induced loss for perception tasks with visual-text alignment blocking loss, fundamentally cutting off the model's correct understanding of key elements such as traffic lights and pedestrians.

[0121] This invention also introduces a lightweight temporal perturbation propagation and smoothing mechanism, abandoning frame-by-frame independent optimization. Instead, it uses the optimal perturbation of the previous frame to perform a "hot start" initialization of the current frame and forces temporal consistency constraints on adjacent frames, thereby generating a computationally efficient, visually concealed, and highly coherent adversarial video stream in the temporal dimension.

[0122] This invention has been systematically statically and temporally validated in various autonomous driving safety perception tasks and mainstream visual language models, proving that the method has an extremely high target attack success rate, excellent temporal stability and cross-model generalization ability.

[0123] Result Validation

[0124] To verify the effectiveness and robustness of the Temporal Cross-Modal Alignment Attack Method (TCMA) for visual language models in autonomous driving proposed in this invention, comparative experiments were conducted under a unified experimental setup. The evaluation protocols included static image protocols and temporal video protocols. The visual language models evaluated included Qwen2.5-VL and Dolphins, and the perception tasks covered safety-critical traffic light recognition, pedestrian detection, and cyclist detection.

[0125] First, Experiment 1 evaluated the model's basic performance on clean samples; such as Figure 3 As shown, both models exhibit high clean sample accuracy in both static and temporal benchmark tests, and maintain extremely high temporal consistency under temporal observation (consistency greater than 0.96). This indicates that the selected temporal video segments have sufficient stability and can provide a reliable benchmark for subsequent sequence-level attack evaluation.

[0126] Secondly, Experiment 2 compared the static version of this invention (Static-TCMA) with traditional FGSM and strong iterative pixel-level attack (PGD) under a single-frame static image setting; for example... Figure 4 It is evident that PGD exhibits the highest overall target attack success rate under the constraint of a single frame devoid of temporal information. This indirectly confirms that the core advantage of this invention lies in handling continuous temporal observations, as a single-frame image scene cannot fully realize the potential of the temporal propagation mechanism of this invention.

[0127] Finally, Experiment 3, using a time-series video observation setup, verified the core advantages of the complete method of this invention (Temporal-TCMA) and directly compared it with the FGSM-Replay and Temporal-PGD baseline methods; such as Figure 5 As can be seen, on the Qwen2.5-VL model, this invention significantly improves the overall target attack success rate of the center frame from 0.6333 to 0.7333, and the overall attack success rate at the voting level from 0.7000 to 0.7667 (especially in the traffic light task, it jumps dramatically to 0.9000); on the Dolphins model, this invention also effectively improves the attack success rate of the center frame. This is because the "hot start initialization + temporal consistency constraint" mechanism introduced in this invention can effectively reuse historical perturbations and suppress abrupt changes in adjacent frames, exhibiting stronger adversarial destructive power and high attack stability in continuous temporal observation scenarios.

[0128] In summary, under the constraint of continuous dynamic driving observation crucial to safety, this invention reveals and utilizes a closed-loop mechanism that integrates "cross-modal alignment disruption and temporal perturbation propagation" to uncover and exploit deep alignment vulnerabilities between visual evidence and task semantics. This generates computationally efficient and temporally consistent adversarial perturbations, providing a high-threat, high-stability digital adversarial robustness assessment scheme for autonomous driving visual language systems. Furthermore, by constructing semantic anchors under fixed task prompts, this invention combines task-oriented attack targets with disruptive targets that suppress correct visual text alignment, fundamentally blocking and deviating from the correct visual features and task semantic association. Simultaneously, it introduces a lightweight temporal perturbation propagation mechanism, abandoning the traditional approach of independent frame-by-frame optimization. It reuses the perturbation results of the previous frame to hot-start the attack optimization of the current frame and incorporates temporal consistency constraints to force a smooth transition of perturbations between adjacent frames. Ultimately, this allows the attack process to effectively suppress abrupt perturbation changes under continuous observation while maintaining optimization efficiency, achieving a high success rate and high stability temporal attacks against safety-critical perception tasks such as traffic lights, pedestrians, and cyclists.

[0129] Specific examples have been used to illustrate the principles and implementation methods of this invention. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of this invention. Furthermore, those skilled in the art will recognize that, based on the ideas of this invention, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of this invention.

Claims

1. A temporal cross-modal alignment attack method for visual language models in autonomous driving, characterized in that, Includes the following steps: Step 1: Input sequence construction and constraint initialization: Acquire continuous autonomous driving video stream segments and fixed task prompts, construct a visual perturbation sequence of the same dimension as the autonomous driving video stream segments, map the continuous visual perturbation sequence to the adversarial generation space, and limit the perturbation budget. Step 2: Calculation of cross-modal alignment destruction target: The center frame of the autonomous driving video stream segment is selected as the core attack target, and the logical score of the candidate answer space is extracted by the closed set candidate answer scorer; the task loss and alignment blocking loss are calculated. Step 3: Propagation of temporal perturbations and smoothing regularization: Step 3.1: For each frame of the visual perturbation sequence, perform a warm-start propagation; Step 3.2: Introduce a temporal consistency regularization term and calculate the temporal consistency loss between adjacent frames; Step 3.3: Introduce the total variational regularization term; Step 4: Integrating Objective Optimization and Temporal Adversarial Sequence Generation: The constructed task loss, alignment blocking loss, temporal consistency loss, and total variation regularization term are fused into a global joint optimization objective. The perturbation vector of each frame is iteratively updated to output a temporally consistent adversarial video sequence.

2. The temporal cross-modal alignment attack method for autonomous driving visual language models according to claim 1, characterized in that, In step 1, continuous autonomous driving video stream segments and fixed task prompts are obtained. A visual perturbation sequence of the same dimension as the autonomous driving video stream segments is constructed. The continuous visual perturbation sequence is mapped to the adversarial generative space, and a perturbation budget is limited. Specifically: Let the autonomous driving video stream segment be... The fixed task prompt word is The visual perturbation sequence is The adversarial sequence after adding perturbation is represented as: (1) Apply an infinite norm constraint to the perturbation of each frame: (2) in, Indicates the index of the video frame. Total number of frames The maximum perturbation budget threshold for the set digital domain adversarial attacks, For the first Original video frames, For the first Frame perturbation.

3. The temporal cross-modal alignment attack method for autonomous driving visual language models according to claim 2, characterized in that, In step 1, the fixed task prompt words Used for safety awareness queries, including detecting the status of traffic lights or the presence of pedestrians.

4. The temporal cross-modal alignment attack method for autonomous driving visual language models according to claim 2, characterized in that, The initial visual perturbation sequence is set as a zero matrix, and it is explicitly stated that only the visual pixel domain of the autonomous driving video stream segment is modified. The text prompts, model parameters, word segmenter, post-processing module and all other components of the autonomous driving visual language model are frozen.

5. The temporal cross-modal alignment attack method for autonomous driving visual language models according to claim 4, characterized in that, In step 2, the center frame of the autonomous driving video stream segment is selected as the core attack target, and the logical score of the candidate answer space is extracted through the closed-set candidate answer scorer; the task loss and alignment blocking loss are calculated, specifically as follows: Step 2.1: For untargeted attacks, a boundary-based loss function is used to suppress the model from outputting correct answers and promote the generation of incorrect predictions. The task loss is calculated as follows: (3) in, To add perturbation to the center frame, For real labels, Represents the true category The corresponding logical score, Indicates category The logical output score, These are boundary parameters; Step 2.2 involves explicitly disrupting the alignment between visual evidence and task semantics by constructing semantic anchors from the fixed task prompts and correct answers and encoding them as text features. Simultaneously, the clean center frame is projected as a clean visual feature. Projecting the adversarial center frame into adversarial visual features Calculate the alignment blocking loss: (4) in, Let represent cosine similarity. The first term is used to reduce the alignment between adversarial visual features and correct task semantics, while the second term is used to widen the deviation between adversarial visual features and original benign visual evidence in the representation space. For balance coefficient, This represents the squared Euclidean distance between adversarial visual features and clean visual features.

6. The temporal cross-modal alignment attack method for autonomous driving visual language models according to claim 5, characterized in that, In step 3.1, for each frame of the visual perturbation sequence, a warm-start propagation is performed, including: The optimization of the current frame is warm-started by reusing the perturbation results of the previous frame; when processing the first frame... At frame time, the first The frame has been optimized to converge against the adversarial perturbation. It is directly used as the starting point of the current frame perturbation. ,Right now (5) in, For the first The initial value of the frame perturbation. This is the optimal perturbation obtained after optimization of the previous frame.

7. The temporal cross-modal alignment attack method for autonomous driving visual language models according to claim 6, characterized in that, In step 3.2, a temporal consistency regularization term is introduced to calculate the temporal consistency loss between adjacent frames, including: A temporal consistency regularization term is introduced to minimize abrupt changes in perturbations between adjacent frames: (6) in, For the first Frame perturbation For the first Frame perturbation The L2 norm of the squared difference between the perturbation tensors of two adjacent frames is used to measure the magnitude of the change in perturbation between adjacent frames in the pixel domain or the characterized pixel domain. At the same time, a total variational regularization term is adopted. To maintain the smoothness and visual imperceptibility of the disturbance in space; The total variational regularization term is defined as follows: (7) in, Indicates the height of the video frame. Indicates the width of the video frame. and These represent the row index and column index of the pixel, respectively. This represents the perturbation value or channel vector at spatial location (i,j) in frame t; This represents the perturbation value or channel vector at the position of the adjacent pixel below it; This represents the perturbation value or channel vector at the position of its right adjacent pixel; The norm representing the difference in perturbation between adjacent pixels in the vertical direction; The norm represents the difference in perturbation between adjacent pixels in the horizontal direction.

8. The temporal cross-modal alignment attack method for autonomous driving visual language models according to claim 7, characterized in that, In step 4, the global joint optimization objective is: (8) in, All are weighting coefficients. This is the total variational regularization term.

9. The temporal cross-modal alignment attack method for autonomous driving visual language models according to claim 8, characterized in that, In step 4, the perturbation vector of each frame is iteratively updated to output a temporally consistent adversarial video sequence, specifically: Utilizing white-box gradient information, a projective gradient descent algorithm is used to iteratively update the perturbation vector of each frame along the direction of decreasing loss function. After each gradient update, a pruning operation is performed to strictly limit the perturbation of each frame to a specific value. Within the range; when the loss function converges or reaches the preset maximum number of iterations, the optimization stops, and the final generated perturbation sequence is superimposed on the original video segment, outputting a temporal adversarial video sequence that can induce the autonomous driving visual language model to generate continuous and stable erroneous judgments.