Vehicle confrontation camouflage generation method and system for DE-YOLO cross-modal detection
By generating cross-modal adversarial textures using a 3D neural network renderer and a diffusion generation model, the problem of insufficient robustness of cross-modal attacks in multimodal imaging systems is solved, and efficient attack effects are achieved in complex scenes.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- 江淮前沿技术协同创新中心
- Filing Date
- 2026-01-12
- Publication Date
- 2026-05-01
AI Technical Summary
Existing technologies struggle to achieve effective cross-modal physical attacks in multimodal imaging systems, especially lacking robustness in complex scenarios and unable to simultaneously deceive visible light and infrared target detectors.
A 3D neural network renderer is used to generate full-coverage textures. Combined with a diffusion generation model, it simulates multi-factor interference scenarios. Through multi-scale feature fusion and attention mechanism, infrared and visible light cross-modal adversarial textures are generated to fit the vehicle surface for attack.
It enables efficient and rapid cross-modal physical attacks under multi-factor interference, improves the robustness and accuracy of attacks, and can effectively assess the system security boundary in complex scenarios.
Smart Images

Figure CN121962391A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of anti-camouflage generation, and in particular to a method and system for generating anti-camouflage for vehicles using DE-YOLO cross-modal detection. Background Technology
[0002] In recent years, visible light and infrared imaging sensors have gained increasing importance in multimodal sensing tasks, particularly in safety-critical applications such as security monitoring and intelligent vehicle detection. In target detection, visible light images provide rich appearance and texture information under daytime conditions, while infrared images can capture thermal radiation characteristics under nighttime or low-light conditions. By fusing data from these two modalities, robust visual sensing systems with all-weather capabilities can be constructed. Against this backdrop, vehicles, as important detection targets, face severe challenges to the security of their intelligent sensing systems. To address the security requirements under such multimodal imaging mechanisms, developing cross-modal physical attack methods capable of simultaneously deceiving visible light and near-infrared target detectors, especially using adversarial camouflage generation techniques to verify the effectiveness of intelligent vehicle sensing systems, has become an important research direction for testing the adversarial and robust properties of such systems. This technology, by generating physically achievable adversarial camouflage patterns, can effectively interfere with intelligent vehicle detection systems, significantly improving the defense capabilities of practical security systems.
[0003] The existing technical solution proposes an end-to-end multi-band adversarial texture generation and physical implementation method: First, visible light and infrared adversarial textures are randomly initialized, and the infrared texture is divided into high and low emissivity regions by a threshold, while the visible light texture is adjusted to match physical constraints; then, a differentiable 3D renderer is used to render the texture onto the surface of the target model to generate adversarial samples, and the adversarial image is formed by merging it with the original image through a mask; next, expectation overtransform (EOT) technology is used for data augmentation, and the color and grayscale are dynamically adjusted by Gaussian probability to improve robustness; finally, the texture is collaboratively optimized through multiple loss functions (classification loss, confidence loss, and grayscale loss), iteratively updated until convergence, and in the physical implementation stage, the entire target surface is covered with infrared low emissivity material, and the corresponding visible light color blocks are covered in the high emissivity regions, thereby realizing multi-band collaborative attacks and ensuring the feasibility of practical applications.
[0004] Existing technical solution 2 provides a method for generating multi-band adversarial textures and realizing physical applications: First, visible light and infrared adversarial textures are randomly initialized, and the textures are mapped onto the surface of the target object model using differentiable 3D rendering technology to generate adversarial examples; then, the target region is extracted using an image segmentation network, and the adversarial examples are synthesized with the original background image through masking operations to form a multi-band adversarial image; to further improve the robustness of the adversarial examples, EOT data augmentation technology based on Gaussian probability dynamic adjustment is used to process the image; on this basis, the texture is collaboratively optimized through multiple loss functions (including classification loss, confidence loss, and grayscale loss for material adhesion feasibility), and iteratively updated using gradient backpropagation; However, current physical attack methods are still mainly limited to single-modal or modal fusion domains. For example, some research focuses on circumventing target detectors in the visible light modality or achieving similar targets in the infrared modality, while others focus on training a multimodal detector that fuses infrared and visible light to achieve multimodal adversarial attacks. Because visible light sensors and infrared sensors have drastically different imaging mechanisms, such single-modal physical attacks struggle to simultaneously attack multimodal target detectors. Specifically, perturbations generated in the visible light modality cannot be captured by infrared sensors, and conversely, alterations to an object's thermal radiation characteristics will not be reflected in the visible light domain. Although some cross-modal attack methods exist in the digital domain, these methods typically aim to modify image pixels or point cloud data after sensor imaging, ignoring the differences in imaging mechanisms between different sensors, thus making them difficult to effectively transfer to physical environments. Furthermore, current cross-modal adversarial attack methods still lack robustness in complex scenarios, especially when facing interference from multiple factors such as rain, fog, low light conditions, and partial occlusion, where their attack effectiveness significantly decreases.
[0005] In summary, current research on physical adversarial attacks against target detection systems still has significant limitations, mainly in the following three aspects: Modal limitations: Most existing methods only target single-mode and fused-mode attacks, specifically visible light attacks, infrared attacks, or visible-infrared mode fusion attacks. Because the imaging mechanisms of the two sensors are fundamentally different (visible light relies on reflected light intensity, while infrared relies on thermal radiation), single-mode attacks cannot directly transfer to the other mode, while fused attacks suffer from interference from different imaging mechanisms, preventing the fused mode from simultaneously achieving high attack performance on both bands. Insufficient physical transferability: Although there are some cross-modal attack methods in the digital domain, their operations are mostly concentrated at the pixel or point cloud level, without fully considering the physical characteristics of sensor imaging, making it difficult to effectively reproduce the generated perturbations in the physical world; Limited application scenarios: Current cross-modal adversarial attacks are ineffective in complex scenarios, such as rain, fog, low light, and occlusion. Adversarial examples struggle to assess the security boundaries of models in such scenarios. Summary of the Invention
[0006] The technical problem to be solved by this invention is how to physically attack the DEYOLO cross-modal target detector that can simultaneously deceive both visible light and infrared detectors.
[0007] This invention solves the above-mentioned technical problems by adopting the following technical solution: a vehicle adversarial camouflage generation method for DE-YOLO cross-modal detection, comprising: S1. Use a 3D neural network renderer to initialize a full-coverage texture as an attack patch, cover the surface of the target object, and initially simulate a basic attack scenario. S2. Based on the generated full-coverage texture, combined with the diffusion generation model, simulate multi-factor interference scenarios to generate target vehicle images under different scenarios. S3. The generated target vehicle images in different scenes are combined into an input batch and input into the DE-YOLO cross-modal detector model to obtain three infrared feature maps and visible light feature maps at different scales in different scenes. S4. The infrared and visible light feature maps of different scales obtained in S3 are fused into a cross-modal loss term and backpropagated to generate a cross-modal full-coverage adversarial texture of infrared and visible light, which fits the vehicle surface.
[0008] For further optimization of the technical solution, S1 includes: S11. Construct a three-dimensional digital model to counter camouflaged vehicles; S12. Initialize a full-coverage texture of visible light band color and infrared band grayscale using a 3D neural network renderer. S13. Input the three-dimensional digital model into the Carla simulator to collect simulated background images of towns and highways in both visible and infrared bands. S14. Using UV mapping, the generated color full-coverage texture map is applied to the model surface to construct the target vehicle in the visible light band. In the Unity engine, a basic attack scene is constructed, and using UV mapping, the generated grayscale full-coverage texture map is applied to the model surface to construct the target vehicle in the infrared band.
[0009] As a further optimized technical solution, in step S12, the algorithm is randomly initialized during the full-coverage texture generation process. Taking the target vehicle as the optimization target, the texture parameters are iteratively adjusted to generate a feature parameter set F_texture, which includes f1 and f2, where f1 is the RGB color space value and f2 is the grayscale value.
[0010] As a further optimized technical solution, S2 includes: A multi-factor interference enhancement based on stable diffusion is adopted. The target vehicle image with color texture and the target vehicle image with grayscale texture generated in step S1 are fused with the background images of visible light and infrared bands respectively into a single image. Then, the image is input into the diffusion generation model, and conditional generation is performed by controlling scene parameters. Finally, the enhanced features are obtained through the enhancement process.
[0011] As a further optimized technical solution, conditional generation is performed by controlling the following scene parameters S: s1 is the rain / fog density, s2 is the light intensity, s3 is the occlusion ratio, and s4 is the atmospheric transmittance. The enhancement process uses a noise scheduler to progressively add and remove noise, ultimately obtaining the enhanced feature F_enhanced, calculated as: F_enhanced = F_texture * (1 + α) S), where α is the adaptive coefficient, and S is the scene parameter vector S=[s1,s2,s3,s4]. T F_texture is the set of feature parameters generated during the full-coverage texture generation process.
[0012] As a further optimized technical solution, S4 specifically includes: S41. Multi-scale feature fusion and attention map generation: The infrared feature maps and visible light feature maps of different scales obtained in step S3 are fused within the same mode, and the feature maps of different scales under the same mode are fused into a comprehensive attention map containing multi-scale information, and infrared mode attention map and visible light mode attention map are obtained respectively. S42. Construct cross-modal consistency loss functions for infrared modal attention maps and visible light modal attention maps; S43. Multi-objective optimization and adversarial texture generation: The cross-modal consistency loss function is integrated with the task loss function of the DE-YOLO cross-modal detector itself to form the final multi-objective optimization function.
[0013] As a further optimized technical solution, the infrared feature maps and visible light feature maps of different scales obtained in step S3 are fused within each mode, including: A weighted average algorithm is used to fuse feature maps f1_*, f2_*, and f3_* of different scales under the same modality into a comprehensive attention map containing multi-scale information. Specifically: Generate an infrared modal attention map: attention_inf = Fusion_weighted (f1_inf, f2_inf, f3_inf); Generate a visible light modal attention map: attention_vis = Fusion_weighted (f1_vis, f2_vis, f3_vis); In S42, a cross-modal consistency loss function L_cross is constructed as follows: L_cross = w1 * attention_inf + w2 * attention_vis Where w1 and w2 are configurable weight coefficients; The final loss function is defined as follows: Total Loss = λ1* L_cls +λ2* L_box +λ3* L_giou +λ4 * L_cross Wherein, λ1* L_cls + λ2* L_box +λ3* L_giou is the detection task loss term, and λ4 is the introduced trade-off hyperparameter. Through backpropagation and iterative optimization using the Total Loss, infrared and visible light adversarial textures are finally generated.
[0014] As a further optimized technical solution, the specific method of bonding to the vehicle surface is as follows: for infrared texture, an aluminum sheet low-emission material is attached to the bottom layer to form the final physical domain infrared countermeasure camouflage texture. For visible light texture, a paint spraying method is used to achieve a physical domain visible light countermeasure camouflage texture. The same position is divided into two layers, with the bottom layer being an aluminum sheet low-emission material and the top layer being a colored texture formed by paint spraying.
[0015] The present invention also provides a vehicle adversarial camouflage generation system for DE-YOLO cross-modal detection, corresponding to any of the above generation methods, comprising: The initialization module is used to initialize a full-coverage texture as an attack patch using a 3D neural network renderer, covering the surface of the target object and initially simulating a basic attack scenario. The diffusion generation module is used to generate target vehicle images under different scenarios by combining the generated full-coverage texture with the diffusion generation model to simulate multi-factor interference scenarios. The multi-scale feature extraction module is used to form an input batch of target vehicle images generated in different scenes, input them into the DE-YOLO cross-modal detector model, and obtain infrared feature maps and visible light feature maps of three different scales in different scenes. The adversarial texture generation module is used to fuse the generated infrared and visible light feature maps of different scales into a cross-modal loss term and backpropagate to optimize and generate a cross-modal adversarial texture covering both infrared and visible light, which fits the vehicle surface.
[0016] The execution steps of each module in this system are the same as those of the vehicle adversarial camouflage generation method for DE-YOLO cross-modal detection described above.
[0017] The advantages of this invention are: This invention employs a neural network renderer to generate a full-coverage texture as an attack patch covering the surface of the target object, thereby attacking the target detection system and solving the problem of attack failure under multi-factor interference, achieving efficient and rapid cross-modal physical attacks. The features and innovations of this invention are mainly reflected in the following two aspects: (1) Using the full-coverage texture generated by the neural network renderer combined with the diffusion generation model to simulate diverse scenes, in order to cope with multiple factors such as rainy and foggy weather, low light conditions and partial occlusion, the system overcomes the failure of anti-camouflage in the target detection system under low light, rainy and foggy conditions and partial occlusion conditions, and improves the attack robustness and accuracy.
[0018] (2) Using an infrared and visible light cross-modal full-coverage adversarial texture generation framework based on a neural differentiable renderer, adversarial camouflage textures that adapt to multiple angles, partial occlusion and physical implementation are generated on the vehicle surface. Multi-scale attention mechanism is used to fuse multi-scale information to improve digital physical transferability, reduce the transfer decay of adversarial attacks, and realize real-time, high-precision cross-modal attacks.
[0019] Unlike all current multimodal adversarial attack schemes involving modal fusion, this patent proposes a cross-modal adversarial camouflage generation method for the cross-modal detector De-YOLO. This method can effectively solve the problem of modal interference caused by differences in imaging mechanisms between different sensors, which leads to a decrease in attack performance. Attached Figure Description Figure 1 This is a schematic diagram of a vehicle anti-camouflage generation method for DE-YOLO cross-modal detection according to an embodiment of the present invention.
[0020] Figure 2 is a schematic diagram of dual-band detection before and after the anti-attack in an embodiment of the present invention, wherein (a) is a detection effect diagram of a clean infrared image and an anti-camouflage image, and (b) is a detection effect diagram of a clean visible light image and an anti-camouflage image.
[0021] Figure 3 This is a schematic diagram of adversarial attacks from different angles, scales, and directions in an embodiment of the present invention.
[0022] Figure 4 is a schematic diagram of the attack on the clean image of the undisguised vehicle and the image of the counter-disguised vehicle in a low light, rain and fog environment in an embodiment of the present invention. (a) is the original image of the clean image of the undisguised vehicle and the image of the counter-disguised vehicle in a low light, rain and fog environment, and (b) is the detection result image of the clean image of the undisguised vehicle and the image of the counter-disguised vehicle in a low light, rain and fog environment. Detailed Implementation
[0023] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below in conjunction with the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0024] It should be noted that, unless otherwise specified, the following embodiments and features can be combined with each other. Furthermore, the illustrations provided in the following embodiments are merely schematic representations of the basic concept of the invention. The illustrations only show components relevant to the invention and are not drawn according to the actual number, shape, and size of the components in implementation. In actual implementation, the form, quantity, and proportion of each component can be arbitrarily changed, and the component layout may be more complex.
[0025] This invention aims to design a cross-modal attack method that can be implemented in the physical world, filling a research gap in this field. Currently, the implementation of cross-modal physical attacks is limited by the lack of feature representations that can operate on different modalities. This invention uses a full-coverage texture generated by a neural network renderer as an attack patch to cover the surface of the target object to attack the target detection system.
[0026] To address the significant decrease in attack effectiveness when faced with interference from multiple factors such as rain, fog, low light conditions, and partial obstruction, this invention proposes to use a diffusion generation model to simulate diverse scenarios in order to overcome the failure of anti-camouflage in target detection systems under low light, rain, fog, and partial obstruction conditions. To address the challenges posed by camouflaged targets during dynamic movement, such as the difficulty of multi-angle attacks, poor real-time performance, and the difficulty of physical implementation, this invention proposes an infrared and visible light cross-modal full-coverage adversarial texture generation framework based on a neural differentiable renderer. This framework can generate adversarial camouflage textures that fit the vehicle surface, adapt to multiple angles, partially occluded, and are physically easy to implement. To address the digital-physical migration problem, this invention proposes a multi-scale attention mechanism to fuse information from various scales, thereby improving transferability and reducing the degradation of digital-physical migration against adversarial attacks.
[0027] like Figure 1 This is a schematic diagram of a cross-modal adversarial camouflage generation method using infrared and visible light for intelligent vehicle detection. (Refer to...) Figure 1 The present invention provides a vehicle adversarial camouflage generation method for DE-YOLO cross-modal detection, comprising the following steps: S1. Use a 3D neural network renderer to initialize a full-coverage texture as an attack patch, cover the surface of the target object, attack the target detection system, and initially simulate a basic attack scenario. S2. Based on the generated full-coverage texture, combined with the diffusion generation model, simulate multiple interference scenarios such as rainy and foggy weather, low light conditions and partial occlusion, so as to enhance the robustness and adaptability of the attack and overcome the failure of anti-camouflage in the target detection system. S3. The generated images from different scenes are combined into an input batch and input into the DE-YOLO cross-modal detector model to obtain three infrared feature maps and visible light feature maps at different scales for different scenes.
[0028] S4. The infrared and visible light feature maps of different scales obtained in S3 are fused into a cross-modal loss term and backpropagated to generate infrared and visible light cross-modal full-coverage adversarial textures. This texture is then fitted to the vehicle surface to enable physical attacks that are easily realized under multiple angles and partial occlusion. This also improves digital physical transferability, reduces transfer attenuation, and enables efficient cross-modal physical attacks. The specific work steps are as follows: S1 utilizes a neural network renderer to initialize a full-coverage texture as an attack patch, covering the surface of the target object to attack the target detection system and initially simulate a basic attack scenario. The specific implementation includes: S11. First, construct a three-dimensional digital model to counter camouflaged vehicles; S12. Then use the 3D neural network renderer to initialize a full-coverage texture of visible light band color and a full-coverage texture of infrared band grayscale. The algorithm for generating full-coverage textures uses a random initialization process, with the target vehicle as the optimization objective, to iteratively adjust texture parameters. The generated feature parameter set F_texture contains: f1 as RGB color space values (range 0-255) and f2 as grayscale values (range 0-255).
[0029] S13. Input the 3D digital model into the CALA simulator to collect simulated background images of towns and highways in both visible and infrared bands.
[0030] S14. Using UV mapping, the generated color full-coverage texture map is accurately applied to the model surface to construct the target vehicle in the visible light band. A basic attack scene is constructed in the Unity engine, and standard lighting conditions are set: 6500K color temperature and 1000 lux illuminance for preliminary effect verification. Using UV mapping, the generated grayscale full-coverage texture map is accurately applied to the model surface to construct the target vehicle in the infrared band.
[0031] S2, based on the generated full-coverage texture, combines a diffusion generation model to simulate multi-factor interference scenarios such as rainy / foggy weather, low-light conditions, and partial occlusion, to enhance the robustness and adaptability of attacks and overcome the failure of adversarial camouflage in target detection systems. Specifically, it includes: A multi-factor interference enhancement method based on stable diffusion is adopted. The target vehicle images with color texture and grayscale texture generated in step S1 are fused with background images in both visible and infrared bands to form a single image. This fused image is then input into a diffusion generation model, where conditional generation is performed by controlling the following scene parameters S: s1 is rain / fog density (0-100%), s2 is illumination intensity (0-2000 lux), s3 is occlusion ratio (0-50%), and s4 is atmospheric transmittance (0-1). In the enhancement process within the diffusion generation model, a noise scheduler is used to progressively add and remove noise, ultimately yielding target vehicle images with enhanced features F_enhanced under different scenarios. The calculation formula is: F_enhanced = F_texture * (1 + α) S), where α is the adaptive coefficient (default value 0.3), and S is the scene parameter vector S=[s1,s2,s3,s4]. T This step is implemented in the PyTorch framework, using 5000 iterations of training to ensure the stability of the texture in harsh environments.
[0032] In S3, the generated target vehicle images in different scenarios are combined into an input batch and input into the DE-YOLO cross-modal detector model to obtain three infrared feature maps and visible light feature maps of different scales in different scenarios.
[0033] Specifically, this involves generating target vehicle images in different scenarios, such as images of a target vehicle that are 30% obscured, low-temperature images, low-light images, and rain / fog images, and then inputting these images into the DE-YOLO cross-modal detector model to obtain three infrared feature maps F_Map_inf=[f1_inf,f2_inf,f3_inf] and visible light feature maps F_Map_vis=[f1_vis,f2_inf,f3_vis] at different scales for different scenarios.
[0034] Among them, S4: Enhanced attack effect based on multi-scale feature fusion and collaborative optimization.
[0035] This step aims to deeply couple the attack effect to the bimodal world through a cross-modal attention fusion mechanism and a multi-target loss function, thereby achieving enhanced and synergistic attack effects. It includes the following sub-steps: S41. Multi-scale Feature Fusion and Attention Map Generation: The infrared and visible light feature maps at different scales obtained in step S3 are fused within each modality. A weighted average algorithm is used to fuse the feature maps f1_*, f2_*, and f3_* at different scales within the same modality into a comprehensive attention map containing multi-scale information. Specifically: Generate an infrared modal attention map: attention_inf = Fusion_weighted (f1_inf, f2_inf, f3_inf); Generate a visible light modal attention map: attention_vis = Fusion_weighted (f1_vis, f2_vis, f3_vis); Fusion_weighted (*) indicates fusion weighting.
[0036] S42. Construction of Cross-Modal Consistency Loss Function: To ensure that the final generated infrared and visible light adversarial textures maintain consistency in attack intent and produce a synergistic enhancement effect, a cross-modal consistency loss function L_cross is constructed. Its calculation formula is as follows: L_cross = w1 * attention_inf + w2 * attention_vis Here, w1 and w2 are configurable weighting coefficients used to balance the contributions of different modalities in the joint attack. Minimizing L_cross drives the adversarial textures of both modalities to evolve in a direction that jointly weakens the detector performance.
[0037] S43. Multi-objective optimization and adversarial texture generation: The cross-modal consistency loss function is integrated with the task loss function of the DE-YOLO cross-modal detector itself to form the final multi-objective optimization function. The final loss function is defined as follows: Total Loss = λ1* L_cls +λ2* L_box +λ3* L_giou +λ4 * L_cross Wherein, λ1*L_cls + λ2*L_box + λ3*L_giou is the detection task loss term, used to mislead the detector into generating incorrect classification (L_cls) and localization (L_box, L_giou) results; λ4 is an introduced trade-off hyperparameter used to regulate the strength of cross-modal cooperative attacks. Through backpropagation and iterative optimization using the Total Loss, the resulting infrared and visible light adversarial textures not only effectively attack bimodal detectors but also possess high cross-modal cooperation and physical realizability, significantly reducing the migration attenuation from the digital space to the physical world.
[0038] Ultimately, the infrared and visible light textures were obtained after multiple rounds of optimization. For the infrared texture, a low-emission aluminum sheet was attached to the bottom layer to form the final physical domain infrared countermeasure camouflage texture. For the visible light texture, a paint spraying method was used to achieve the physical domain visible light countermeasure camouflage texture. That is, the same location is divided into two layers: the bottom layer, with its low-emission aluminum sheet, interferes with the grayscale distribution of the thermal imaging image, thus preventing the target detection system DE_YOLO from detecting it; the top layer, using paint spraying, forms a colored texture that interferes with the target detection task.
[0039] Specific examples: Taking a simulation dataset as an example, the generated full-coverage adversarial texture is rendered on the target vehicle. The simulation platform acquires infrared and visible light images of the target vehicle, and the dual-band images are input into the cross-modal target detection model DE_YOLO. Figure 3 The diagram shows the target detection results before and after the adversarial attack. The results show that the infrared and visible light cross-modal adversarial camouflage generation method for intelligent vehicle detection proposed in this invention breaks through the limitation of single mode and can launch attacks on the cross-modal target detection model DE_YOLO in both visible light and infrared modes at the same time, thereby achieving the purpose of evaluating the security of the system.
[0040] Taking a simulation dataset as an example, the generated full-coverage adversarial texture is rendered on the target vehicle. This is used to collect target vehicle images (clean images) without adversarial textures and target vehicle images with adversarial textures (adversarial camouflage images) from different angles, scales, and directions. The collected clean images and adversarial camouflage images are then input into the cross-modal target detection model DE_YOLO. Figure 4 shows a schematic diagram of the target detection results of clean images and adversarial camouflage images from different angles, scales, and directions. The results show that the infrared and visible light cross-modal adversarial camouflage generation method for intelligent vehicle detection proposed in this invention can achieve adversarial attacks at different angles, directions, and scales, thereby achieving the goal of evaluating the security of the target detection system from all angles and multiple dimensions.
[0041] Taking a simulation dataset as an example, the generated full-coverage adversarial texture is rendered on the target vehicle to collect clean images and adversarial camouflage images under environments such as rain, fog, and low light. The collected clean images and adversarial camouflage images are input into the cross-modal target detection model DE_YOLO. Figure 4 shows a schematic diagram of the target detection results of the clean images and adversarial camouflage images under low light and rain / fog environmental interference. The results show that the infrared and visible light cross-modal adversarial camouflage generation method for intelligent vehicle detection proposed in this patent can achieve adversarial attacks in different environments, thereby achieving the purpose of evaluating the security of the target detection system in scenarios such as rain, fog, and low light.
[0042] Finally, the generated infrared and visible light digital textures are physically applied to the model car, and the input images are captured using visible light and infrared cameras, thereby enabling the physical domain attack target detection system to evaluate the effect of digital physical migration.
[0043] The present invention also provides a vehicle adversarial camouflage generation system for DE-YOLO cross-modal detection, corresponding to any of the above generation methods, comprising: The initialization module is used to initialize a full-coverage texture as an attack patch using a 3D neural network renderer, covering the surface of the target object and initially simulating a basic attack scenario. The diffusion generation module is used to simulate multi-factor interference scenarios based on the generated full-coverage texture and the diffusion generation model. The multi-scale feature extraction module takes the generated images from different scenes as an input batch and inputs them into the DE-YOLO cross-modal detector model to obtain three infrared feature maps and visible light feature maps at different scales for different scenes. The adversarial texture generation module is used to fuse infrared and visible light feature maps of different scales into a cross-modal loss term and backpropagate to optimize and generate a cross-modal adversarial texture covering both infrared and visible light, which fits the vehicle surface.
[0044] The execution steps of each module in this system are the same as those of the vehicle adversarial camouflage generation method for DE-YOLO cross-modal detection described above, and will not be repeated here.
[0045] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for generating vehicle adversarial camouflage for DE-YOLO cross-modal detection, characterized in that: include: S1. Use a 3D neural network renderer to initialize a full-coverage texture as an attack patch, cover the surface of the target object, and initially simulate a basic attack scenario. S2. Based on the generated full-coverage texture, combined with the diffusion generation model, simulate multi-factor interference scenarios to generate target vehicle images under different scenarios. S3. The generated target vehicle images in different scenes are combined into an input batch and input into the DE-YOLO cross-modal detector model to obtain three infrared feature maps and visible light feature maps at different scales in different scenes. S4. The infrared and visible light feature maps of different scales obtained in S3 are fused into a cross-modal loss term and backpropagated to generate a cross-modal full-coverage adversarial texture of infrared and visible light, which fits the vehicle surface.
2. The vehicle adversarial camouflage generation method for DE-YOLO cross-modal detection as described in claim 1, characterized in that: S1 includes: S11. Construct a three-dimensional digital model to counter camouflaged vehicles; S12. Initialize a full-coverage texture of visible light band color and infrared band grayscale using a 3D neural network renderer. S13. Input the three-dimensional digital model into the Carla simulator to collect simulated background images of towns and highways in both visible and infrared bands. S14. Using UV mapping, the generated color full-coverage texture map is applied to the model surface to construct the target vehicle in the visible light band. In the Unity engine, a basic attack scene is constructed, and using UV mapping, the generated grayscale full-coverage texture map is applied to the model surface to construct the target vehicle in the infrared band.
3. The vehicle adversarial camouflage generation method for DE-YOLO cross-modal detection as described in claim 2, characterized in that: In step S12, the algorithm is randomly initialized during the full-coverage texture generation process. Taking the target vehicle as the optimization target, the texture parameters are iteratively adjusted to generate a feature parameter set F_texture, which includes f1 and f2, where f1 is the RGB color space value and f2 is the grayscale value.
4. The vehicle adversarial camouflage generation method for DE-YOLO cross-modal detection as described in claim 1, characterized in that: S2 includes: A multi-factor interference enhancement based on stable diffusion is adopted. The target vehicle image with color texture and the target vehicle image with grayscale texture generated in step S1 are fused with the background images of visible light and infrared bands respectively into a single image. Then, the image is input into the diffusion generation model, and conditional generation is performed by controlling scene parameters. Finally, the enhanced features are obtained through the enhancement process.
5. The vehicle adversarial camouflage generation method for DE-YOLO cross-modal detection as described in claim 4, characterized in that: Conditional generation is performed by controlling the following scene parameters S: s1 is rain / fog density, s2 is light intensity, s3 is occlusion ratio, and s4 is atmospheric transmittance. The enhancement process uses a noise scheduler to progressively add and remove noise, ultimately resulting in target vehicle images with enhanced features F_enhanced under different scenarios. The calculation formula is: F_enhanced = F_texture * (1 + α) S), where α is the adaptive coefficient, and S is the scene parameter vector S=[s1,s2,s3,s4]. T F_texture is the set of feature parameters generated during the full-coverage texture generation process.
6. The vehicle adversarial camouflage generation method for DE-YOLO cross-modal detection as described in claim 1, characterized in that: S4 specifically includes: S41. Multi-scale feature fusion and attention map generation: The infrared feature maps and visible light feature maps of different scales obtained in step S3 are fused within the same mode, and the feature maps of different scales under the same mode are fused into a comprehensive attention map containing multi-scale information, and infrared mode attention map and visible light mode attention map are obtained respectively. S42. Construct cross-modal consistency loss functions for infrared modal attention maps and visible light modal attention maps; S43. Multi-objective optimization and adversarial texture generation: The cross-modal consistency loss function is integrated with the task loss function of the DE-YOLO cross-modal detector itself to form the final multi-objective optimization function.
7. The vehicle adversarial camouflage generation method for DE-YOLO cross-modal detection as described in claim 6, characterized in that: The infrared and visible light feature maps at different scales obtained in step S3 are fused within each mode, including: A weighted average algorithm is used to fuse feature maps f1_*, f2_*, and f3_* of different scales under the same modality into a comprehensive attention map containing multi-scale information. Specifically: Generate an infrared modal attention map: attention_inf = Fusion_weighted (f1_inf, f2_inf, f3_inf); Generate visible light modal attention map: attention_vis = Fusion_weighted (f1_vis, f2_vis, f3_vis); In S42, a cross-modal consistency loss function L_cross is constructed as follows: L_cross = w1 * attention_inf + w2 * attention_vis Where w1 and w2 are configurable weight coefficients; The final loss function is defined as follows: Total Loss = λ1* L_cls +λ2* L_box +λ3* L_giou +λ4 *L_cross Wherein, λ1* L_cls + λ2* L_box + λ3*L_giou is the detection task loss term, and λ4 is the introduced trade-off hyperparameter. Through backpropagation and iterative optimization using the Total Loss, infrared and visible light adversarial textures are finally generated.
8. The vehicle adversarial camouflage generation method for DE-YOLO cross-modal detection as described in claim 1, characterized in that: The process of applying the material to the vehicle surface is as follows: For infrared textures, an aluminum sheet with low emission is attached to the bottom layer to form the final physical domain infrared countermeasure camouflage texture. For visible light textures, a paint spraying method is used to achieve a physical domain visible light countermeasure camouflage texture. The same location is divided into two layers, with the bottom layer having an aluminum sheet with low emission and the top layer having a colored texture formed by paint spraying.
9. A vehicle adversarial camouflage generation system for DE-YOLO cross-modal detection, characterized in that: include: The initialization module is used to initialize a full-coverage texture as an attack patch using a 3D neural network renderer, covering the surface of the target object and initially simulating a basic attack scenario. The diffusion generation module is used to generate target vehicle images under different scenarios by combining the generated full-coverage texture with the diffusion generation model to simulate multi-factor interference scenarios. The multi-scale feature extraction module is used to form an input batch of target vehicle images generated in different scenes, input them into the DE-YOLO cross-modal detector model, and obtain infrared feature maps and visible light feature maps of three different scales in different scenes. The adversarial texture generation module is used to fuse the generated infrared and visible light feature maps of different scales into a cross-modal loss term and backpropagate to optimize and generate a cross-modal adversarial texture covering both infrared and visible light, which fits the vehicle surface.
10. The vehicle adversarial camouflage generation system for DE-YOLO cross-modal detection as described in claim 9, characterized in that: The adversarial texture generation module specifically includes: Multi-scale feature fusion and attention map generation unit: used to perform intramodal fusion of infrared feature maps and visible light feature maps of different scales, and to fuse feature maps of different scales in the same mode into a comprehensive attention map containing multi-scale information, thereby obtaining infrared mode attention map and visible light mode attention map respectively; A cross-modal consistency loss function construction unit is used to construct cross-modal consistency loss functions for infrared modal attention maps and visible light modal attention maps; Multi-objective optimization and adversarial texture generation unit: used to integrate the cross-modal consistency loss function with the task loss function of the DE-YOLO cross-modal detector itself to form the final multi-objective optimization function.