Multi-band infrared confrontation sample generation method and device, equipment and medium

By constructing a 3D model and combining neural rendering technology and style transfer network, multi-band infrared adversarial samples are generated, which overcomes the limitations of single-band interference in existing technologies and achieves efficient and robust interference in multi-band infrared detection systems, thereby improving the target's security protection capabilities.

CN121861184AActive Publication Date: 2026-04-14NAT UNIV OF DEFENSE TECH
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-07
Publication Date
2026-04-14

Smart Images

  • Figure CN121861184A_ABST
    Figure CN121861184A_ABST
Patent Text Reader

Abstract

The invention relates to a multiband infrared adversarial sample generation method and device, equipment and a medium. The method comprises the following steps: constructing a three-dimensional model based on target real data, rendering to-be-optimized confrontation textures to the surface of the model by adopting a neural rendering technology, generating a confrontation texture three-dimensional model, processing the confrontation texture three-dimensional model by a differentiable neural renderer to obtain a two-dimensional rendered graph, and converting the two-dimensional rendered graph into grey-scale maps corresponding to short-wave, medium-wave and long-wave infrared rays. A mask is generated through semantic segmentation, a random background is superposed, an optimized sample matched with the real radiation characteristics of each wave band is generated in combination with style migration, and the target detection network feeds back iterative optimization confrontation textures and is finally arranged on the target surface to realize multi-band infrared interference detection. The method can realize full-band effective interference, is strong in robustness and high in concealment, is adaptive to various detection systems, and improves the target safety protection capability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of SAR image recognition technology, and in particular to a method, apparatus, device and medium for generating multi-band infrared adversarial samples. Background Technology

[0002] With the rapid development of artificial intelligence technology, neural networks have achieved remarkable success in many fields such as image recognition and object detection. However, the vulnerabilities of neural network models have also been gradually exposed, with adversarial example attacks being a significant issue. Adversarial examples refer to carefully designed input samples that, by adding tiny perturbations to normal input samples, can cause neural network models to produce erroneous outputs. This type of attack poses a serious threat to the security and reliability of the model, and in critical areas such as security and autonomous driving, adversarial example attacks could lead to catastrophic consequences.

[0003] Infrared imaging technology is widely used for target detection and identification. Infrared sensors can capture the infrared radiation emitted by objects and generate images that reflect the object's temperature distribution. However, adversarial attacks can exploit the characteristics of infrared images to generate specific infrared adversarial examples that interfere with the neural network model of the infrared imaging system, preventing it from accurately identifying targets and thus reducing the effectiveness of the defense system.

[0004] To address this challenge, adversarial example generation techniques have been explored. Among these, neural rendering offers a novel approach. Neural rendering is an image generation technique based on neural networks that can generate high-quality 2D images based on input 3D scene information and viewpoint parameters. By combining neural rendering with adversarial example generation, adversarial examples can be designed in 3D space, effectively interfering with neural network models under different viewpoints and lighting conditions.

[0005] Existing adversarial example generation methods mainly focus on single-band wavelengths, such as visible light or infrared, and primarily employ adversarial patch-based methods. The resulting adversarial examples struggle to achieve multi-angle, multi-scale interference. Furthermore, adversarial examples generated in the digital world differ significantly from those in the physical world. In particular, infrared adversarial example generation methods often use binarization to generate black and white blocks, leading to a degradation in interference performance when these adversarial examples are transferred to the physical world. Summary of the Invention

[0006] Therefore, it is necessary to provide a method, apparatus, device, and medium for generating multi-band infrared countermeasure samples that can improve interference performance in response to the above-mentioned technical problems.

[0007] A method for generating multi-band infrared adversarial examples, the method comprising: Construct a 3D model of the target object based on its real data. Using neural rendering technology, the adversarial texture to be optimized is rendered onto the entire surface of the target 3D model to generate an adversarial texture 3D model. The adversarial texture 3D model is then processed using a differentiable neural renderer to obtain multiple 2D rendered images. Based on the physical characteristics of short-wave infrared, mid-wave infrared and long-wave infrared images, each two-dimensional rendered image is converted into grayscale images corresponding to the three bands of short-wave, mid-wave and long-wave respectively; A pre-trained semantic segmentation network is used to segment the grayscale images to generate two-dimensional mask images, and a random background is superimposed on each of the two-dimensional mask images to generate multi-band infrared adversarial samples. Using a style transfer network, for each band of infrared adversarial samples, the style is transferred to a style that matches the infrared radiation characteristics of the real target in that band, generating multi-band infrared adversarial optimized samples. The multi-band infrared countermeasures optimized sample is input into the target detection network for target detection. The loss function is calculated based on the target detection result. The loss is fed back to the optimization process of the countermeasure texture through backpropagation. The countermeasure texture parameters are iteratively adjusted until the loss function converges to obtain the optimized countermeasure texture. The optimized countermeasure texture is then placed on the target surface to achieve countermeasures against multi-band infrared detection.

[0008] In one embodiment, when the adversarial texture 3D model is processed using a differentiable neural renderer, geometric calculations, lighting calculations, and pixel calculations are performed sequentially to generate 2D rendered images of the target at different angles and scales in the image.

[0009] In one embodiment, when converting each two-dimensional rendered image into grayscale images corresponding to the shortwave, midwave, and longwave bands respectively, based on the image physical characteristics of shortwave infrared, midwave infrared, and longwave infrared: Based on the characteristics of infrared wavelengths and radiation patterns, the correlation between gray values ​​and target attributes in each band is determined, thereby constructing conversion models for different bands. By using conversion models for different wavebands for each two-dimensional rendered image, and through dynamic range mapping and noise models, grayscale images of the corresponding shortwave, medium wave, and long wave bands are obtained.

[0010] In one embodiment, when a random background is superimposed on each of the two-dimensional mask images, for each band, the two-dimensional mask image corresponding to each band is superimposed on the random background image of the same band.

[0011] In one embodiment, during the training of the semantic segmentation network: Obtain a training sample dataset, which includes short-wave infrared, mid-wave infrared, and long-wave infrared images of the target object under different scenes, angles, and lighting conditions, as well as target annotation masks for each image. The semantic segmentation network is trained using the training sample dataset to minimize the difference between the predicted target mask and the corresponding labeled target mask, thereby optimizing the parameters in the semantic segmentation network.

[0012] In one embodiment, the loss function used when iteratively adjusting the adversarial texture parameters by feeding the loss back to the adversarial texture optimization process via backpropagation is expressed as:

[0013] In the above formula, For confidence loss function, Let the style transfer loss function be... Here, the weights of the style transfer loss function are denoted as , where the confidence loss function is denoted as . Represented as:

[0014] In the above formula, Optimize samples for multi-band infrared countermeasures. For target detection networks, The confidence score for the target.

[0015] This application also provides a multi-band infrared adversarial sample generation device, the device comprising: The target 3D model building module is used to build a target 3D model based on the real data of the target object; The 2D rendering image acquisition module is used to render the adversarial texture to be optimized onto the overall surface of the target 3D model using neural rendering technology, generate an adversarial texture 3D model, and process the adversarial texture 3D model using a differentiable neural renderer to obtain multiple 2D rendering images. The grayscale image construction module is used to convert each two-dimensional rendered image into grayscale images corresponding to the shortwave, midwave, and longwave bands, respectively, based on the image physical characteristics of shortwave infrared, midwave infrared, and longwave infrared. A multi-band infrared adversarial sample generation module is used to segment targets in each grayscale level image using a pre-trained semantic segmentation network, generate a two-dimensional mask image, and superimpose a random background on each two-dimensional mask image to generate multi-band infrared adversarial samples. The multi-band infrared adversarial sample optimization module is used to utilize a style transfer network to transfer the style of each band of infrared adversarial sample to a style that matches the infrared radiation characteristics of the real target in that band, thereby generating multi-band infrared adversarial optimized samples. The adversarial texture optimization module is used to input the multi-band infrared adversarial optimization sample into the target detection network for target detection, calculate the loss function based on the target detection result, feed the loss back to the adversarial texture optimization process through backpropagation, iteratively adjust the adversarial texture parameters until the loss function converges, obtain the optimized adversarial texture, and place the optimized adversarial texture on the target surface to achieve adversarial interference against multi-band infrared detection.

[0016] A computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the steps in the multi-band infrared adversarial sample generation method described above.

[0017] A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps in the above-described multi-band infrared adversarial sample generation method.

[0018] The aforementioned multi-band infrared adversarial example generation method, apparatus, device, and medium employ neural rendering technology to render the adversarial texture to be optimized onto the overall surface of a 3D model constructed based on real target data, generating an adversarial texture 3D model. A differentiable neural renderer processes the adversarial texture 3D model, and based on the image physical characteristics of short-wave infrared, mid-wave infrared, and long-wave infrared, each 2D rendered image is converted into grayscale images corresponding to the short-wave, mid-wave, and long-wave bands, respectively. Furthermore, a pre-trained semantic segmentation network is used to segment the target in each grayscale image, generating 2D mask images, which are then overlaid on each 2D mask image. A random background is added to generate multi-band infrared adversarial samples. Then, a style transfer network is used to transfer the style of each infrared adversarial sample to a style that matches the infrared radiation characteristics of the real target in that band, generating multi-band infrared adversarial optimized samples. The multi-band infrared adversarial optimized samples are input into a target detection network for target detection. The loss function is calculated based on the target detection results, and the loss is fed back to the optimization process of the adversarial texture through backpropagation. The adversarial texture parameters are iteratively adjusted until the loss function converges, resulting in the optimized adversarial texture. The optimized adversarial texture is then placed on the target surface to achieve adversarial interference against multi-band infrared detection. This method generates adversarial samples covering short-wave, mid-wave, and long-wave infrared bands, enabling effective interference across all bands in multi-band collaborative detection scenarios. Its neural rendering technology based on 3D models ensures the robustness of adversarial textures under different viewing angles and lighting conditions. The superposition of random backgrounds and style transfer optimization make it highly consistent with the real infrared radiation characteristics, making it difficult to be identified as an interference signal. At the same time, through iterative optimization based on detection network feedback, it can be adapted to various infrared detection systems, effectively breaking through the multi-band cross-validation mechanism, significantly reducing the probability of target identification, and comprehensively improving the target's security protection capability in complex infrared detection environments. Attached Figure Description

[0019] Figure 1 This is a flowchart illustrating a multi-band infrared adversarial sample generation method in one embodiment; Figure 2 This is a schematic diagram illustrating the specific implementation process of the method in one embodiment; Figure 3 This is a schematic diagram illustrating the effect of multi-band infrared adversarial samples in one embodiment; Figure 4 This is a structural block diagram of a multi-band infrared adversarial sample generation device in one embodiment; Figure 5 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation

[0020] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0021] Existing adversarial example generation methods mainly target detectors in a single band. For example, the IACS jamming method, although it generates adversarial examples at different angles and scales based on 3D rendering and can evade infrared detection models from multiple angles and scales, is mainly focused on a single band. While it can generate effective adversarial examples in a specific band, its limitations are becoming increasingly apparent in practical applications, especially in complex multi-band environments.

[0022] In one embodiment, such as Figure 1 As shown, a method for generating multi-band infrared adversarial examples is provided, including the following steps: Step S100: Construct a three-dimensional model of the target object based on its real data.

[0023] Step S110: Using neural rendering technology, the adversarial texture to be optimized is rendered onto the entire surface of the target 3D model to generate an adversarial texture 3D model. The adversarial texture 3D model is then processed using a differentiable neural renderer to obtain multiple 2D rendered images.

[0024] Step S120: Based on the image physical characteristics of short-wave infrared, mid-wave infrared and long-wave infrared, each two-dimensional rendered image is converted into grayscale images corresponding to the three bands of short-wave, mid-wave and long-wave respectively.

[0025] Step S130: Use a pre-trained semantic segmentation network to segment the target in each grayscale level image, generate a two-dimensional mask image, and superimpose a random background on each two-dimensional mask image to generate multi-band infrared adversarial samples. Step S140: Using a style transfer network, the style of each infrared adversarial sample in each band is transferred to a style that matches the infrared radiation characteristics of the real target in that band, generating multi-band infrared adversarial optimized samples.

[0026] Step S150: Input the multi-band infrared countermeasures optimization sample into the target detection network for target detection, calculate the loss function based on the target detection result, feed the loss back to the optimization process of the countermeasure texture through backpropagation, iteratively adjust the countermeasure texture parameters until the loss function converges, obtain the optimized countermeasure texture, and place the optimized countermeasure texture on the target surface to achieve countermeasures interference against multi-band infrared detection.

[0027] In step S100, a three-dimensional model of the target is constructed based on the real data of the target object.

[0028] In this embodiment, a target 3D model containing geometric features and basic infrared radiation properties is constructed based on the target object's real structural data (such as 3D scan point cloud) and infrared characteristic data (such as temperature distribution of real infrared images).

[0029] To construct a 3D model of a target containing geometric features and basic infrared radiation properties, the process involves several steps. First, a high-precision laser scanner is used to acquire 3D point cloud data of the target object. Then, multi-view photogrammetry is employed to capture 2D images of the target from different angles, supplementing the point cloud data with details. Simultaneously, an infrared camera is used to acquire infrared images of the target at different wavelengths, recording its temperature distribution and radiation characteristics. Next, the point cloud data undergoes noise reduction and filtering, the 2D images are distorted and color corrected, and the point cloud and 2D image data are fused to generate a textured 3D model. Then, the Delaunay triangulation algorithm is used to convert the point cloud into a high-quality 3D mesh model, and UV mapping technology is used to map the infrared images as texture onto the model surface, ensuring that the model's geometry and infrared radiation characteristics are consistent with the target object. Finally, the generated 3D model is optimized, including removing redundant faces, smoothing the mesh, and optimizing the topology to improve rendering efficiency and accuracy. Through this series of techniques, a high-precision, highly realistic 3D target model can be constructed, providing a solid foundation for subsequent adversarial example generation.

[0030] In step S110, the adversarial texture to be optimized can be generated by neural rendering technology according to preset parameters. An adversarial texture is a carefully designed pattern or texture whose core function is to mislead the infrared detection system, causing the target to exhibit non-realistic features in the infrared image that differ from its true characteristics. In the multi-band infrared adversarial scenario where this method is applied, the design of the adversarial texture needs to combine the physical characteristics of different bands (such as reflectivity and emissivity) with the weaknesses of the target detection algorithm.

[0031] Specifically, adversarial textures are represented in virtual texture form, or digital form. They primarily utilize the texture set of a 3D model through gradient backpropagation, adjusting the texture set using a loss function. The texture set is designed to address weaknesses in target detection algorithms, reducing target detectability through high-frequency noise or specific geometric patterns. Using a differentiable neural renderer, the adversarial texture 3D model is rendered under different viewpoints, lighting conditions, and scene parameters, generating diverse 2D rendered images that simulate changes in real-world scenes. Furthermore, the parameters of the adversarial texture (such as color, intensity, and contrast) are dynamically optimized during generation to adapt to different environmental conditions, ensuring effective interference with target detection in multi-band infrared imaging systems while maintaining a high degree of consistency with the real scene. This bridges the gap between the digital and physical worlds.

[0032] In this embodiment, when processing the adversarial texture 3D model using a differentiable neural renderer, geometric calculations, lighting calculations, and pixel calculations are performed sequentially under different viewpoints, lighting conditions, and scene parameters to generate 2D rendered images of the target at different angles and scales in the image. This results in diverse 2D rendered images that correspond to changes in the real scene, providing more realistic input for subsequent multi-terminal conversion and scene fusion.

[0033] In step S120, when converting each two-dimensional rendered image into grayscale images corresponding to the short-wave, mid-wave, and long-wave bands based on the image physical characteristics of short-wave infrared, mid-wave infrared, and long-wave infrared: based on the infrared wavelength characteristics and radiation laws, the correlation between the grayscale values ​​of each band and the target attributes is determined, thereby constructing conversion models for different bands. Using the conversion models for different bands, for each two-dimensional rendered image, through dynamic range mapping and noise models, grayscale images corresponding to the short-wave, mid-wave, and long-wave bands are obtained.

[0034] In this embodiment, in order to convert each two-dimensional rendered image into grayscale images corresponding to the three bands of shortwave, midwave, and longwave, a series of processing was performed based on the image physical characteristics of shortwave infrared, midwave infrared, and longwave infrared.

[0035] Specifically, firstly, different infrared bands possess different physical characteristics. Short-wave infrared primarily reflects the reflectivity of a target, with grayscale values ​​closely related to the target's reflectivity. Mid-wave infrared primarily reflects the thermal radiation characteristics of a target, with grayscale values ​​related to the target's temperature and emissivity; long-wave infrared primarily reflects the thermal radiation characteristics of a target, with grayscale values ​​related to the target's temperature and emissivity, but more strongly dependent on ambient temperature. Based on these characteristics, a conversion model for different bands was constructed to convert the pixel values ​​of a two-dimensional rendered image into corresponding grayscale levels.

[0036] For short-wave infrared, the following formula is used for pixel value conversion:

[0037] In the above formula, and It is a coefficient determined based on the target reflectivity and environmental conditions.

[0038] For mid-wave infrared, the following formula is used for pixel value conversion:

[0039] In the above formula, and It is a coefficient determined based on the target emissivity and temperature. It is a nonlinear adjustment parameter used to simulate the radiation characteristics of mid-wave infrared radiation.

[0040] For long-wave infrared, the following formula is used for pixel value conversion:

[0041] In the above formula, and It is a coefficient determined based on the target emissivity and ambient temperature. The logarithmic function is used to simulate the radiation characteristics of long-wave infrared radiation.

[0042] Subsequently, the grayscale was dynamically adjusted.

[0043] In the above formula, and These are the minimum and maximum gray values ​​for that band, respectively.

[0044] Through the above steps, grayscale images that are highly consistent with real infrared images can be generated based on the physical characteristics of shortwave, midwave, and longwave infrared, providing high-quality input for subsequent adversarial sample generation.

[0045] In step S130, to enhance the scene adaptability and realism of the adversarial examples, diverse natural or complex environmental backgrounds (such as infrared radiation characteristics under different terrain, climate, and lighting conditions) are introduced. This makes the generated multi-band infrared adversarial examples closer to the complex environment in which the target is located in actual detection scenarios, avoiding the adversarial textures being easily identified as abnormal interference signals by the detection system due to a single background. At the same time, random backgrounds can simulate various background interferences that targets may encounter in real applications, forcing the adversarial textures to not only be designed to interfere with the target's own characteristics during the optimization process, but also to adapt to the changing background radiation environment. This improves the robustness of the adversarial examples against interference from multi-band infrared detection systems in different scenarios during actual deployment, ensuring that they can still effectively confuse the boundary between the target and the background in complex real environments, reducing the probability of being accurately identified.

[0046] In this embodiment, after segmenting the image at each grayscale level using a semantic segmentation network, the boundary between the target region and the non-target region is clearly defined. Based on this, the background superimposed can be precisely controlled to act only on the region outside the target. This ensures that the adversarial texture of the target region is not disturbed and its interference characteristics are fully preserved, while allowing the background and target region to blend naturally, simulating the coexistence relationship between the target and the background in a real scene. This approach avoids the destruction of the target's adversarial texture by background information, ensuring the stability of the core interference capability of the adversarial sample. Furthermore, through mask constraints, the superimposed random background better matches the physical scene of "target embedding environment" in actual detection, making the generated adversarial sample more difficult to distinguish as anomaly in a multi-band infrared detection system, further improving the concealment and effectiveness of the adversarial interference.

[0047] Furthermore, when overlaying a random background onto each two-dimensional mask image, for each band, the two-dimensional mask image corresponding to each band is overlaid onto the random background image of the same band.

[0048] In this embodiment, when training the semantic segmentation network: a training sample dataset is acquired, which includes short-wave infrared, mid-wave infrared, and long-wave infrared images of the target object under different scenes, angles, and lighting conditions, as well as target annotation masks for each image. The semantic segmentation network is trained using the training sample dataset to minimize the difference between the predicted target mask and the corresponding target annotation mask, thereby optimizing the parameters in the semantic segmentation network. During training, the cross-entropy loss function is used to optimize the network parameters, and the Adam optimizer is employed for gradient descent.

[0049] Furthermore, to improve the robustness of the semantic segmentation network, data augmentation techniques, such as random cropping, rotation, flipping, and color jittering, were added to the training data to simulate different environmental conditions and perspective changes.

[0050] Specifically, the semantic segmentation network adopts the DeepLabv3+ architecture, which combines an encoder-decoder structure with dilated convolutions, effectively capturing contextual information and detailed features in images.

[0051] In step S140, to further enhance the realism and concealment of adversarial examples, a style transfer network is used to reduce the difference between multi-band infrared adversarial examples and actual infrared images. This makes the optimized multi-band infrared adversarial examples closer to real targets in terms of texture, grayscale distribution, and radiation intensity, avoiding being identified as abnormal signals by the detection system due to style deviations from real scenes. Simultaneously, it allows for deep integration of adversarial textures with the infrared physical characteristics of the target itself, ensuring that the interference effect is preserved in multi-band detection while conforming to the natural radiation patterns of the target in that band. This overcomes the detection system's identification mechanism for "abnormal radiation patterns," enhancing the applicability and interference effectiveness of adversarial examples in real-world environments and reducing the probability of being detected by the multi-band infrared detection system.

[0052] In this embodiment, when training the style transfer network, a large number of images containing multi-band infrared features of the target are collected. These images should cover short-wave infrared, mid-wave infrared and long-wave infrared bands. The style transfer network is trained and the network parameters are optimized to minimize the difference between the generated image and the target style image.

[0053] Specifically, the style transfer network adopts the StyleGAN architecture, which uses the images obtained above for analysis and processing to train the style transfer network and optimize the network parameters to minimize the difference between the generated image and the target style image.

[0054] In step S150, the generated multi-band infrared adversarial optimized samples are input into a multi-band infrared target detector. Based on feedback from the target detector, the loss function is updated using gradient backpropagation to further optimize the adversarial texture. This closed-loop optimization process ensures that the generated multi-band infrared adversarial samples possess stronger interference capabilities across multiple bands, including short-wave infrared, mid-wave infrared, and long-wave infrared, significantly improving their anti-interference capability and deception effectiveness in complex environments. Through this process, the generated multi-band infrared adversarial samples can effectively interfere with the target detection model in a multi-band infrared imaging system, improving the security and reliability of the infrared imaging system.

[0055] In this embodiment, the multi-band infrared target detector can adopt any model architecture capable of infrared image target detection, and is not limited to a specific algorithm type, including but not limited to deep learning-based YOLO series, Faster R-CNN, SSD and other networks, as well as traditional detection methods based on feature matching or threshold segmentation.

[0056] In this embodiment, the loss function used when iteratively adjusting the adversarial texture parameters by feeding back the loss to the adversarial texture optimization process through backpropagation is expressed as:

[0057] In the above formula, For confidence loss function, Let the style transfer loss function be... Here, the weights of the style transfer loss function are denoted as , where the confidence loss function is denoted as . Represented as:

[0058] In the above formula, Optimize samples for multi-band infrared countermeasures. For target detection networks, Let be the confidence score of the target. The input and output of the target detection network can be represented as:

[0059] In the above formula, The confidence score for the target. The category score of the target. For the bounding box information of the target, , To predict the center point coordinates of the target, To predict the width and height of the target boundary.

[0060] like Figure 2 The diagram shown illustrates the specific implementation process of this method, with a car as the target object.

[0061] This paper also demonstrates the effectiveness of the proposed method through experiments. In these experiments, vehicles were used as the target object. Furthermore, to more effectively evaluate the interference performance of adversarial examples, Average Precision (AP) and Attack Success Rate (ASR) were introduced as key evaluation metrics. Average Precision (AP) is a commonly used performance metric in object detection tasks, used to measure the overall detection capability of a model under different confidence thresholds. Attack Success Rate (ASR) measures the effectiveness of adversarial examples in practical applications, i.e., the proportion of times an adversarial example prevents the object detection model from correctly identifying the target under specific conditions. A higher ASR indicates a stronger attack performance of the adversarial example and a better interference effect on the object detection model.

[0062] This method employs 3D neural rendering technology and style transfer networks, achieving significant attack effectiveness in the short-wave infrared, mid-wave infrared, and long-wave infrared bands, as shown in Table 1. In the short-wave infrared band, the jamming success rate reaches 52.3%; in the mid-wave infrared band, the jamming success rate is as high as 97.1%; and in the long-wave infrared band, the jamming success rate reaches 81.8%. These multi-band infrared adversarial examples effectively interfere with the target detection model and reduce the target recognition capability of the infrared imaging system in these bands. Figure 3The diagram shown illustrates the effectiveness of multi-band infrared adversarial examples. In summary, this method overcomes the limitations of existing technologies in multi-band environments by generating high-quality and robust multi-band infrared adversarial examples. It provides a more comprehensive testing and improvement tool for the security and reliability of infrared imaging systems, and has significant practical application value and importance.

[0063] Table 1. Quantitative analysis of the attack effects on clean and adversarial samples using multi-band infrared technology.

[0064] In the aforementioned multi-band infrared adversarial sample generation method, a 3D model is constructed based on real target data and combined with neural rendering technology to achieve accurate mapping and multi-view rendering of adversarial textures in 3D space. This ensures that adversarial samples can adapt to different detection angles. By utilizing a differentiable neural renderer and multi-band physical feature transformation, the generated samples simultaneously cover short-wave, mid-wave, and long-wave infrared bands, overcoming the limitations of single-band adversarial methods. By using semantic segmentation masks and superimposing random backgrounds, the adaptability and realism of the samples in complex scenes are enhanced. By optimizing radiation characteristics through style transfer networks, the stealth of adversarial samples is improved, preventing them from being identified as abnormal signals by the detection system. Combined with feedback iterative optimization of the target detection network, the adversarial textures can accurately target the weaknesses of the detection model. Ultimately, this achieves efficient, robust, and stealthy interference against multi-band infrared detection systems, significantly improving the security protection capability of targets in multi-band collaborative detection environments. Moreover, the method has strong universality and can be adapted to various detection models and application scenarios.

[0065] It should be understood that, although Figure 1 The steps in the flowchart are shown sequentially as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order in which these steps are executed, and they can be performed in other orders. Figure 1 At least some of the steps in the process may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed in turn or alternately with other steps or at least some of the sub-steps or stages of other steps.

[0066] In one embodiment, such as Figure 4 As shown, a multi-band infrared adversarial example generation device is provided, including: a target 3D model construction module 200, a 2D rendering image acquisition module 210, a grayscale image construction module 220, a multi-band infrared adversarial example generation module 230, a multi-band infrared adversarial example optimization module 240, and an adversarial texture optimization module 250, wherein: The target 3D model construction module 200 is used to construct a target 3D model based on the real data of the target object; The 2D rendering image acquisition module 210 is used to render the adversarial texture to be optimized onto the overall surface of the target 3D model using neural rendering technology, generate an adversarial texture 3D model, and process the adversarial texture 3D model using a differentiable neural renderer to obtain multiple 2D rendering images. The grayscale image construction module 220 is used to convert each two-dimensional rendered image into grayscale images corresponding to the three bands of shortwave infrared, midwave infrared and longwave infrared, respectively, based on the image physical characteristics of shortwave infrared, midwave infrared and longwave infrared. The multi-band infrared adversarial sample generation module 230 is used to segment the grayscale images of each grayscale level using a pre-trained semantic segmentation network, generate a two-dimensional mask image, and superimpose a random background on each two-dimensional mask image to generate multi-band infrared adversarial samples. The multi-band infrared adversarial sample optimization module 240 is used to utilize a style transfer network to transfer the style of each band of infrared adversarial sample to a style that matches the infrared radiation characteristics of the real target in that band, thereby generating multi-band infrared adversarial optimized samples. The adversarial texture optimization module 250 is used to input the multi-band infrared adversarial optimization sample into the target detection network for target detection, calculate the loss function based on the target detection result, feed the loss back to the adversarial texture optimization process through backpropagation, iteratively adjust the adversarial texture parameters until the loss function converges, obtain the optimized adversarial texture, and arrange the optimized adversarial texture on the target surface to achieve adversarial interference against multi-band infrared detection.

[0067] Specific limitations regarding the multi-band infrared adversarial example generation device can be found in the limitations of the multi-band infrared adversarial example generation method described above, and will not be repeated here. Each module in the aforementioned multi-band infrared adversarial example generation device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the corresponding operations of each module.

[0068] In one embodiment, a computer device is provided, which may be a terminal, and its internal structure diagram may be as follows: Figure 5As shown, the computer device includes a processor, memory, network interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The network interface is used to communicate with external terminals via a network connection. When the computer program is executed by the processor, it implements a multi-band infrared adversarial sample generation method. The display screen can be an LCD screen or an e-ink screen. The input devices can be a touch layer covering the display screen, buttons, a trackball, or a touchpad mounted on the computer device casing, or an external keyboard, touchpad, or mouse.

[0069] Those skilled in the art will understand that Figure 5 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0070] In one embodiment, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to perform the following steps: Construct a 3D model of the target object based on its real data. Using neural rendering technology, the adversarial texture to be optimized is rendered onto the entire surface of the target 3D model to generate an adversarial texture 3D model. The adversarial texture 3D model is then processed using a differentiable neural renderer to obtain multiple 2D rendered images. Based on the physical characteristics of short-wave infrared, mid-wave infrared and long-wave infrared images, each two-dimensional rendered image is converted into grayscale images corresponding to the three bands of short-wave, mid-wave and long-wave respectively; A pre-trained semantic segmentation network is used to segment the grayscale images to generate two-dimensional mask images, and a random background is superimposed on each of the two-dimensional mask images to generate multi-band infrared adversarial samples. Using a style transfer network, for each band of infrared adversarial samples, the style is transferred to a style that matches the infrared radiation characteristics of the real target in that band, generating multi-band infrared adversarial optimized samples. The multi-band infrared countermeasures optimized sample is input into the target detection network for target detection. The loss function is calculated based on the target detection result. The loss is fed back to the optimization process of the countermeasure texture through backpropagation. The countermeasure texture parameters are iteratively adjusted until the loss function converges to obtain the optimized countermeasure texture. The optimized countermeasure texture is then placed on the target surface to achieve countermeasures against multi-band infrared detection.

[0071] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, the computer program performing the following steps when executed by a processor: Construct a 3D model of the target object based on its real data. Using neural rendering technology, the adversarial texture to be optimized is rendered onto the entire surface of the target 3D model to generate an adversarial texture 3D model. The adversarial texture 3D model is then processed using a differentiable neural renderer to obtain multiple 2D rendered images. Based on the physical characteristics of short-wave infrared, mid-wave infrared and long-wave infrared images, each two-dimensional rendered image is converted into grayscale images corresponding to the three bands of short-wave, mid-wave and long-wave respectively; A pre-trained semantic segmentation network is used to segment the grayscale images to generate two-dimensional mask images, and a random background is superimposed on each of the two-dimensional mask images to generate multi-band infrared adversarial samples. Using a style transfer network, for each band of infrared adversarial samples, the style is transferred to a style that matches the infrared radiation characteristics of the real target in that band, generating multi-band infrared adversarial optimized samples. The multi-band infrared countermeasures optimized sample is input into the target detection network for target detection. The loss function is calculated based on the target detection result. The loss is fed back to the optimization process of the countermeasure texture through backpropagation. The countermeasure texture parameters are iteratively adjusted until the loss function converges to obtain the optimized countermeasure texture. The optimized countermeasure texture is then placed on the target surface to achieve countermeasures against multi-band infrared detection.

[0072] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0073] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0074] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims.

Claims

1. A method for generating multi-band infrared adversarial examples, characterized in that, The method includes: Construct a 3D model of the target object based on its real data. Using neural rendering technology, the adversarial texture to be optimized is rendered onto the entire surface of the target 3D model to generate an adversarial texture 3D model. The adversarial texture 3D model is then processed using a differentiable neural renderer to obtain multiple 2D rendered images. Based on the physical characteristics of short-wave infrared, mid-wave infrared and long-wave infrared images, each two-dimensional rendered image is converted into grayscale images corresponding to the three bands of short-wave, mid-wave and long-wave respectively; A pre-trained semantic segmentation network is used to segment the grayscale images to generate two-dimensional mask images, and a random background is superimposed on each of the two-dimensional mask images to generate multi-band infrared adversarial samples. Using a style transfer network, for each band of infrared adversarial samples, the style is transferred to a style that matches the infrared radiation characteristics of the real target in that band, generating multi-band infrared adversarial optimized samples. The multi-band infrared countermeasures optimized sample is input into the target detection network for target detection. The loss function is calculated based on the target detection result. The loss is fed back to the optimization process of the countermeasure texture through backpropagation. The countermeasure texture parameters are iteratively adjusted until the loss function converges to obtain the optimized countermeasure texture. The optimized countermeasure texture is then placed on the target surface to achieve countermeasures against multi-band infrared detection.

2. The multi-band infrared adversarial sample generation method according to claim 1, characterized in that, When processing the adversarial texture 3D model using a differentiable neural renderer, geometric calculations, lighting calculations, and pixel calculations are performed sequentially to generate 2D rendered images of the target at different angles and scales in the image.

3. The multi-band infrared adversarial sample generation method according to claim 2, characterized in that, When converting each 2D rendered image into grayscale images corresponding to the shortwave, midwave, and longwave bands based on the image physical characteristics of shortwave infrared, midwave infrared, and longwave infrared: Based on the characteristics of infrared wavelengths and radiation patterns, the correlation between gray values ​​and target attributes in each band is determined, thereby constructing conversion models for different bands. By using conversion models for different wavebands for each two-dimensional rendered image, and through dynamic range mapping and noise models, grayscale images of the corresponding shortwave, medium wave, and long wave bands are obtained.

4. The multi-band infrared adversarial sample generation method according to claim 3, characterized in that, When superimposing a random background on each of the two-dimensional mask images, for each band, the two-dimensional mask image corresponding to each band is superimposed on the random background image of the same band.

5. The multi-band infrared adversarial sample generation method according to claim 4, characterized in that, When training the semantic segmentation network: Obtain a training sample dataset, which includes short-wave infrared, mid-wave infrared, and long-wave infrared images of the target object under different scenes, angles, and lighting conditions, as well as target annotation masks for each image. The semantic segmentation network is trained using the training sample dataset to minimize the difference between the predicted target mask and the corresponding labeled target mask, thereby optimizing the parameters in the semantic segmentation network.

6. The multi-band infrared adversarial sample generation method according to claim 5, characterized in that, The loss function used in the optimization process of the adversarial texture, which iteratively adjusts the adversarial texture parameters by feeding the loss back into the adversarial texture through backpropagation, is expressed as: In the above formula, For confidence loss function, For style transfer loss function, Here, the weights of the style transfer loss function are denoted as , where the confidence loss function is denoted as . Represented as: In the above formula, Optimize samples for multi-band infrared countermeasures. For target detection networks, The confidence score for the target.

7. A multi-band infrared adversarial sample generation device, characterized in that, The device includes: The target 3D model building module is used to build a target 3D model based on the real data of the target object; The 2D rendering image acquisition module is used to render the adversarial texture to be optimized onto the entire surface of the target 3D model using neural rendering technology, generate an adversarial texture 3D model, and process the adversarial texture 3D model using a differentiable neural renderer to obtain multiple 2D rendering images. The grayscale image construction module is used to convert each two-dimensional rendered image into grayscale images corresponding to the shortwave, midwave, and longwave bands, respectively, based on the image physical characteristics of shortwave infrared, midwave infrared, and longwave infrared. A multi-band infrared adversarial sample generation module is used to segment targets in each grayscale level image using a pre-trained semantic segmentation network, generate two-dimensional mask images, and overlay random backgrounds on each two-dimensional mask image to generate multi-band infrared adversarial samples. The multi-band infrared adversarial sample optimization module is used to utilize a style transfer network to transfer the style of each band of infrared adversarial sample to a style that matches the infrared radiation characteristics of the real target in that band, thereby generating multi-band infrared adversarial optimized samples. The adversarial texture optimization module is used to input the multi-band infrared adversarial optimization sample into the target detection network for target detection, calculate the loss function based on the target detection result, feed the loss back to the adversarial texture optimization process through backpropagation, iteratively adjust the adversarial texture parameters until the loss function converges, obtain the optimized adversarial texture, and place the optimized adversarial texture on the target surface to achieve adversarial interference against multi-band infrared detection.

8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Federal learning method and system based on multi-agent and knowledge distillation, and medium

    CN120354255A

  • Multi-band antenna

    US20050134509A1

  • Method and system for building short-wave, medium-wave and long-wave infrared spectrum dictionary

    US20240044715A1