Multi-modal adversarial patch attack system and method
By designing a multimodal adversarial patch attack system, using multimodal fusion and local distortion excitation technology, an adversarial sample that can cause significant prediction deviations in deep neural networks is generated, which solves the problem of poor multimodal attack effects in the existing technology and achieves a more efficient adversarial attack effect.
Patent Information
- Application Number
- CN202510302615.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-06-13
AI Technical Summary
The shortcomings of multimodal attacks and overall attacks in existing research, especially when fusing multimodal input information, are difficult to generate large prediction errors, and significant information between visible light and thermal imaging data cannot be fully considered.
A multimodal adversarial patch attack system is designed, including an image acquisition module, an adversarial sample initialization module, a multimodal fusion module, a local distortion excitation multimodal significant information extraction module and an attack optimization supervision module. Through the combination of these modules, an adversarial sample that can cause significant prediction deviations in deep neural networks can be generated.
By integrating visible and thermal image information, extracting features and iteratively identifying the most vulnerable multimodal information patches, the generated multimodal adversarial samples can effectively deceive the multimodal network model, significantly interfere with the normal function of the visual system, improve the adversarial effect, and provide a new perspective for further studying the fragility of deep learning models.
Smart Images

Figure CN120147784A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and in particular to a multimodal adversarial patch attack system and method. Background Art
[0002] Deep neural networks have shown great potential in achieving cross-modal learning and multimodal information fusion, such as processing visible light modal images and thermal imaging modalities simultaneously. However, most existing adversarial attack research focuses on single-modal deep learning models, and conducts in-depth analysis of the vulnerabilities of visible light images, for example. Although this single-modal attack research provides a basis for understanding the security of deep learning, it fails to fully consider the complex interactions and potential security threats that may occur in multimodal scenarios. In multimodal systems, various data sources often have interdependencies and complementary relationships, and traditional attack methods may not be able to be directly extended to this new input structure. Therefore, when using the fusion of visible light and thermal imaging data, attackers can manipulate the decision-making process of the model by designing adversarial samples for specific modalities. This makes the design of attacks more complex and hidden in multimodal learning environments. In addition, many attacks only consider attacks on overall information, without thinking and designing attack algorithms from a local perspective. Summary of the invention
[0003] The technical problem to be solved by the present invention is to provide a multimodal adversarial patch attack system and method. It aims to solve the deficiencies of multimodal attacks and overall attacks in existing research, especially how to produce large prediction errors when fusing multimodal input information. It can deeply mine the significant information between visible light and thermal imaging data and distort and aggregate the data in dense areas to design attack vectors, so that deep neural networks produce significant prediction deviations when facing such carefully constructed inputs.
[0004] To solve the above technical problems, the technical solution adopted by the present invention is: a multimodal adversarial patch attack system, comprising an image acquisition module, an adversarial sample initialization module, a multimodal fusion module, a local distortion-excited multimodal salient information extraction module, an attack optimization supervision module guided by a built-in loss coefficient, and an adversarial sample generation module for patch embedding operations with built-in pixel value replacement, which are connected in sequence.
[0005] The further improvement of the technical solution of the present invention is that each module is specifically composed as follows:
[0006] The image acquisition module includes a visible light image acquisition unit and a thermal imaging image acquisition unit which are independent of each other;
[0007] The adversarial sample initialization module includes a visible light adversarial sample initialization unit and an infrared thermal imaging adversarial sample initialization unit respectively connected to the visible light image acquisition unit and the infrared thermal imaging image acquisition unit, as well as 4 element multiplication operations and 2 addition operations;
[0008] The local distortion excitation multi-modal salient information extraction module includes a local distortion excitation unit and a multi-modal salient information extraction unit.
[0009] A further improvement of the technical solution of the present invention is that: the visible light adversarial sample initialization unit includes a visible light random mask generation branch, a visible light random noise generation branch, and a visible light inverse mask branch. The visible light inverse mask branch performs an element inversion operation on the mask generated by the visible light random mask generation branch to generate a visible light inverse mask. The first element multiplication operation is used to multiply each corresponding element of the visible light inverse mask and the visible light image obtained by the visible light acquisition unit. The second element multiplication operation is used to multiply each corresponding element of the visible light mask generated by the visible light random mask generation branch and the visible light random noise generated by the visible light random noise generation branch. The first addition operation is used to add the results obtained by the first element multiplication operation and the second element multiplication operation element by element to obtain the initial visible light adversarial sample;
[0010] The infrared thermal imaging adversarial sample initialization unit includes an infrared thermal imaging random mask generation branch, an infrared thermal imaging random noise generation branch, and an infrared thermal imaging inverse mask branch. The infrared thermal imaging inverse mask branch performs an element inversion operation on the mask generated by the infrared thermal imaging random mask generation branch to generate an infrared thermal imaging inverse mask. The third element multiplication operation is used to multiply each corresponding element of the infrared thermal imaging inverse mask and the infrared thermal imaging image obtained by the infrared thermal imaging acquisition unit. The fourth element multiplication operation is used to multiply each corresponding element of the infrared thermal imaging mask generated by the infrared thermal imaging random mask generation branch and the infrared thermal imaging random noise generated by the infrared thermal imaging random noise generation branch. The second addition operation is used to add the results obtained by the first element multiplication operation and the second element multiplication operation element by element to obtain the initial infrared thermal imaging adversarial sample.
[0011] A further improvement of the technical solution of the present invention is that: a spatial attention amplifier is arranged in the local distortion excitation unit, and the intensity of the spatial attention amplifier is controlled by a constant distortion factor k; the multi-modal salient information extraction unit realizes the calculation of class activation mapping based on the improved Grad-CAM by connecting the output layer of the target detection network and the feature map layer through a gradient backpropagation calculator.
[0012] A multi-modal adversarial patch attack method includes the following steps:
[0013] Step 1: Obtain visible light images and infrared thermal imaging images in complex scenes;
[0014] Step 2: Input the visible light image into the visible light adversarial sample initialization unit to obtain the initial visible light adversarial sample.
[0015] Input the thermal imaging image into the thermal imaging adversarial sample initialization unit to obtain the initial thermal imaging adversarial sample.
[0016] Step 3: Input the initial visible light adversarial sample and the initial thermal imaging adversarial sample into the multi-modal fusion module to obtain the multi-modal fusion feature map.
[0017] Step 4: Input the multi-modal fusion feature map into the local distortion excitation multi-modal significant information extraction module to obtain the dense abnormal information feature map and the attention thermal significant feature map.
[0018] Step 5: Input the dense abnormal information feature map, the attention thermal significant feature map, the initial visible light adversarial sample, and the initial thermal infrared adversarial sample into the attack optimization supervision module to obtain the final adversarial patch.
[0019] Step 6: Put the final patch into the adversarial sample generation module to generate the final visible light adversarial sample and the final thermal infrared adversarial sample.
[0020] A further improvement of the technical solution of the present invention lies in: The specific steps of Step 2 are as follows:
[0021] Step 2.1: Input the visible light image into the visible light random mask generation branch and the visible light random noise generation branch to obtain the visible light mask and the visible light patch respectively; input the visible light mask into the visible light reverse mask branch to obtain the visible light reverse mask.
[0022] Step 2.2: Add the result of element-wise multiplication of the visible light reverse mask and the visible light image and the result of element-wise multiplication of the visible light mask and the visible light patch to obtain the initial visible light adversarial sample.
[0023] Step 2.3: Input the thermal imaging image into the thermal imaging random mask generation branch and the thermal imaging random noise generation branch to obtain the thermal imaging mask and the thermal imaging patch respectively; input the thermal imaging mask into the thermal imaging reverse mask branch to obtain the thermal imaging reverse mask.
[0024] Step 2.4: Add the result of element-wise multiplication of the thermal imaging reverse mask and the thermal imaging image and the result of element-wise multiplication of the thermal imaging mask and the thermal imaging patch to obtain the initial thermal imaging adversarial sample.
[0025] A further improvement of the technical solution of the present invention lies in: The specific steps of Step 4 are as follows:
[0026] Step 4.1: Input the multi-modal fusion feature map into the local distortion excitation unit to obtain a dense anomaly information feature map;
[0027] Step 4.2: Input the multi-modal fusion feature map into the multi-modal significant information extraction unit to obtain an attention heat map significant feature map.
[0028] A further improvement of the technical solution of the present invention is that the specific steps of Step 5 are as follows:
[0029] Input the dense anomaly information feature map, the attention heat map significant feature map, the visible light initial adversarial sample, and the thermal infrared initial adversarial sample into the attack optimization supervision module to calculate the loss function for patch generation;
[0030] Use the loss function to determine the error of the optimal patch, and use the error backpropagation algorithm to backpropagate the error to guide patch learning until the loss function is minimized, so that the adversarial patch masters the important information of the image, updates the patch step by step, and finally outputs the final adversarial patch.
[0031] The calculation formula of the patch attack supervision mechanism is:
[0032]
[0033] L total = L LDI + λL MSIE
[0034] Where, x out represents the value predicted by the model, f Θ represents the multi-modal model, L LDI represents the local distortion excitation loss function, n and m represent the length and width of the image, x out (i,j) represents the value predicted by the neural network model at (i,j), and x target represents the ground truth value at (i,j), k represents the magnification factor of local distortion excitation, L MSIE represents the multi-modal significant extraction loss function, S i,j represents the pixel value of (i,j) in the attention heat map, represents the visible light adversarial sample, represents the thermal imaging adversarial sample, L total represents the total loss function, and λ represents the proportion of the multi-modal significant extraction loss function;
[0035] The calculation formula for guiding the update of the patch via the loss function is:
[0036]
[0037] Where, δ tdenote the adversarial patch at the t-th iteration, α denote the update step size, and L total denote the total loss function, and δ t+1 denote the (t + 1)-th iteration adversarial sample after updating the adversarial patch at the t-th iteration.
[0038] A further improvement of the technical solution of the present invention lies in: The specific steps of step 6 are as follows: Replace the pixel values of the local patch of the original visible light initial adversarial sample with the pixel values of the final adversarial patch to obtain the visible light final adversarial sample;
[0039] Replace the pixel values of the local patch of the original thermal infrared initial adversarial sample with the pixel values of the final adversarial patch to obtain the thermal infrared final adversarial sample.
[0040] Due to the adoption of the above technical solution, the technical progress achieved by the present invention is: Different from other adversarial attack methods for single-modal models, the present invention integrates visible light and thermal imaging image information to extract features and iteratively identify the most vulnerable multi-modal information patches. The generated multi-modal adversarial samples solve many limitations faced by single-modal network models and increase the attention to significant regions through comprehensive information analysis. It can create effective patches from training data and use these adversarial samples to attack multi-modal network models. By adding optimized patches to the input data, these patches mimic the patterns or features in the original training data to effectively deceive the model. This strategy can more effectively attack the weaknesses of the model. Adversarial patches can significantly interfere with the normal function of the visual system, thus posing challenges to the security and robustness of the system. This new patch attack method not only improves the adversarial effect but also provides a new perspective for further research on the vulnerability of deep learning models. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings;
[0042] Figure 1 is a schematic structural diagram of a multi-modal adversarial patch attack system of the present invention;
[0043] Figure 2 is a schematic flow diagram of a multi-modal adversarial patch attack method of the present invention;
[0044] Figure 3 is an explanatory diagram of the implementation structure content of a multi-modal adversarial patch attack method of the present invention;
[0045] Figure 4 This is the initial adversarial sample generation framework diagram in the present invention. Specific implementation manner
[0046] The present invention will be further described in detail below in conjunction with embodiments:
[0047] As Figure 1 shown, it is a structural schematic diagram of a multi-modal adversarial patch attack system, including an image acquisition module, an adversarial sample initialization module, a multi-modal fusion module, a local distortion excitation multi-modal significant information extraction module, an attack optimization supervision module guided by an internal loss coefficient, and an adversarial sample generation module with an internal pixel value replacement patch embedding operation, which are connected in sequence. Among them, the image acquisition module includes an independent visible light image acquisition unit and a thermal imaging image acquisition unit; the adversarial sample initialization module includes a visible light adversarial sample initialization unit and a thermal imaging adversarial sample initialization unit respectively connected to the visible light image acquisition unit and the thermal imaging image acquisition unit, as well as 4 element multiplication operations and 2 addition operations; the visible light adversarial sample initialization unit includes a visible light random mask generation branch, a visible light random noise generation branch, and a visible light inverted mask branch. The visible light inverted mask branch performs an element inversion operation on the mask generated by the visible light random mask generation branch to generate a visible light inverted mask. The first element multiplication operation is used to multiply each corresponding element of the visible light inverted mask and the visible light image obtained by the visible light acquisition unit. The second element multiplication operation is used to multiply each corresponding element of the visible light mask generated by the visible light random mask generation branch and the visible light random noise generated by the visible light random noise generation branch. The first addition operation is used to add the results obtained by the first element multiplication operation and the second element multiplication operation element by element to obtain a visible light initial adversarial sample; the thermal imaging adversarial sample initialization unit includes a thermal imaging random mask generation branch, a thermal imaging random noise generation branch, and a thermal imaging inverted mask branch. The thermal imaging inverted mask branch performs an element inversion operation on the mask generated by the thermal imaging random mask generation branch to generate a thermal imaging inverted mask. The third element multiplication operation is used to multiply each corresponding element of the thermal imaging inverted mask and the thermal imaging image obtained by the thermal imaging acquisition unit. The fourth element multiplication operation is used to multiply each corresponding element of the thermal imaging mask generated by the thermal imaging random mask generation branch and the thermal imaging random noise generated by the thermal imaging random noise generation branch. The second addition operation is used to add the results obtained by the first element multiplication operation and the second element multiplication operation element by element to obtain a thermal imaging initial adversarial sample.
[0048] The local distortion excitation multi-modal salient information extraction module includes a local distortion excitation unit and a multi-modal salient information extraction unit; a spatial attention amplifier is provided in the local distortion excitation unit, and the intensity of the spatial attention amplifier is controlled by a constant distortion factor k; the multi-modal salient information extraction unit realizes the calculation of class activation mapping based on the improved Grad-CAM by connecting the output layer of the target detection network with the feature map layer through a gradient backpropagation calculator.
[0049] The adversarial sample generation module contains a patch embedding operation that can change the pixel values of the local image.
[0050] As Figure 2 shown is a schematic flow chart of a multi-modal adversarial patch attack method implemented by means of the above system, and an explanatory diagram of the corresponding results generated by each module during the adversarial patch attack method shown as Figure 3 is used for explanation. The specific steps are as follows:
[0051] Step 1: Obtain visible light images and thermal imaging images in complex scenes;
[0052] Step 2: As Figure 4 shown, input the visible light image into the visible light adversarial sample initialization unit to obtain the initial visible light adversarial sample;
[0053] Step 2.1: Input the visible light image into the visible light random mask generation branch and the visible light random noise generation branch to obtain the visible light mask and the visible light patch respectively; input the visible light mask into the visible light inverted mask branch to obtain the visible light inverted mask;
[0054] Step 2.2: Add the result of element-wise multiplication of the visible light inverted mask and the visible light image and the result of element-wise multiplication of the visible light mask and the visible light patch to obtain the initial visible light adversarial sample;
[0055] Step 2.3: Input the thermal imaging image into the thermal imaging random mask generation branch and the thermal imaging random noise generation branch to obtain the thermal imaging mask and the thermal imaging patch respectively; input the thermal imaging mask into the thermal imaging inverted mask branch to obtain the thermal imaging inverted mask;
[0056] Step 2.4: Add the result of element-wise multiplication of the thermal imaging inverted mask and the thermal imaging image and the result of element-wise multiplication of the thermal imaging mask and the thermal imaging patch to obtain the initial thermal imaging adversarial sample.
[0057] Input the thermal imaging image into the thermal imaging adversarial sample initialization unit to obtain the initial thermal imaging adversarial sample;
[0058] Step 3: Input the visible light initial adversarial sample and the thermal imaging initial adversarial sample into the multi-modal fusion module to obtain a multi-modal fusion feature map;
[0059] Step 4: Input the multi-modal fusion feature map into the local distortion excitation multi-modal significant information extraction module to obtain a dense anomaly information feature map and an attention thermal significant feature map;
[0060] Step 4.1: Input the multi-modal fusion feature map into the local distortion excitation unit to obtain a dense anomaly information feature map;
[0061] Step 4.2: Input the multi-modal fusion feature map into the multi-modal significant information extraction unit to obtain an attention thermal significant feature map.
[0062] Step 5: Input the dense anomaly information feature map, the attention thermal significant feature map, the visible light initial adversarial sample, and the initial adversarial sample of thermal infrared into the attack optimization supervision module to obtain the final adversarial patch;
[0063] Send the dense anomaly information feature map, the attention thermal significant feature map, the visible light initial adversarial sample, and the initial adversarial sample of thermal infrared into the attack optimization supervision module, and calculate the loss function for patch generation;
[0064] Use the loss function to determine the error of the best patch, and use the error backpropagation algorithm to backpropagate the error to guide patch learning until the loss function is minimized, so that the adversarial patch masters the important information of the image, updates the patch step by step, and finally outputs the final adversarial patch.
[0065] The calculation formula of the patch attack supervision mechanism is:
[0066]
[0067]
[0068] L total = L LDI + λL MSIE
[0069] where x out represents the value predicted by the model, f Θ represents the multi-modal model, L LDI represents the local distortion excitation loss function, n and m represent the length and width of the image, x out (i, j) represents the value predicted by the neural network model at (i, j), and x target represents the ground truth value at (i, j), k represents the magnification factor of local distortion excitation, L MSIE represents the multi-modal significant extraction loss function, and S i,jRepresents the pixel value at (i, j) in the attention heat map, Represents the visible light adversarial sample, Represents the thermal imaging adversarial sample, L total Denotes the total loss function, and λ represents the proportion of the multi-modal saliency extraction loss function;
[0070] The calculation formula for updating the patch under the guidance of the loss function is:
[0071]
[0072] Among them, δ t Represents the adversarial patch at the t-th iteration, α represents the update step size, L total Represents the total loss function, and δ t+1 Represents the adversarial sample at the (t + 1)-th iteration after updating the adversarial patch at the t-th iteration.
[0073] Step 6: Put the final patch into the adversarial sample generation module to generate the final visible light adversarial sample and the final thermal infrared adversarial sample. Replace the pixel values of the local patch of the original visible light initial adversarial sample with the pixel values of the final adversarial patch to obtain the final visible light adversarial sample;
[0074] Replace the pixel values of the local patch of the original thermal infrared initial adversarial sample with the pixel values of the final adversarial patch to obtain the final thermal infrared adversarial sample.
[0075] The local distortion excitation unit designed by the local distortion excitation multi-modal saliency information extraction module of the present invention makes the designed patch more likely to obtain salient features by expanding the proportion of key attention and reducing the attention to irrelevant regions during the process of attacking the deep neural model. And the designed multi-modal saliency information extraction unit can mislead the model in reverse by adaptively learning the salient regions of different modalities during the patch generation process, thereby producing an attack effect, and can concentrate the attention of the multi-modal model on the salient regions. Specifically, if the features of the salient region are learned and the designed attack patch can also contain the features of the salient region, then a large error can be generated between the prediction result and the true value, thus achieving the purpose of the attack. And multi-modal involves inputs from multiple modalities, which is more challenging than single-modal tasks.
[0076] The embodiments described above are only used to describe the preferred embodiments of the present invention, and do not limit the scope of the present invention. Without departing from the design spirit of the present invention, various deformations and improvements made by those of ordinary skill in the art to the technical solutions of the present invention shall fall within the protection scope determined by the claims of the present invention.
Claims
1. A multimodal adversarial patch attack system, characterized by: It includes an image acquisition module, an adversarial sample initialization module, a multimodal fusion module, a local distortion-excited multimodal salient information extraction module, an attack optimization supervision module guided by a built-in loss coefficient, and an adversarial sample generation module for patch embedding operations with built-in pixel value replacement, which are connected in sequence.
2. A multimodal adversarial patch attack system according to claim 1, characterized in that: The specific components of each module are as follows: The image acquisition module includes a visible light image acquisition unit and a thermal imaging image acquisition unit which are independent of each other; The adversarial sample initialization module includes a visible light adversarial sample initialization unit and a thermal imaging adversarial sample initialization unit respectively connected to the visible light image acquisition unit and the thermal imaging image acquisition unit, as well as 4 element multiplication operations and 2 addition operations; The local distortion excitation multimodal salient information extraction module includes a local distortion excitation unit and a multimodal salient information extraction unit; The adversarial example generation module contains a patch embedding operation that can change the local image pixel values.
3. A multimodal adversarial patch attack system according to claim 2, characterized in that: The visible light adversarial sample initialization unit includes a visible light random mask generation branch, a visible light random noise generation branch, and a visible light inversion mask branch. The visible light inversion mask branch performs an element inversion operation on the mask generated by the visible light random mask generation branch to generate a visible light inversion mask. The first element multiplication operation is used to multiply each corresponding element of the visible light inversion mask and the visible light image obtained by the visible light acquisition unit. The second element multiplication operation is used to multiply each corresponding element of the visible light mask generated by the visible light random mask generation branch and the visible light random noise generated by the visible light random noise generation branch. The first addition operation is used to perform element-wise addition of the result obtained by the first element multiplication operation and the result obtained by the second element multiplication operation to obtain a visible light initial adversarial sample. The thermal imaging adversarial sample initialization unit includes a thermal imaging random mask generation branch, a thermal imaging random noise generation branch, and a thermal imaging inversion mask branch. The thermal imaging inversion mask branch performs an element inversion operation on the mask generated by the thermal imaging random mask generation branch to generate a thermal imaging inversion mask. The third element multiplication operation is used to multiply each corresponding element of the thermal imaging inversion mask and the thermal imaging image obtained by the thermal imaging acquisition unit. The fourth element multiplication operation is used to multiply each corresponding element of the thermal imaging mask generated by the thermal imaging random mask generation branch and the thermal imaging random noise generated by the thermal imaging random noise generation branch. The second addition operation is used to add the result obtained by the first element multiplication operation and the result obtained by the second element multiplication operation to obtain the thermal imaging initial adversarial sample.
4. A multimodal adversarial patch attack system according to claim 2, characterized in that: A spatial attention amplifier is set in the local distortion excitation unit, and the strength of the spatial attention amplifier is controlled by a constant distortion factor k; the multimodal salient information extraction unit is composed of a gradient back-propagation calculator that connects the output layer of the target detection network with the feature layer layer to realize the class activation mapping calculation based on the improved Grad-CAM.
5. A multimodal adversarial patch attack method, implemented by the attack system according to any one of claims 1 to 4, characterized in that: The steps include: Step 1: Acquire visible light images and thermal imaging images in complex scenes; Step 2: Input the visible light image into the visible light adversarial sample initialization unit to obtain the visible light initial adversarial sample; Input the thermal imaging image into the thermal imaging adversarial sample initialization unit to obtain the thermal imaging initial adversarial sample; Step 3: Input the visible light initial adversarial sample and the thermal imaging initial adversarial sample into the multimodal fusion module to obtain a multimodal fusion feature map; Step 4: Input the multimodal fusion feature map into the local distortion-excited multimodal salient information extraction module to obtain a dense abnormal information feature map and an attention thermal salient feature map; Step 5: Input the dense anomaly information feature map, the attention thermal salient feature map, the visible light initial adversarial sample, and the thermal infrared initial adversarial sample into the attack optimization supervision module to obtain the final adversarial patch; Step 6: Put the final patch into the adversarial sample generation module to generate the visible light final adversarial sample and the thermal infrared final adversarial sample.
6. A multimodal adversarial patch attack method according to claim 5, characterized in that: Step 2 The specific steps are as follows: Step 2.1: Input the visible light image into the visible light random mask generation branch and the visible light random noise generation branch to obtain a visible light mask and a visible light patch respectively; input the visible light mask into the visible light inversion mask branch to obtain a visible light inversion mask; Step 2.2: Add the result of element-wise multiplication of the visible light inversion mask and the visible light image and the result of element-wise multiplication of the visible light mask and the visible light patch to obtain the visible light initial adversarial sample; Step 2.3: Input the thermal imaging image into the thermal imaging random mask generation branch and the thermal imaging random noise generation branch to obtain the thermal imaging mask and thermal imaging patch respectively; Inputting the thermal imaging mask into the thermal imaging inversion mask branch to obtain a thermal imaging inversion mask; Step 2.4: Add the result of element-wise multiplication of the thermal imaging inversion mask and the thermal imaging image and the result of element-wise multiplication of the thermal imaging mask and the thermal imaging patch to obtain the thermal imaging initial adversarial sample.
7. A multimodal adversarial patch attack method according to claim 5, characterized in that: Step 4 The specific steps are as follows: Step 4.1: Input the multimodal fusion feature map into the local distortion excitation unit to obtain a dense abnormal information feature map; Step 4.2: Input the multimodal fusion feature map into the multimodal salient information extraction unit to obtain the attention thermal salient feature map.
8. The multimodal adversarial patch attack method according to claim 5, characterized in that: Step 5 The specific steps are as follows: The dense anomaly information feature map, the attention thermal salient feature map, the visible light initial adversarial sample, and the thermal infrared initial adversarial sample are fed into the attack optimization supervision module to calculate the loss function for patch generation. The error of the best patch is determined by using the loss function, and the error back propagation algorithm is used to back propagate the error to guide patch learning until the loss function is minimized, so that the adversarial patch can grasp important information of the image, and the patch is updated step by step, and finally the final adversarial patch is output. The calculation formula of the patch attack supervision mechanism is: THE total =L LDI +λL MSIE Among them, x out represents the value predicted by the model, f Θ represents the multimodal model, L LDI represents the local distortion excitation loss function, n and m represent the length and width of the image, x out (i, j) represents the value estimated by the neural network model at (i, j), and x target represents the ground truth value at (i, j), k represents the amplified local distortion excitation multiple, L MSIE represents the multimodal significant extraction loss function, S i,j Represents the pixel value of (i,j) in the attention heat map, represents the visible light adversarial sample, represents the thermal imaging adversarial sample, L total represents the total loss function, and λ represents the proportion of multimodal significant extraction loss function; The calculation formula for updating the patch guided by the loss function is: Among them, δ t represents the t-th iteration adversarial patch, α represents the update step size, and L total represents the total loss function, δ t+1 It represents the t+1th iteration adversarial sample after updating the tth iteration adversarial patch.
9. A multimodal adversarial patch attack method according to claim 5, characterized in that: The specific steps of step 6 are as follows: replace the pixel values of the local patch of the original visible light initial adversarial sample with the pixel values of the final adversarial patch to obtain the visible light final adversarial sample; The pixel value of the final adversarial patch is used to replace the pixel value of the local patch of the original thermal infrared initial adversarial sample to obtain the thermal infrared final adversarial sample.