A method and device for generating adversarial samples based on multi-feature saliency map fusion
Through a multi-feature significant graph fusion method, combined with gradient and perturbability interpretability algorithms, high-quality adversarial samples are generated, which solves the problems of contingency and bias in the generation of adversarial samples in the prior art, and improves the concealment of adversarial samples and the robustness of artificial intelligence models.
Patent Information
- Application Number
- CN202510176897.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-18
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-02-18
AI Technical Summary
There are contingencies and biases in the existing adversarial sample generation methods based on salient graphs, and it is difficult to ensure the quality of generated adversarial samples.
The method based on multi-feature significant graph fusion is adopted to fuse the interpretability algorithm results of gradients and perturbations to generate a fused significant feature map, reducing the chances and biases brought by a single algorithm, and improving the concealment of the adversarial samples by adding a small number of local adversarial perturbations multiple times.
It effectively reduces the chance and bias in the generation of adversarial samples, improves the concealment and quality of adversarial samples, and helps improve the robustness of artificial intelligence models.
Smart Images

Figure CN119672472B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence security, and specifically to a method and device for generating adversarial samples based on multi-feature saliency map fusion. Background Art
[0002] In recent years, artificial intelligence technologies represented by generative artificial intelligence and deep learning have been gradually widely used in the fields of situational awareness and target detection. The emergence of adversarial samples has posed a great threat to the security and stability of artificial intelligence algorithm models. Adversarial samples are highly concealed, highly targeted, and ubiquitous. Once spread, they may greatly weaken the effectiveness of one's own artificial intelligence model. Fortunately, retraining artificial intelligence models with adversarial samples can improve the robustness of the model to a certain extent and effectively suppress the negative impact of adversarial samples. Therefore, research on generating high-quality adversarial samples is crucial to conducting adversarial training to improve the robustness and security of the model.
[0003] However, existing methods mainly obtain sample salient areas (areas with the greatest contribution to model recognition) based on heat maps, and then add adversarial perturbations to the areas to obtain adversarial samples. Among them, the quality of the generated heat map determines the quality of the adversarial sample. Most existing methods are based on a single heat map generation algorithm, which is random and biased, and it is difficult to ensure that the selected salient feature points are consistent with the actual situation, that is, it is difficult to ensure the quality of the generated adversarial samples.
[0004] In view of this, in view of the defects and shortcomings of the existing adversarial sample generation methods based on saliency maps, how to generate higher-quality adversarial samples to help improve the robustness of artificial intelligence models is a problem that needs to be solved. Summary of the invention
[0005] The embodiments of the present application provide a method and device for generating adversarial samples based on multi-feature saliency map fusion, so as to solve the defects and shortcomings of the adversarial sample generation method based on saliency map in the prior art and generate higher quality adversarial samples.
[0006] In order to achieve the above objectives, this application adopts the following technical solutions:
[0007] In a first aspect, the present application provides an adversarial sample generation method based on the fusion of multiple feature saliency maps, the method comprising: obtaining multiple image sample data, and dividing the multiple image sample data into a training set, a test set and a validation set, training an artificial intelligence model through the training set, adjusting the parameters of the artificial intelligence model through the validation set, and selecting the optimal model parameters through the test set; inputting any original image sample into the artificial intelligence model, obtaining a first feature saliency map based on a gradient interpretability method, obtaining a second feature saliency map based on a perturbation interpretability method, fusing and enhancing the first feature saliency map and the second feature saliency map to obtain an enhanced feature saliency map; truncating and normalizing the enhanced feature saliency map to obtain a processed enhanced feature saliency map; based on a global gradient attack algorithm, obtaining a global adversarial perturbation of the original image sample; superimposing the global adversarial perturbation and the processed enhanced feature saliency map to obtain a local adversarial perturbation of the original image sample; adding the local adversarial perturbation to the original image sample to generate an adversarial sample.
[0008] A possible design scheme, the method of the first aspect also includes dividing the multiple image sample data into a training set, a test set and a validation set, including: stratified sampling of the multiple image sample data, and dividing them into a training set, a test set and a validation set according to a preset ratio.
[0009] A possible design scheme, the method of the first aspect also includes obtaining a first feature saliency map based on a gradient interpretability method, including: forward propagating the original image sample in the artificial intelligence model to obtain the output feature map of the last convolutional layer; backpropagating based on the prediction score of a specific category to obtain the gradient of the prediction score in the output feature map; performing global average pooling on the gradient to obtain the weight of the output feature map in each channel; based on the weight of each channel, weighted summing the output feature map to obtain a class activation map; filtering the class activation map based on an activation function to remove negative values; interpolating the processed class activation map so that the processed class activation map is consistent with the original image sample size to obtain the first feature saliency map.
[0010] A possible design scheme, the method of the first aspect also includes obtaining a second feature saliency map based on a perturbation interpretability method, including: randomly generating multiple occlusion templates, wherein each occlusion template covers a portion of the original image sample and the rest remains unchanged; based on the occlusion templates, obtaining multiple occluded original image samples; using an artificial intelligence model, predicting each occluded original image sample and recording the prediction results; comparing the prediction results of the original image sample and the occluded original image sample to obtain an importance score for each feature point or area in the original image sample to the artificial intelligence model; and obtaining a second feature saliency map based on the importance score of each feature point or area to the artificial intelligence model.
[0011] A possible design scheme, the method of the first aspect also includes fusing and enhancing the first feature saliency map and the second feature saliency map to obtain an enhanced feature saliency map, including: fusing and enhancing the first feature saliency map and the second feature saliency map through a weighted average fusion formula to obtain an enhanced feature saliency map, wherein the weighted average fusion formula is S=λ·S1+(1-λ)·S2, S1 is the first feature saliency map, S2 is the second feature saliency map, S is the enhanced feature saliency map, λ is the weight coefficient, and the value range is [0,1].
[0012] In a possible design scheme, the method of the first aspect further includes truncating and normalizing the enhanced feature saliency map, including: performing threshold truncation processing on the enhanced feature saliency map, including: if the feature saliency value in the enhanced feature saliency map is less than a preset threshold, setting the feature saliency value to zero; if the feature saliency value in the enhanced feature saliency map is greater than the preset threshold, retaining the feature saliency value;
[0013] The feature saliency map after threshold truncation is normalized, including: obtaining the maximum feature saliency value and the minimum feature saliency value in the enhanced feature saliency map, and normalizing each feature saliency value by a normalization formula, wherein the normalization formula is p=(pp min ) / (p max -p min ), p max is the maximum eigensignificant value, p min is the minimum feature significance value.
[0014] A possible design scheme, the method of the first aspect also includes, based on the global gradient attack algorithm, obtaining a global adversarial perturbation of the original image sample, including: calculating the gradient of the artificial intelligence model loss function with respect to the original image sample, wherein the loss function is J(θ, x, y), θ is a parameter of the artificial intelligence model, x is the original image sample, y is the true label of the original image sample, and the gradient of the original image sample is represents the derivative of the loss function; the gradient of the original image sample is signed by the sign function and multiplied by a preset constant to obtain the global adversarial perturbation of the original image sample, where the sign function is sign(·).
[0015] A possible design scheme, the method of the first aspect also includes adding local adversarial perturbations to the original image samples to generate adversarial samples, including: adding local adversarial perturbations to the original image samples to generate adversarial samples; if the decision result of the artificial intelligence model for the adversarial sample is inconsistent with the true label of the original image sample, it is determined that the adversarial sample is successfully generated; if the decision result of the artificial intelligence model for the adversarial sample is consistent with the true label of the original image sample, then adding local adversarial perturbations again until the decision result of the artificial intelligence model for the adversarial sample is inconsistent with the true label of the original image sample, and determining that the adversarial sample is successfully generated.
[0016] In a second aspect, an adversarial sample generation device based on multi-feature saliency map fusion is provided, and the adversarial sample generation device based on multi-feature saliency map fusion includes a module for executing the method of the first aspect mentioned above.
[0017] In a possible design scheme, the adversarial sample generation device based on multi-feature saliency map fusion of the second aspect may further include a transceiver. The transceiver may be a transceiver circuit or an interface circuit. The transceiver may be used for the adversarial sample generation device based on multi-feature saliency map fusion of the second aspect to communicate with other devices.
[0018] In a possible design scheme, the adversarial sample generation device based on multi-feature saliency map fusion of the second aspect may also include a memory. The memory may be integrated with the processor or may be separately provided. The memory may be used to store instructions involved in the method of the first aspect.
[0019] In an embodiment of the present application, by fusing the results of the gradient-based and perturbation-based interpretability algorithms to generate a fused significant feature map, feature points with strong significance can be determined, effectively reducing the randomness and bias brought about by a single interpretable algorithm; thereafter, by adding local adversarial perturbations in small amounts and multiple times, combined with judging whether the decision results of the artificial intelligence model for the adversarial sample are consistent with the true labels of the original image samples, redundant adversarial perturbations are reduced, and the concealment of the adversarial samples is effectively improved.
[0020] Other features and advantages of the present application will be described in detail in the subsequent detailed description. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.
[0022] Figure 1A schematic diagram of a process for generating adversarial samples based on multi-feature saliency map fusion provided in an embodiment of the present application;
[0023] Figure 2 A schematic diagram of a process for obtaining an enhanced feature saliency map provided in an embodiment of the present application;
[0024] Figure 3 A flowchart of adversarial sample generation provided in an embodiment of the present application;
[0025] Figure 4 Schematic diagram of the structure of the adversarial sample generation device based on multi-feature saliency map fusion provided in the embodiment of the present application Figure 1 ;
[0026] Figure 5 Schematic diagram of the structure of the adversarial sample generation device based on multi-feature saliency map fusion provided in the embodiment of the present application Figure 2 . DETAILED DESCRIPTION
[0027] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without making creative work are within the scope of protection of the present application. At the same time, in the description of the embodiments of the present application, the terms "first", "second", etc. are only used to distinguish the descriptions and cannot be understood as indicating or implying relative importance. Thus, the features defined as "first" and "second" may explicitly or implicitly include one or more features. In the description of the embodiments of the present application, the meaning of "multiple" is two or more, unless otherwise clearly and specifically defined.
[0028] Figure 1 A flowchart of an adversarial sample generation method based on multi-feature saliency map fusion provided in an embodiment of the present application.
[0029] The process of the adversarial sample generation method based on multi-feature saliency map fusion is as follows:
[0030] Step S101, obtain multiple image sample data, and divide the multiple image sample data into a training set, a test set and a validation set, train the artificial intelligence model through the training set, adjust the parameters of the artificial intelligence model through the validation set, and select the optimal model parameters through the test set.
[0031] That is to say, the obtained multiple image sample data are stratified and sampled, and divided into training set, test set and validation set according to a preset ratio.
[0032] For example, assuming that 10,000 image samples are obtained, stratified sampling is performed according to category, and the images are divided into a training set of 8,000, a test set of 1,000, and a test set of 1,000. Then, the artificial intelligence model is trained using the training set of 8,000 images, and the model parameters are adjusted using the validation set during the training process. Finally, the model is tested using the test set of 1,000 images. When the performance in the test set no longer improves, the model parameters are saved and the training process ends.
[0033] Step S102: input any original image sample into the artificial intelligence model, obtain a first feature saliency map based on a gradient interpretability method, obtain a second feature saliency map based on a perturbation interpretability method, and fuse and enhance the first feature saliency map and the second feature saliency map to obtain an enhanced feature saliency map.
[0034] In this application, the gradient interpretability method and the perturbation interpretability method are used to obtain an enhanced feature saliency map (or a fused feature saliency map). The specific process is as follows: Figure 2 shown.
[0035] Step S201, obtain the first feature saliency map S based on the gradient interpretability method 1 .
[0036] Step S202, obtaining a second feature saliency map S based on the perturbation interpretability method 2 .
[0037] Step S203: convert the first feature saliency map S 1 and the second feature saliency map S 1 Perform fusion enhancement to obtain the enhanced feature saliency map S.
[0038] Among them, the specific process of step S201 includes: forward propagating the original image sample x in the artificial intelligence model F to obtain the output feature map of the last convolutional layer; backpropagating based on the prediction score of a specific category to obtain the gradient of the prediction score in the output feature map; globally average pooling the gradient to obtain the weight of the output feature map in each channel; based on the weight of each channel, weighted summing the output feature map to obtain a class activation map; filtering the class activation map based on the activation function to remove negative values; interpolating the processed class activation map to make the processed class activation map consistent with the original image sample size to obtain a first feature saliency map.
[0039] That is to say, after the original image sample x is forward propagated, the output feature map of the last convolutional layer in channel k is A k ;
[0040] The AI model F predicts a score y for category c c, y c Perform back propagation and get y c In A k The gradient on ), Represents the output feature map A k The data at coordinate (i, j) in channel k;
[0041] Then, after global average pooling, the weight of the output feature map on channel k is obtained, that is, Here, Z is equal to the width × height of the output feature map.
[0042] Then perform the above processing on each channel, and then perform weighted summation on the processed results to obtain the class activation map, that is,
[0043] Then the negative values are removed through the ReLU activation function, that is,
[0044] Finally, the processed class activation map is interpolated to make it consistent with the original image sample size, and superimposed with the original image sample x to obtain the first feature saliency map S 1 .
[0045] The specific process of step S202 includes: randomly generating multiple occlusion templates, wherein each occlusion template covers a portion of the original image sample, and the rest remains unchanged; based on the occlusion templates, obtaining multiple occluded original image samples; using the artificial intelligence model, predicting each occluded original image sample, and recording the prediction results; comparing the prediction results of the original image sample and the occluded original image sample to obtain the importance score of each feature point or area in the original image sample to the artificial intelligence model; based on the importance score of each feature point or area to the artificial intelligence model, obtaining a second feature saliency map.
[0046] For example, assuming that the size of the input original image sample x is 3×224×224, 1000 occlusion templates are generated to randomly occlude x, and 1000 different occlusion samples x1, ..., x1000 are obtained;
[0047] Input x1, ..., x1000 into the model in sequence and observe the model's score in category c
[0048] If y c The value of decreases, which proves that the positive importance of the occluded part to the model decision is greater, otherwise it is smaller; through y c The rate of change of the value is used to calculate the importance of the obscured part;
[0049] For pixel p, its importance score is:
[0050] The mask p is an indicator function that takes the value 1 if the pixel p is covered, and 0 otherwise. represents the expectation for all samples.
[0051] In order to facilitate the comparison of the importance of all pixels, the importance scores of all pixels are normalized:
[0052] Generate the second feature saliency map S according to the importance score of each pixel 2 , in the second feature saliency map, the highlighted area can be used to represent the part that plays an important role in the model decision.
[0053] The specific process of step S203 includes: 1 and the second feature saliency map S 2 , through the weighted average fusion formula, fusion enhancement is performed to obtain the enhanced feature saliency map S, where the weighted average fusion formula is S = λ·S1+(1-λ)·S2, S1 is the first feature saliency map, S2 is the second feature saliency map, S is the enhanced feature saliency map, λ is the weight coefficient, and the value range is [0,1].
[0054] Furthermore, the value of λ can be adjusted as needed to reflect the importance of different inputs.
[0055] For example, assuming that the partial area of the saliency map S1 of the image x is [(1,0,1), (1,1,0), (0,1,1)], the corresponding area of the saliency map S2 and S1 is [(1,0,0), (1,1,1), (0,1,1)], and λ = 0.4, then the corresponding area of the enhanced feature saliency map S is:
[0056] S=0.4×[(1,0,1),(1,1,0),(0,1,1)]+0.6×[(1,0,0),(1,1,1),(0,1,1)]=[(1,0,0.4),(1,1,0.6),(0,1,1)].
[0057] Step S103, truncate and normalize the enhanced feature saliency map to obtain a processed enhanced feature saliency map.
[0058] The enhanced feature saliency map is subjected to a threshold truncation process, including: if the feature saliency value in the enhanced feature saliency map is less than a preset threshold value ω, the feature saliency value is set to zero; if the feature saliency value in the enhanced feature saliency map is greater than the preset threshold value ω, the feature saliency value is retained;
[0059] The feature saliency map after threshold truncation is normalized, including: obtaining the maximum feature saliency value p in the enhanced feature saliency map max and the minimum eigensignificant value p min , and normalize each feature significance value through the normalization formula, where the normalization formula is p = (pp min ) / (p max -p min ).
[0060] Corresponding to the example of step S203, further processing is performed, assuming ω = 0.5, threshold truncation processing is performed, and then normalization processing is performed: p max =1, p min =0, and the normalized result is: S′=[(1,0,0),(1,1,0.6),(0,1,1)].
[0061] It should be noted that since the example image is a black and white image (the eigenvalues are 1 and 0), the normalization result is consistent with that before normalization.
[0062] Step S104: obtaining a global adversarial perturbation of the original image sample based on a global gradient attack algorithm.
[0063] Calculate the gradient of the AI model loss function J(θ,x,y) for the original image sample Among them, θ is the parameter of the artificial intelligence model, x is the original image sample, and y is the true label of the original image sample. Represents the derivative of the loss function. The gradient of the original image sample is obtained by the sign function sign(·). Perform a sign operation and multiply it by the preset constant ε to obtain the global adversarial disturbance η.
[0064] The specific formula is Among them, the preset constant ε can be understood as; ε is a hyperparameter used to control the disturbance amplitude, which is set according to actual conditions and is not restricted here.
[0065] It can be understood that the cross entropy loss function is used in the artificial intelligence model F, that is, if there are C categories, the true label is y, and the model prediction is Then the cross entropy loss function of a single sample is:
[0066] It calculates the gradient of the input original image sample x: Among them, δ ii is the Kronecker product, which is 1 when i=j and 0 otherwise.
[0067] For example, suppose that the partial gradient of x is obtained
[0068] The sign operation of the gradient is sign{[(0.8,0.3,-1.1),(1.3,0.5,0.8),(-1.6,-1.1,0.9)]}=[(1,1,-1),(1,1,1),(-1,-1,1)].
[0069] Assume ε = 0.03, so the global adversarial perturbation is: η = 0.03 × [(1, 1, -1), (1, 1, 1), (-1, -1, 1)] = [(0.03, 0.03, -0.03), (0.03, 0.03, 0.03), (-0.03, -0.03, 0.03)].
[0070] Step S105, superimposing the global adversarial perturbation and the processed enhanced feature saliency map to obtain a local adversarial perturbation of the original image sample.
[0071] Multiply the global adversarial perturbation η by the processed enhanced feature saliency map S′ to obtain the local adversarial perturbation η′=η·S′.
[0072] Illustratively, η′=[(0.03, 0.03, -0.03), (0.03, 0.03, 0.03), (-0.03, -0.03, 0.03)]·[(1, 0, 0), (1, 1, 0.6), (0, 1, 1)]=[(0.03, 0, 0), (0.03, 0.03, 0.018), (0, -0.03, 0.03)].
[0073] It should be noted that the superposition processing of the present application takes multiplication as an example, and addition or other processing methods are also possible, which are not limited here.
[0074] Step S106: Add the local adversarial perturbation to the original image sample to generate an adversarial sample.
[0075] Add the local adversarial perturbation η′ to the original image sample x to generate the adversarial sample x adv =x+η′;
[0076] In one possible implementation, if the decision result of the artificial intelligence model on the adversarial sample is inconsistent with the true label y of the original image sample, then the adversarial sample x is determined adv Generate successfully;
[0077] Another possible implementation scenario is that if the decision result of the artificial intelligence model for the adversarial sample is consistent with the true label y of the original image sample, then the local adversarial perturbation η′ is added again until the decision result of the artificial intelligence model for the adversarial sample is inconsistent with the true label of the original image sample, and it is determined that the adversarial sample is generated successfully.
[0078] For example, η′=[(0.03, 0, 0), (0.03, 0.03, 0.018), (0, -0.03, 0.03)], x adv =x+η′, if x adv The decision result of x is inconsistent with y. adv is an adversarial sample. If the first generated x adv The decision result (that is, F(x adv )) is consistent with y, then add a local adversarial perturbation η′, then x adv =x+η′+η′, and if the comparison is still consistent, add local adversarial perturbation η′, then x adv =x+η′+η′+η′, until they are inconsistent, and adversarial samples are generated. For details, please refer to Figure 3 understand.
[0079] In summary, in this application, by fusing the results of the gradient-based and perturbation-based interpretable algorithms to generate a fused significant feature map, it is possible to determine feature points with strong significance, effectively reducing the randomness and bias brought by a single interpretable algorithm; then, by adding local adversarial perturbations in small amounts and multiple times, combined with judging whether the decision results of the artificial intelligence model for the adversarial samples are consistent with the true labels of the original image samples, redundant adversarial perturbations are reduced, and the concealment of adversarial samples is effectively improved.
[0080] Combination of the above Figure 1-Figure 3 The adversarial sample generation method based on multi-feature saliency map fusion provided by the embodiment of the present application is described in detail. Figure 4 A detailed description is given of an adversarial sample generation device based on multi-feature saliency map fusion provided in an embodiment of the present application.
[0081] Figure 4 This is a schematic diagram of the structure of the adversarial sample generation device based on multi-feature saliency map fusion provided in the embodiment of the present application. Figure 1 For example, Figure 4 As shown, the adversarial sample generation device 400 based on multi-feature saliency map fusion includes: a transceiver module 401 and a processing module 402. For ease of description, Figure 4 Only the main components of the adversarial sample generation device based on multi-feature saliency map fusion are shown.
[0082] Among them, the transceiver module 401 is used to perform the transceiver function of the above-mentioned adversarial sample generation method based on multi-feature saliency map fusion, and the processing module 402 is used to perform other functions of the above-mentioned adversarial sample generation method based on multi-feature saliency map fusion except the transceiver function.
[0083] Optionally, the transceiver module 401 may include a sending module ( Figure 4 ) and a receiving module ( Figure 4 (not shown in the figure). The sending module is used to implement the sending function of the adversarial sample generation device 400 based on the fusion of multiple feature saliency maps, and the receiving module is used to implement the receiving function of the adversarial sample generation device 400 based on the fusion of multiple feature saliency maps.
[0084] Optionally, the adversarial sample generation device 400 based on multi-feature saliency map fusion may further include a storage module ( Figure 4 ), the storage module stores a program or instruction. When the processing module 402 executes the program or instruction, the adversarial sample generation device 400 based on multi-feature saliency map fusion can execute the adversarial sample generation method based on multi-feature saliency map fusion in the embodiment of the present application.
[0085] Combine the following Figure 5 Each component of the adversarial sample generation device 500 based on multi-feature saliency map fusion is specifically introduced:
[0086] The processor 501 is the control center of the adversarial sample generation device 500 based on multi-feature saliency map fusion, which can be a processor or a general term for multiple processing elements. For example, the processor 501 is one or more central processing units (CPUs), or an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application, such as one or more microprocessors (digital signal processors, DSPs), or one or more field programmable gate arrays (FPGAs).
[0087] Optionally, the processor 501 can perform various functions of the adversarial sample generation device 500 based on multi-feature saliency map fusion by running or executing a software program stored in the memory 502, and calling data stored in the memory 502, such as executing the adversarial sample generation method based on multi-feature saliency map fusion in the embodiment of the present application.
[0088] In a specific implementation, as an embodiment, the processor 501 may include one or more CPUs, such as Figure 5 CPU0 and CPU1 are shown in FIG.
[0089] In a specific implementation, as an embodiment, the adversarial sample generation device 500 based on multi-feature saliency map fusion may also include multiple processors, such as Figure 5The processor 501 and processor 504 shown in . Each of these processors can be a single-core processor (single-CPU) or a multi-core processor (multi-CPU). The processor here can refer to one or more devices, circuits, and / or processing cores for processing data (such as computer program instructions). Among them, the memory 502 is used to store the software program for executing the solution of the present application, and the execution is controlled by the processor 501. The specific implementation method can refer to the above method embodiment, which will not be repeated here.
[0090] Optionally, the memory 502 may be a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM) or other types of dynamic storage devices that can store information and instructions, or an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compressed optical disc, laser disc, optical disc, digital versatile disc, Blu-ray disc, etc.), a magnetic disk storage medium or other magnetic storage device, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 502 may be integrated with the processor 501, or may exist independently, and may be connected to the processor 501 through the interface circuit ( Figure 5 (not shown) is coupled to the processor 501, which is not specifically limited in the embodiment of the present application.
[0091] The transceiver 503 is used for communication with other communication devices. For example, the adversarial sample generation device 500 based on multi-feature saliency map fusion is a first device, and the transceiver 503 can be used for communication with a second device or a third device.
[0092] Optionally, the transceiver 503 may include a receiver and a transmitter ( Figure 5 The receiver is used to implement a receiving function, and the transmitter is used to implement a sending function.
[0093] Optionally, the transceiver 503 may be integrated with the processor 501, or may exist independently, and may be connected to the adversarial sample generation device 500 based on multi-feature saliency map fusion through an interface circuit ( Figure 5 (not shown) is coupled to the processor 501, which is not specifically limited in the embodiment of the present application.
[0094] Understandably, Figure 5 The structure of the adversarial sample generation device 500 based on multi-feature saliency map fusion shown in the figure does not constitute a limitation of the adversarial sample generation device based on multi-feature saliency map fusion. The actual adversarial sample generation device based on multi-feature saliency map fusion may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently.
[0095] In addition, the technical effects of the adversarial sample generation device 500 based on multi-feature saliency map fusion can refer to the technical effects of the method described in the above method embodiment, and will not be repeated here.
[0096] It should be understood that the processor in the embodiment of the present application may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field programmable gate arrays (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. It should also be understood that the memory in the embodiment of the present application may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic random access memory (DRAM), synchronous DRAM (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link DRAM (SLDRAM), and direct rambus RAM (DR RAM).
[0097] The above embodiments can be implemented in whole or in part by software, hardware (such as circuits), firmware or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website site, computer, server or data center to another website site, computer, server or data center by wired (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that contains one or more available media sets. The available medium can be a magnetic medium (for example, a floppy disk, a hard disk, a tape), an optical medium (for example, a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state hard disk.
Claims
1. A method for generating adversarial samples based on multi-feature saliency map fusion, characterized in that: The method comprises: Acquire multiple image sample data, and divide the multiple image sample data into a training set, a test set, and a validation set, train the artificial intelligence model through the training set, adjust the parameters of the artificial intelligence model through the validation set, and select the optimal model parameters through the test set; Input any original picture sample into the artificial intelligence model, obtain a first feature saliency map based on a gradient interpretability method, obtain a second feature saliency map based on a perturbation interpretability method, and fuse and enhance the first feature saliency map and the second feature saliency map to obtain an enhanced feature saliency map; performing truncation and normalization processing on the enhanced feature saliency map to obtain a processed enhanced feature saliency map; Based on a global gradient attack algorithm, a global adversarial perturbation of the original image sample is obtained; Superimposing the global adversarial perturbation and the processed enhanced feature saliency map to obtain a local adversarial perturbation of the original image sample; The local adversarial perturbation is added to the original image sample to generate an adversarial sample.
2. The adversarial sample generation method based on multi-feature saliency map fusion according to claim 1 is characterized in that: The dividing the plurality of image sample data into a training set, a test set, and a validation set comprises: The plurality of image sample data are subjected to stratified sampling and divided into a training set, a test set and a validation set according to a preset ratio.
3. The adversarial sample generation method based on multi-feature saliency map fusion according to claim 1 is characterized in that: The gradient-based interpretability method obtains a first feature saliency map, including: Forward propagation of the original image sample in the artificial intelligence model to obtain the output feature map of the last convolutional layer; Perform back propagation based on the prediction score of a specific category to obtain the gradient of the prediction score in the output feature map; Performing global average pooling on the gradient to obtain the weight of the output feature map in each channel; Based on the weight of each channel, weighted summation is performed on the output feature map to obtain a class activation map; Filtering the class activation map based on an activation function to remove negative values; Interpolation processing is performed on the processed class activation map so that the processed class activation map has the same size as the original image sample, thereby obtaining a first feature saliency map.
4. The adversarial sample generation method based on multi-feature saliency map fusion according to claim 1 is characterized in that: The perturbation-based interpretability method obtains a second feature saliency map, including: Randomly generate multiple occlusion templates, where each occlusion template covers a portion of the original image sample and the rest remains unchanged; Based on the occlusion template, a plurality of occluded original image samples are obtained; Using the artificial intelligence model, predict each occluded original image sample and record the prediction results; Comparing the prediction results of the original picture sample and the occluded original picture sample to obtain the importance score of each feature point or area in the original picture sample to the artificial intelligence model; Based on the importance score of each feature point or region to the artificial intelligence model, a second feature saliency map is obtained.
5. The adversarial sample generation method based on multi-feature saliency map fusion according to claim 1 is characterized in that: The fusing and enhancing the first feature saliency map and the second feature saliency map to obtain an enhanced feature saliency map includes: The first feature saliency map and the second feature saliency map are fused and enhanced by a weighted average fusion formula to obtain an enhanced feature saliency map. Among them, the weighted average fusion formula is S=λ·S1+(1-λ)·S2, S1 is the first feature saliency map, S2 is the second feature saliency map, S is the enhanced feature saliency map, λ is the weight coefficient, and the value range is [0,1].
6. The adversarial sample generation method based on multi-feature saliency map fusion according to claim 1 is characterized in that: The truncating and normalizing the enhanced feature saliency map comprises: The enhanced feature saliency map is subjected to a threshold truncation process, comprising: If the feature saliency value in the enhanced feature saliency map is less than a preset threshold, setting the feature saliency value to zero; If the feature saliency value in the enhanced feature saliency map is greater than a preset threshold, retaining the feature saliency value; The feature saliency map after threshold truncation is normalized, including: The maximum and minimum feature saliency values in the enhanced feature saliency map are obtained, and each feature saliency value is normalized by a normalization formula, wherein the normalization formula is p=(pp min ) / (p max -p min ), p max is the maximum eigensignificant value, p min is the minimum feature significance value.
7. The adversarial sample generation method based on multi-feature saliency map fusion according to claim 1 is characterized in that: The step of obtaining a global adversarial perturbation of the original image sample based on a global gradient attack algorithm includes: Calculate the gradient of the artificial intelligence model loss function to the original image sample, where the loss function is J(θ, x, y), θ is the parameter of the artificial intelligence model, x is the original image sample, y is the true label of the original image sample, and the gradient of the original image sample is represents the derivation of the loss function; A sign operation is performed on the gradient of the original image sample through a sign function, and the gradient is multiplied by a preset constant to obtain a global adversarial perturbation of the original image sample, wherein the sign function is sign(·).
8. The method for generating adversarial samples based on multi-feature saliency map fusion according to claim 1, characterized in that: The adding the local adversarial perturbation to the original image sample to generate an adversarial sample includes: Adding the local adversarial perturbation to the original image sample to generate an adversarial sample; If the decision result of the artificial intelligence model for the adversarial sample is inconsistent with the true label of the original image sample, it is determined that the adversarial sample is successfully generated; If the decision result of the artificial intelligence model for the adversarial sample is consistent with the true label of the original image sample, the local adversarial perturbation is added again until the decision result of the artificial intelligence model for the adversarial sample is inconsistent with the true label of the original image sample, and it is determined that the adversarial sample is generated successfully.
9. A device for generating adversarial samples based on multi-feature saliency map fusion, characterized in that: The apparatus comprises: a module for executing the method according to any one of claims 1-8.
Citation Information
Patent Citations
Adversarial sample generation method and device of specified label, electronic equipment and medium
CN111340180A
White-box attack resisting method based on gradient and triple difference fusion
CN118821117A