Decoupling adversarial sample generation method

By defining the decoupling principle and building multi-objective optimization functions, the generative model is trained to generate adversarial samples that meet the decoupling principle, the problem of undecoupling of adversarial samples in the existing technology is solved, and the difference and aggressiveness of adversarial samples are improved.

CN120047769APending Publication Date: 2025-05-27NAT INNOVATION INST OF DEFENSE TECH PLA ACAD OF MILITARY SCI
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510089521.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The adversarial samples generated in the prior art cannot be decoupled, especially on the model side, resulting in high similarity in distributions of adversarial samples and does not help to understand adversarial samples.

Method used

By defining the decoupling principle, decoupling adversarial samples into the sum of clean samples and adversarial perturbations, multi-objective optimization functions and loss functions are constructed, and the generative model is trained to generate adversarial samples that meet the decoupling principle.

Benefits of technology

The decoupling of adversarial samples is achieved, the difference and aggressiveness of adversarial samples are improved, and the problem of undecoupling of adversarial samples in the existing technology is solved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047769A_ABST
    Figure CN120047769A_ABST
Patent Text Reader

Abstract

The invention discloses a decoupling adversarial sample generation method. The method comprises the following steps: decoupling an adversarial sample by taking a condition of decoupling the adversarial sample and obtaining a sum of a clean sample and adversarial disturbance as a decoupling principle; constructing a multi-objective optimization function, wherein the multi-objective optimization function comprises a plurality of constraint conditions; constructing a plurality of loss functions to enable the constructed plurality of loss functions to correspond to a plurality of constraint conditions, and training and updating the generative model by using the weighted sum of the plurality of loss functions as a final loss function to obtain a generative model meeting a decoupling principle confrontation sample; evaluating the obtained generative model meeting the decoupling principle confrontation sample by using an evaluation index until the evaluation result meets the decoupling principle and successfully attacks the target model at the same time; and generating a confrontation sample which can simultaneously meet the decoupling principle and aggressiveness by using the trained generative model. According to the invention, the problem that the adversarial sample cannot be decoupled, especially the problem that the adversarial sample cannot be decoupled at the model side, can be solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of adversarial examples and deep neural networks, and particularly relates to a decoupled adversarial example generation method. Background Art

[0002] In recent years, artificial intelligence technologies centered around deep neural networks have been widely developed and are widely applied in real life, such as object detection and object tracking technologies in autonomous driving vehicles, face recognition and face matching technologies in face payment, object detection technologies in video surveillance, and so on. However, deep neural networks have been proven to be vulnerable to subtle noises introduced into the input. For example, adding noises invisible to the human eye to pictures can mislead the classifier into inferring incorrect results, which poses great security risks to the widespread deployment of systems based on deep neural network technologies in real life, especially for scenarios with high security requirements, such as face payment and video surveillance. Some studies have shown that using adversarial examples as a supplement to training data and fine-tuning the deep neural network can greatly improve the adversarial robustness of the network. Therefore, developing adversarial attack technologies with greater attack intensity and greater differences is more helpful for the robustness of the model.

[0003] Current methods for generating adversarial examples can be divided into white-box attack methods and black-box attack methods. White-box attack methods assume that the attacker can obtain all the information of the target victim model, such as model structure, model parameters, and training data. In contrast, black-box attack methods assume that the attacker cannot obtain the information of the target victim model and only allow querying the inference results of the model. Under the above assumptions, most white-box attack methods are based on model gradient methods. A typical method is the Fast Gradient Sign Method (FGSM), which generates adversarial perturbations in a single step along the gradient ascent direction of the model loss. On this basis, many improved versions have emerged, such as the method IFGSM that generates adversarial perturbations in multiple steps with smaller step sizes, the method PGD that updates from noisy data samples instead of clean samples, and the method MIFGSM that introduces a momentum term to stabilize the update steps. In addition, some researchers have also constructed adversarial examples by using generative models, such as generative adversarial networks, autoencoders, diffusion models, etc. Relatively speaking, due to only being able to obtain the inference results of the target victim model, black-box attack methods usually use evolutionary algorithms to solve. Common methods include searching for adversarial perturbations using differential evolution, particle swarm optimization algorithms, gradient estimation algorithms, and so on. Among these adversarial example attack methods, white-box attack methods can achieve better attack performance, such as higher attack success rates and higher attack success rates against unknown models, that is, attack transferability.

[0004] Although there are already various adversarial example generation methods, these methods only focus on a single metric, i.e., the success rate of adversarial attacks. Therefore, the generated adversarial examples have high similarity in distribution and are not helpful for understanding adversarial examples. In addition, there is currently an issue of non-decoupling in adversarial examples. Summary of the Invention

[0005] To solve some or all of the technical problems existing in the above-mentioned prior art, the present invention provides a decoupled adversarial example generation method, which can solve the problem of non-decoupling of current adversarial examples in the prior art, especially the problem of non-decoupling of current adversarial examples on the model side.

[0006] The technical solution of the present invention is as follows:

[0007] Provided is a decoupled adversarial example generation method, including:

[0008] Taking the condition of decoupling an adversarial example into the sum of a clean sample and an adversarial perturbation as the decoupling principle to decouple the adversarial example;

[0009] Constructing a multi-objective optimization function, where the multi-objective optimization function includes multiple constraint conditions;

[0010] Constructing multiple loss functions such that the constructed multiple loss functions respectively correspond to multiple said constraint conditions, and using the weighted sum of the multiple loss functions as the final loss function to train and update the generation model to obtain a generation model of adversarial examples that satisfies the decoupling principle;

[0011] Using the trained generation model to generate adversarial examples that can simultaneously satisfy the decoupling principle and the aggressiveness.

[0012] Further, in the above-mentioned decoupled adversarial example generation method, the multi-objective optimization function includes a dual-objective optimization function, the dual-objective optimization function includes three constraint conditions, and the dual-objective optimization function includes:

[0013]

[0014] Where: minδ represents the optimization objective, aiming to find an adversarial perturbation; f(x)≠f(x + δ) is the first constraint condition, indicating that the adversarial perturbation needs to mislead the neural network into making a wrong judgment; F(F(x + δ)) = f(F(x) + F(δ)) is the second constraint condition, indicating that the adversarial perturbation needs to satisfy the decoupling principle; ||δ|| p ≤∈ is the third constraint condition, indicating that the adversarial perturbation needs to be constrained within a preset amplitude to ensure that it cannot be detected by humans, ||·|| pLet ∥·∥ₚ denote the p-norm, δ denote the adversarial perturbation, ∈ denote the maximum magnitude used to constrain the adversarial perturbation, F(x) denote the predicted logical value of the model for x, F(δ) denote the predicted logical value of the model for δ, f(x) denote the deep neural network model containing the input sample x, and f(x + δ) denote the deep neural network model containing the input sample x and the adversarial perturbation δ.

[0015] Further, in the above decoupled adversarial sample generation method, the generation model includes an autoencoder or a diffusion model.

[0016] Further, in the above decoupled adversarial sample generation method, the loss function includes an adversarial loss function, a decoupling loss function, and an imperceptibility loss function. The final loss function is the sum of the adversarial loss function, the decoupling loss function, and the imperceptibility loss function, where

[0017] The adversarial loss function includes:

[0018]

[0019] In the above formula, denotes the adversarial loss function, F y (·) denotes the response value of the model for the input sample at index y, max i≠y F i (x + δ) denotes the maximum response value except for index y, and c denotes the attack strength;

[0020] The decoupling loss function includes:

[0021]

[0022] In the above formula, denotes the decoupling loss function, denotes the index of the maximum response value in the clean sample, denotes the index of the maximum response value in the adversarial sample;

[0023] The imperceptibility loss function includes:

[0024]

[0025] In the above formula, denotes the imperceptibility loss function;

[0026] The final loss function includes:

[0027]

[0028] In the above formula, denotes the final loss function, and α and β denote weights.

[0029] Furthermore, in the above decoupled adversarial sample generation method, multiple loss functions are constructed such that the constructed multiple loss functions respectively correspond to multiple said constraint conditions, and the weighted sum of the multiple loss functions is used as the final loss function. Training and updating the generation model includes:

[0030] Training the generation model using a two-stage training method of a first stage and a second stage;

[0031] The first stage includes optimizing the generation model with an adversarial loss and an imperceptibility loss, such that the generation model can generate adversarial strict samples with adversarial properties;

[0032] The second stage includes training the generation model using the final loss function, such that the generation model can generate adversarial samples that are both adversarial and satisfy the decoupling principle.

[0033] Furthermore, in the above decoupled adversarial sample generation method, it also includes evaluating the generation model of the adversarial samples that satisfy the decoupling principle using evaluation metrics until the evaluation results simultaneously satisfy the decoupling principle and the deception rate of the decoupled adversarial samples.

[0034] Furthermore, in the above decoupled adversarial sample generation method, the evaluation metrics include:

[0035]

[0036] where FA represents the evaluation metric, ACC ops represents the decoupling principle metric of the decoupled adversarial samples, and FR represents the deception rate of the decoupled adversarial samples.

[0037] Furthermore, in the above decoupled adversarial sample generation method, the decoupling principle metric of the decoupled adversarial samples is calculated and determined by the following formula:

[0038]

[0039] where i represents the i-th sample, N represents the total number of test samples, represents the sign function. If the condition in is true, then 1 is output, otherwise 0 is output. x i represents the clean sample obtained from the decoupled adversarial sample, F(x i ) represents the logit output corresponding to the model, δ i represents the adversarial perturbation obtained from the i-th decoupled adversarial sample that satisfies the decoupling principle, F(δ i ) represents the logit output corresponding to the adversarial perturbation, represents the i-th adversarial sample generated using the generation model, Represents the logit output corresponding to its model.

[0040] Furthermore, in the above decoupled adversarial sample generation method, the deception rate of the decoupled adversarial sample is calculated and determined by the following formula:

[0041]

[0042] In the above formula, f(x i ) represents the predicted class output by the model.

[0043] The main advantages of the technical solution of the present invention are as follows:

[0044] A decoupled adversarial sample generation method of the present invention first defines a decoupling principle, that is, an adversarial sample can be decoupled into the sum of a clean sample and an adversarial perturbation at the model prediction end, so as to be consistent with the construction method of the adversarial sample. Secondly, the present invention designs a decoupling loss and constructs a multi-objective optimization function. Then, the present invention designs a generation model, and under the constraints of the adversarial loss and the decoupling loss, successfully trains a generation model that can generate adversarial samples satisfying the decoupling principle. Finally, an evaluation index that can simultaneously measure the adversarial attackability and the satisfaction of the decoupling principle is designed. Through the above method, the problem that the current adversarial samples in the prior art cannot be decoupled, especially the problem that the current adversarial samples cannot be decoupled on the model side, can be solved. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] The drawings described herein are used to provide a further understanding of the embodiments of the present invention and constitute a part of the present invention. The illustrative embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the attached

[0046] In the figure:

[0047] Figure 1 is a schematic flowchart of a decoupled adversarial sample generation method according to an embodiment of the present invention.

[0048] Figure 2 is a schematic diagram of the principle of adversarial sample generation in a decoupled adversarial sample generation method provided by an embodiment of the present invention; DETAILED DESCRIPTION OF THE EMBODIMENTS

[0049] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below in conjunction with the specific embodiments of the present invention and the corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0050] The following will, in conjunction with the accompanying Figure 1-2 , elaborate in detail on the technical solution provided by the embodiments of the present invention.

[0051] As shown in the accompanying Figure 1 , the embodiments of the present invention provide a decoupled adversarial sample generation method, which can generate adversarial samples with a greater distribution difference from existing adversarial samples. The method includes the following steps S1 - S4:

[0052] To enable a clearer understanding and description of the technical solution of the present invention, first introduce the definition of adversarial samples in the present invention: Given a trained deep neural network model f, for an input sample x with label y, construct a subtle adversarial perturbation δ, and add it to the clean sample to construct an adversarial sample x adv = x + δ, which will not attract human attention visually but can mislead the deep neural network model f to make an incorrect identification. The optimization method of the adversarial perturbation δ can be obtained by solving the following optimization objective:

[0053]

[0054] where L represents the training loss function of the deep neural network, ||·|| p represents the p - norm, and ∈ represents the maximum amplitude used to constrain the adversarial perturbation. The meaning of the above formula is to optimize the adversarial perturbation δ that can maximize the loss of the deep neural network while ensuring that the amplitude of the adversarial perturbation δ is within a given threshold.

[0055] The following will elaborate in detail on the technical solution of the present invention, including:

[0056] Step S1: Take the condition of decoupling the adversarial sample to obtain the sum of the clean sample and the adversarial perturbation as the decoupling principle to decouple the adversarial sample;

[0057] As can be seen from the above definition, the adversarial sample x adv is the sum of the clean sample x and the adversarial perturbation δ. The present invention defines the following decoupling principle:

[0058] Decoupling principle: At the model prediction end, an adversarial sample that satisfies f(x + δ) = f(x) + f(δ) is called a sample that satisfies the decoupling principle.

[0059] It can be seen that the decoupling principle, at the model prediction end, formally decouples the prediction result of the adversarial sample into the sum of the prediction results of the clean sample and the adversarial perturbation, making it consistent with the definition of the adversarial sample.

[0060] For a classification task based on a deep neural network, since f(·) represents the predicted class, when the predicted class of either f(x) or f(δ) is 1000, as long as the predicted class of the other item exceeds 0, the decoupling principle cannot hold. Therefore, the present invention relaxes the decoupling principle for the classification task and re-expresses the decoupling principle in the following form: At the model prediction end, an adversarial sample that satisfies f(F(x+δ)) = f(F(x)+F(δ)) is called a sample that satisfies the decoupling principle.

[0061] In the redefined decoupling principle, F(x) represents the predicted logical value of the model for x, which is a 1000-dimensional vector used to characterize the model's response to the input, and satisfies f(x) = argmaxF(x). The argmax(x) function is used to find the index corresponding to the maximum value in the input variable x in the vector.

[0062] Step S2: Construct a multi-objective optimization function, and the multi-objective optimization function includes multiple constraint conditions;

[0063] As an example, in order to generate an adversarial sample that is both adversarially attackable and satisfies the decoupling principle, the present invention sets the multi-objective optimization function as a dual-objective optimization function. This dual-objective optimization function includes three constraint conditions, and the dual-objective optimization function includes:

[0064]

[0065] Among them: minδ represents the optimization objective, aiming to find an adversarial perturbation; f(x)≠f(x+δ) is the first constraint condition, indicating that the adversarial perturbation needs to mislead the neural network to make a wrong judgment; f(F(x+δ)) = f(F(x)+F(δ)) is the second constraint condition, indicating that the adversarial perturbation needs to satisfy the decoupling principle; ||δ|| p ≤∈ is the third constraint condition, indicating that the adversarial perturbation needs to be constrained within a preset amplitude to ensure that it cannot be detected by humans. ||·|| p represents the p-norm, δ represents the adversarial perturbation, ∈ represents the maximum amplitude used to constrain the adversarial perturbation, F(x) represents the predicted logical value of the model for x, F(δ) represents the predicted logical value of the model for δ, f(x) represents the deep neural network model including the input sample x, and f(x+δ) represents the deep neural network model including the input sample x and the adversarial perturbation δ.

[0066] Step S3: Construct multiple loss functions so that the constructed multiple loss functions respectively correspond to multiple constraint conditions, and use the weighted sum of the multiple loss functions as the final loss function to train and update the generation model to obtain a generation model of adversarial samples that satisfies the decoupling principle;

[0067] Specifically, the loss function includes an adversarial loss function, a decoupling loss function, and an imperceptible loss function. The final loss function is the sum of the adversarial loss function, the decoupling loss function, and the imperceptible loss function. Among them,

[0068] The adversarial loss function includes:

[0069]

[0070] In the above formula, represents the adversarial loss function, and F y (·) represents the response value of the model to the input sample at index y, and max i≠y F i (x + δ) represents the maximum response value except for index y, and c represents the attack strength.

[0071] Thus, when the adversarial perturbation is not aggressive, the response value at index y is the maximum value, that is, the predicted value of the model is equal to the label of the data, and the loss function is greater than 0 at this time. When the adversarial perturbation is aggressive, the response value at index y is no longer the maximum value, and the loss function becomes negative. At this time, the parameter c is used to control the amplitude of pushing the clean sample to the other side of the decision boundary by the adversarial perturbation, that is, the attack strength.

[0072] Analyzing the decoupling principle, it can be found that it is only necessary to ensure that the maximum response value of the model for the adversarial sample is consistent with the maximum response value of the sum of the clean sample and the adversarial perturbation. Therefore, the decoupling loss function includes:

[0073]

[0074] In the above formula, represents the decoupling loss function, represents the index of the maximum response value in the clean sample, represents the index of the maximum response value in the adversarial sample.

[0075] Thus, by minimizing the decoupling loss defined by the above decoupling loss function , the predicted value of can be made consistent with the adversarial sample, that is, the decoupling principle is satisfied.

[0076] To ensure the imperceptibility of the adversarial perturbation, the present invention designs an imperceptible loss based on the second norm. The imperceptible loss function includes:

[0077]

[0078] In the above formula, represents the imperceptible loss function.

[0079] Thus, using the above formula The defined imperceptible loss function can ensure that during the optimization process, the amplitude of the adversarial perturbation will not be too large to cause human perception due to the perturbation.

[0080] The final loss function includes:

[0081]

[0082] In the above formula, represents the final loss function, and α and β represent weights.

[0083] Specifically, training and updating the generation model includes:

[0084] Training the generation model using a two-stage training method in the first stage and the second stage;

[0085] The first stage includes optimizing the generation model with adversarial loss and imperceptible loss, so that the generation model can generate adversarial strict samples with adversarial properties;

[0086] The second stage includes training the generation model using the final loss function, so that the generation model can generate adversarial samples that are both adversarial and satisfy the decoupling principle.

[0087] Specifically, the training process of the generation model of the present invention is carried out by the following method:

[0088] Input: Training dataset (X, Y), the training dataset includes public datasets such as ImageNet, including training sample images and labels, target victim model f, batch input size m, maximum number of iterations MaxEpoch, alternating training steps n;

[0089] Sample m pairs of training samples (x, y) from the training data (each time, sample m pairs of training samples and labels from the training dataset), and generate an adversarial perturbation G(x);

[0090] Construct an adversarial sample x adv = x + G(x)

[0091] Convert the adversarial perturbation to the interval [0, 1] through (δ + 1) * 0.5;

[0092] Obtain F(x adv ), F(x), F(δ)

[0093]

[0094] If epoch ≤ n then

[0095] Use to update the generation model loss or use to update the generation model loss.

[0096] To measure the performance of the adversarial examples generated by the present invention, the present invention uses evaluation metrics to evaluate the obtained generation model of adversarial examples that satisfy the decoupling principle.

[0097] Step S4: Use the trained generation model to generate adversarial examples that can satisfy both the decoupling principle and the aggressiveness.

[0098] Specifically, the adversarial example generation method includes:

[0099] Input the clean sample into the generation model to generate an adversarial perturbation;

[0100] Add the generated adversarial perturbation to the clean sample to obtain an adversarial example.

[0101] Specifically, to solve the optimization problem defined by the dual objective function The present invention proposes to use a generation model to generate adversarial examples. Specifically, as Figure 2 shown, the present invention first inputs the clean sample x into the generation model G to generate an adversarial perturbation G(x), and adds the adversarial perturbation to the clean sample x to form an adversarial example x adv = G(x) + x, input the adversarial example into the deep neural network and calculate the loss, and finally update the parameters of the generation model G according to the value of the loss function.

[0102] As an example, input the adversarial example into the classification neural network to calculate the corresponding loss through the above three loss functions, and update the parameters of the generation model G according to the calculated value.

[0103] In some alternative implementation manners of this embodiment, the generation model includes an autoencoder or a diffusion model.

[0104] In order to make the decoupled adversarial example generation method of the present invention satisfy both the decoupling principle and the deception rate of the decoupled adversarial examples, the present invention also sets evaluation metrics to evaluate the obtained adversarial examples. The methods for evaluation include:

[0105] Use the evaluation metrics to evaluate the obtained generation model of adversarial examples that satisfy the decoupling principle until the evaluation results satisfy both the decoupling principle and the deception rate of the decoupled adversarial examples;

[0106] Specifically, the evaluation metrics include:

[0107]

[0108] Among them, FA represents the evaluation metric, ACC ops represents the decoupling principle metric of the decoupled adversarial example, and FR represents the deception rate of the decoupled adversarial example.

[0109] Among them, the decoupling principle index of the decoupled adversarial samples is calculated and determined by the following formula:

[0110]

[0111] Among them, i represents the i-th sample, and N represents the total number of test samples. denotes the sign function. If the condition in is true, then 1 is output; otherwise, zero is output. x i represents the clean sample obtained from the decoupled adversarial sample, and F(x i ) represents the logit output corresponding to the model. δ i represents the adversarial perturbation obtained from the i-th adversarial sample that satisfies the decoupling principle, and F(δ i ) represents the logit output corresponding to the adversarial perturbation. represents the i-th adversarial sample generated by using the generative model. represents the logit output corresponding to its model.

[0112] Among them, the deception rate of the decoupled adversarial sample is calculated and determined by the following formula:

[0113]

[0114] In the above formula, f(x i ) represents the predicted class output by the model.

[0115] It should be noted that FR and ACC ops are a pair of contradictory indicators. When the adversarial perturbation has no effect, the prediction result of the clean sample is not equal to the prediction result of the adversarial sample. At this time, the latter can achieve a very high attack success rate. However, it cannot be guaranteed that all adversarial samples with aggressiveness satisfy the decoupling principle. For this reason, the present invention designs the following indicators to simultaneously reflect the relationship between the two.

[0116] Thus, for a method for generating decoupled adversarial samples of the present invention, first, the decoupling principle is defined, that is, at the model prediction end, the adversarial sample can be decoupled into the sum of a clean sample and an adversarial perturbation, so as to keep it consistent with the construction method of the adversarial sample. Secondly, the present invention designs a decoupling loss and constructs a multi-objective optimization function. Then, the present invention designs a generative model, and under the constraints of the adversarial loss and the decoupling loss, successfully trains a generative model that can generate adversarial samples satisfying the decoupling principle. Finally, an evaluation index that can simultaneously measure the adversarial aggressiveness and the satisfaction of the decoupling principle is designed. Through the method of the present invention, the problem that the current adversarial samples in the prior art cannot be decoupled, especially the problem that the current adversarial samples cannot be decoupled on the model side, can be solved.

[0117] It should be noted that in this text, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. In addition, in this text, "front", "rear", "left", "right", "upper" and "lower" are all referenced with respect to the placement state shown in the drawings.

[0118] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A decoupled adversarial sample generation method, characterized in that: include: Decoupling adversarial samples and obtaining the sum of clean samples and adversarial perturbations is used as the decoupling principle to decouple adversarial samples. Constructing a multi-objective optimization function, wherein the multi-objective optimization function includes a plurality of constraints; Constructing multiple loss functions so that the constructed multiple loss functions correspond to the multiple constraints respectively, and using the weighted sum of the multiple loss functions as the final loss function to train and update the generation model to obtain a generation model that satisfies the decoupling principle and the adversarial sample; Using the trained generative model, adversarial samples are generated that can satisfy both the decoupling principle and aggressiveness.

2. The decoupled adversarial sample generation method according to claim 1, characterized in that: The multi-objective optimization function includes a dual-objective optimization function, and the dual-objective optimization function includes three constraints. The dual-objective optimization function includes: minδ stf(x)≠f(x+δ) f(F(x+δ))=f(F(x)+F(δ)) ||d|| p ≤∈ Among them: minδ represents the optimization goal, which aims to find an adversarial disturbance; f(x)≠f(x+δ) is the first constraint, which means that the adversarial disturbance needs to mislead the neural network to make wrong judgments; f(F(x+δ))=f(F(x)+F(δ)) is the second constraint, which means that the adversarial disturbance needs to satisfy the decoupling principle; ||δ|| p ≤∈ is the third constraint, indicating that the adversarial perturbation needs to be constrained within a preset amplitude to ensure that it is not perceived by humans, ||·|| p represents the p-norm, δ represents the adversarial perturbation, ∈ represents the maximum amplitude used to constrain the adversarial perturbation, F(x) represents the model's predicted logical value for x, F(δ) represents the model's predicted logical value for δ, f(x) represents a deep neural network model containing input sample x, and f(x+δ) represents a deep neural network model containing input sample x and adversarial perturbation δ.

3. The decoupled adversarial sample generation method according to claim 1, characterized in that: The generative model includes an autoencoder or a diffusion model.

4. The decoupled adversarial sample generation method according to claim 1, characterized in that: The loss function includes an adversarial loss function, a decoupling loss function and an imperceptible loss function, and the final loss function includes the sum of the adversarial loss function, the decoupling loss function and the imperceptible loss function, wherein, The adversarial loss function includes: In the above formula, Denotes the adversarial loss function, F y (·) represents the model’s response value to the input sample at index y, max i≠y F i (x+δ) represents the maximum response value except for the index y, and c represents the attack intensity; The decoupled loss function includes: In the above formula, represents the decoupled loss function, represents the index of the maximum response value in the clean sample, Represents the index of the maximum response value in the adversarial example; Imperceptible loss functions include: In the above formula, represents the imperceptible loss function; The final loss function includes: In the above formula, represents the final loss function, and α and β represent weights.

5. The decoupled adversarial sample generation method according to claim 4, characterized in that: Constructing a plurality of loss functions so that the constructed plurality of loss functions correspond to the plurality of constraint conditions respectively, and using the weighted sum of the plurality of loss functions as the final loss function, and training and updating the generation model comprises: The generative model is trained using a two-stage training method consisting of the first and second stages; The first stage includes optimizing the generative model with adversarial loss and imperceptible loss, so that the generative model can generate adversarial samples with adversarial properties; The second stage includes training the generative model using the final loss function so that the generative model can generate adversarial samples that are both adversarially aggressive and satisfy the decoupling principle.

6. The decoupled adversarial sample generation method according to claim 1, characterized in that: It also includes using evaluation indicators to evaluate the generative model of the adversarial sample that satisfies the decoupling principle, until the evaluation result satisfies both the decoupling principle and the deception rate of the decoupled adversarial sample.

7. The decoupled adversarial sample generation method according to claim 6, characterized in that: The evaluation indicators include: Among them, FA represents the evaluation index, ACC ops It represents the decoupling principle indicator of the decoupled adversarial sample, and FR represents the deception rate of the decoupled adversarial sample.

8. The decoupled adversarial sample generation method according to claim 6, characterized in that: The decoupling principle index of the decoupled adversarial sample is calculated and determined by the following formula: Among them, i represents the i-th sample, N represents the total number of test samples, represents a symbolic function, if If the condition is true, then output 1, otherwise output 0, x i represents the clean sample obtained by decoupling the adversarial sample, F(x i ) represents the logit output corresponding to the model, δ i represents the adversarial perturbation obtained by the i-th adversarial sample that satisfies the decoupling principle, F(δ i ) represents the logit output corresponding to the adversarial perturbation, represents the i-th adversarial sample generated by the generative model, Represents the logit output corresponding to the model.

9. The decoupled adversarial sample generation method according to claim 7, characterized in that: The deception rate of the decoupled adversarial sample is calculated by the following formula: In the above formula, f(x i ) represents the predicted category of the model output.

Citation Information

Cited By

  • Sharpness perception integrated antagonism test method for multi-modal large model

    CN121616912A