Training method and device for adversarial disturbance generation model

By pre-training and fine-tuning the generation model, and combining the loss function to adjust the parameters, the inefficiency problem of anti-perturbation generation under the black box target model is solved, and efficient anti-attack effect is achieved.

CN120279583APending Publication Date: 2025-07-08ANT BLOCKCHAIN TECHNOLOGY (SHANGHAI) CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510248206.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-03
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The prior art is difficult to efficiently generate adversarial perturbations to carry out effective adversarial attacks when the target model is a black box, resulting in slow and expensive attacks.

Method used

The generative model is pre-trained using the known first predictive model and fine-tuning the generative model through a small number of queries of the target model, combining two loss functions to adjust the pending parameters to generate an adversarial perturbation.

Benefits of technology

It improves the effectiveness and efficiency of adversarial attacks, and can quickly generate identification results for counter-perturbation to mislead the target model, while minimizing observable changes to target information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279583A_ABST
    Figure CN120279583A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a training method and a training device for an anti-disturbance generation model, which can be used for resisting an attack by adding disturbance to a picture in order to resist the recognition attack of an attacker on the picture by using an image recognition model. The disturbance noise added to the picture is predicted through a generative model, and the target of adding disturbance comprises the following steps: enabling an image recognition model of an attacker to output a wrong prediction result; and observers are difficult to find. Therefore, an image recognition model of an attacker is regarded as a confrontation target model, a known prediction model is used for replacing the target model to assist in pre-training a generation model, then the target model is used for rapid fine tuning of the generation model, and the obtained generation model can rapidly generate disturbance noise for any picture for confrontation attack. Therefore, the effectiveness of attack resistance can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] One or more embodiments of this specification relate to the field of image recognition technology, and in particular, to a method and apparatus for training an adversarial perturbation generation model. Background Art

[0002] Image recognition is an important field of artificial intelligence. It usually uses a computer to process, analyze, and understand images to identify various different patterns of targets and objects. Among them, in image recognition technologies such as face recognition and fingerprint recognition, privacy information such as face images and fingerprint images is involved. In recent years, privacy information recognition technologies such as faces and fingerprints have been widely applied to various fields, which also poses a certain risk of privacy leakage. For example: A malicious attacker illegally invades the image acquisition systems of privacy places such as hospitals and community access controls, obtains relevant face images, and uses an image recognition model to identify the information of patients and residents from them, and so on. Here, the attacker is equivalent to using an image recognition model to attack the image, and the countermeasure against this attack is an adversarial attack. The idea of an adversarial attack is usually to add noise to the image so that the attacker's image recognition model outputs an incorrect prediction result. Since the attacker's image recognition model is like a black box, therefore, how to effectively generate the noise information for adversarial attacks is an important technical problem in adversarial attacks. Summary of the Invention

[0003] One or more embodiments of this specification describe a method and apparatus for training an adversarial perturbation generation model to solve one or more problems mentioned in the background art.

[0004] According to a first aspect, a method for training an adversarial perturbation generation model is provided. The generation model is used to process the following data and output a predicted perturbation: target information, correct prediction result of the target information, target prediction result, current perturbation; the predicted perturbation is an adversarial perturbation for countering an identification attack of a target model against the target information, and the target model generates an attack by identifying the target information; the method includes: pre-training the generation model based on a first prediction model, where in a single iteration of pre-training, the model loss includes a first loss and a second loss. The first loss is determined by a first gap between a prediction result obtained by processing a perturbed sample target information, which is obtained by adding the predicted perturbation to the sample target information, through the first prediction model and the target prediction result. The second loss is determined by a second gap between the perturbed sample target information and the sample target information. The first prediction model is a known model having the same prediction task as the target model; fine-tuning the undetermined parameters in the generation model by querying the target model.

[0005] In one embodiment, the target information includes at least one of a picture, an audio, and a text.

[0006] In one embodiment, in a single iteration of pre-training, the model loss is determined by processing the sample target information of the current batch via the generative model; the pre-training of the generative model based on the first prediction model further includes: adjusting the undetermined parameters in the generative model with the goal of minimizing the model loss; updating the current perturbation with a prediction perturbation.

[0007] In one embodiment, the sample target information after adding the prediction perturbation is the sum of the sample target information and the prediction perturbation, and the first loss is positively correlated with the first gap; when the prediction result and the target prediction result are in numerical form, the first gap is the difference between the prediction result and the target prediction result; when the prediction result and the target prediction result are in vector form, the first gap is the distance between the prediction result and the target prediction result, and the distance is measured by at least one of cosine similarity, cross entropy, Jaccard coefficient, Euclidean distance, Chebyshev distance, Mahalanobis distance, and variance.

[0008] In one embodiment, the second loss is positively correlated with the second gap, and the second gap is the 2-norm between the tensor of the perturbed sample target information and the tensor of the sample target information, and the second loss is positively correlated with the second gap.

[0009] In one embodiment, the processing of the target prediction result and the correct prediction result of the target information by the generative model is the processing of the following fusion tensor: embedding the target prediction result and the correct prediction result of the target information to obtain the corresponding target embedding tensor and correct embedding tensor; fusing the target embedding tensor and the correct embedding tensor by one of addition, averaging, subtraction, and concatenation to obtain the fusion tensor.

[0010] In one embodiment, the adjusting of the undetermined parameters in the generative model based on the target model includes: sampling in multiple directions within a predetermined distance range of the sample target information, and querying the respective prediction results output by the target model for each sampling result; comparing each prediction result with the target prediction result respectively, and determining the prediction result closest to the target prediction result as the optimal prediction result; determining the estimated gradient direction of the prediction perturbation of the generative model according to the sampling direction corresponding to the optimal prediction result; and backpropagating the estimated gradient direction to the generative model, so as to adjust the undetermined parameters of the generative model with the goal of minimizing the model loss.

[0011] According to a second aspect, there is provided a method for adversarial attack, including: in response to detecting an identification attack of a target model on first target information, generating a first adversarial perturbation for the first target information through T iterations by using a generation model trained in the manner described in the first aspect; adding the first adversarial perturbation to the first target information to counter the identification attack of the target model.

[0012] According to a third aspect, there is provided a training device for an adversarial perturbation generation model, where the generation model is used to process the following data and output a predicted perturbation: target information, correct prediction result of the target information, target prediction result, current perturbation; the predicted perturbation is an adversarial perturbation predicted for the target information to counter the identification attack of the target model, and the target model generates an attack by identifying the target information.

[0013] The device includes:

[0014] A pre-training unit configured to pre-train the generation model based on a first prediction model. Wherein, in a single iteration round of pre-training, the model loss includes a first loss and a second loss. The first loss is determined by a first gap between a prediction result obtained by processing a sample target information after adding the predicted perturbation through the first prediction model and the target prediction result. The second loss is determined by a second gap between the perturbed sample target information after adding the predicted perturbation and the sample target information. The first prediction model is a known model having the same prediction task as the target model.

[0015] A fine-tuning unit configured to fine-tune undetermined parameters in the generation model based on the target model. During the fine-tuning process, the prediction result in the first loss is determined by the target model.

[0016] According to a fourth aspect, there is provided a computer-readable storage medium, on which a computer program is stored. When the computer program is executed on a computer, the computer is made to execute the method described in the first aspect or the second aspect.

[0017] According to a fifth aspect, there is provided a computing device, including a memory and a processor. It is characterized in that an executable code is stored in the memory, and when the processor executes the executable code, the method described in the first aspect or the second aspect is implemented.

[0018] Through the methods and devices provided in the embodiments of this specification, when the target model used by an attacker to conduct an identification attack on target information is a black-box model, in order to train a generative model that can effectively generate adversarial perturbations, first, a known first prediction model is used to pre-train the generative model, and then the generative model is fine-tuned through a small number of queries to the target model. During the training process, two aspects of restrictions are imposed on the prediction noise output by the generative model: as much as possible, no observable changes are made to the target information; as much as possible, the target model is made to mispredict towards the target prediction result. In this way, the effectiveness of adversarial attacks can be improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.

[0020] Figure 1 Shows a schematic diagram of the architecture of a specific implementation scenario of an adversarial attack in this specification;

[0021] Figure 2 Shows a schematic diagram of the training process of an adversarial perturbation generation model according to an embodiment;

[0022] Figure 3 Shows a schematic diagram of a specific pre-training process when the generative model is a U-Net;

[0023] Figure 4 Shows a schematic diagram of the process of an adversarial attack according to an embodiment;

[0024] Figure 5 Shows a schematic block diagram of a training device for an adversarial perturbation generation model according to an embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0025] The technical solutions provided in this specification will be described below in conjunction with the accompanying drawings.

[0026] Figure 1 Shows a schematic diagram of the architecture of a specific adversarial attack application scenario where the target model is an image recognition model. Figure 1 Shows an example of the attack and adversarial attack architecture in the recognition scenario of a face image.

[0027] Such as Figure 1As shown, in an example of an attack on image recognition, a face image is illegally obtained by an attacker, and the corresponding image recognition model is used as the target model to predict and classify the face image, such as identifying a specific person, etc. The adversarial attack process can be to use the image with perturbations added to make the image recognition model make mistakes. For example, the holder or protector of the face image adds noise (usually noise that is difficult to identify with the naked eye) to the face image. In the case where the attacker obtains the face image with noise added, through the corresponding image recognition model, an incorrect classification result is output, such as identifying the wrong person, etc. The noise added during the adversarial attack process can also be called adversarial perturbation, perturbation noise, adversarial noise, etc.

[0028] Figure 1 The example given shows an application scenario of an identification attack that generates an identification attack by using a picture as the target information. In practice, the target information of the attack can also be audio, text (such as natural language description information), etc., which is not limited here. In fact, for identification attacks on various target information such as text, audio, and pictures, adversarial attacks can be carried out by generating adversarial perturbations.

[0029] Conventional adversarial attack methods usually rely on accessing the gradients of the target model (in this specification, it refers to the recognition model used for the identification attack, such as a classifier or Figure 1 the image recognition model in Figure 1 ), and iteratively update the adversarial perturbation (such as the noise in

[0030] Under the technical concept of this specification, the adversarial attacker can use a generative model to predict adversarial perturbations. Specifically, the generative model can process the following data: the target information x, the correct prediction result of the target information (such as the sample class label) y, the target prediction result t, and the current perturbation z, and the output data can be the predicted perturbation z0. In order to train the generative model when the target model structure and parameters are unknown, the technical concept adopted in this specification is: first, use a known first prediction model to pre-train the generative model, and then use the target model to fine-tune the generative model by estimating the gradient. During the training process of the generative model, considering the predicted perturbation z0, it is desired to achieve the following goals as much as possible: the observed difference (such as visual difference) between the target information x' after adding the perturbation and the original target information x is as small as possible; the prediction result obtained by the target model is as close as possible to the target prediction result t (the desired misclassification result, not the correct prediction result y).

[0031] That is to say, during the training process of the generative model, the model loss needs to be determined by the prediction result of the target model. Thus, during the gradient transmission process, it is difficult to directly obtain the gradient information backpropagated by the attacker's target model (which will be described in detail later). Under the technical concept of this specification, a known (existing or trained by the attacker) first prediction model can be used to replace the attacker's target model to determine the model loss and the backpropagated gradient information, and pre-train the generative model. For the generative model pre-trained using the known first prediction model, the prediction result can be obtained from the attacker's target model through a small number of queries, so as to estimate the corresponding gradient direction and fine-tune the generative model.

[0032] In this way, the generative model can be fine-tuned based on pre-training and a small number of queries to the attacker's target model, and can adapt to various images, and can quickly generate corresponding adversarial perturbations for any target information of the same type (pictures, audio, or text) that has not been trained as a sample, greatly improving the practicality and efficiency of adversarial attacks.

[0033] The following combines the Figure 2 illustrated embodiments to describe in detail the technical concept of the generative model training in this specification.

[0034] Figure 2The training process of an adversarial perturbation generation model according to an embodiment is shown. The execution entity of this process can be any computer, device, or server with certain computing capabilities. Among them, the execution entity can be a trusted device of the attacked party, that is, a device that can provide sample target information and usually does not cause privacy leakage. Among them, the generation model is used to generate noise data for adversarial attacks, that is, adversarial perturbations, against the target model being attacked. The target model generates attacks by identifying target information. For example, it can be an image recognition model, a voiceprint recognition model, a text recognition model, etc., which are used to identify who the specific person is, who the audio belongs to, whether the target object (such as cultural relics, commodities, etc.) is included in the image, etc. When the target model is a classification model, the predicted classification category can be determined according to the output result. For example, the output result is an identifier representing the predicted category (such as a person identifier), the probability on each classification category (the category with the highest probability is the predicted classification category), etc. In this specification, the target model can be a model used by the attacker, usually a black box model, that is, a model whose parameters and structure cannot be known by the adversarial attack party. It can be understood that when the target model is a white box model, the technical solution of this specification is still applicable.

[0035] As Figure 2 shown, the process of training the generation model for adversarial attacks may include the following steps: Step 201, pre-train the generation model based on the first prediction model. Among them, in a single iteration of pre-training, the model loss includes a first loss and a second loss. The first loss is determined by the first gap between the prediction result t' obtained by processing the perturbed sample target information x' after adding the prediction perturbation z0 by the first prediction model and the target prediction result t. The second loss is determined by the gap between the perturbed sample target information x' and the sample target information x. The first prediction model is a known model with the same prediction task as the target model; Step 202, fine-tune the undetermined parameters in the generation model by querying the target model.

[0036] First, in step 201, pre-train the generation model based on the first prediction model.

[0037] Among them, the first prediction model is a model used in the pre-training stage to determine the model loss and perform gradient transmission in the process of adjusting the undetermined parameters of the generation model instead of the target model. The first prediction model often has the same prediction task as the target model (such as identifying the face in an image, identifying the user corresponding to the voiceprint in the audio), and its structure is known and the model parameter values are known. The first prediction model can be a conventional target information recognition model, such as a convolutional neural network for processing pictures, a natural language processing model, etc.

[0038] The generative model is a model used to generate attack-resistant adversarial perturbations (i.e., noise), which can be a GAN (Generative Adversarial Networks), a diffusion model, a feedforward CNN (Convolutional Neural Network) model, VAEs (Variational Autoencoders), and so on. In this specification, the diffusion model is taken as an example for description. As a generative model, the diffusion model can generate data by simulating the progressive denoising process of data. The core idea of the diffusion model is that first, noise is gradually added to the data until it completely becomes random noise, and then the noise is gradually removed to restore the original data. In the application of adversarial attacks, the process of removing noise can be adapted to the task process of generating adversarial perturbations from the input target information.

[0039] To generate adversarial perturbations, the generative model can process the following data: target information x, the correct prediction result (such as a sample label) y corresponding to the target information x, the target prediction result t, and the current perturbation z. The output data of the generative model can be denoted as the predicted perturbation z0. Here, the notations such as x, y, z, t, z0, etc. are set for convenience in distinguishing and reading during the description process and do not constitute a substantial limitation on the relevant things. Among them, the target information used in the pre-training process can be denoted as sample target information, the correct prediction result y of the target information can be the label corresponding to the sample target information, and the target prediction result t can be the prediction result expected to be output by the attacker's target model, such as an incorrect recognition result (for example, the recognition result of identifying a picture of a cat as a dog). The current perturbation z is the perturbation added to the target information x in the current iteration cycle.

[0040] During the pre-training process of the generative model, the undetermined parameters in the generative model can be adjusted with the aim of reducing the model loss. Under the technical concept of this specification, the model loss can include a first loss and a second loss. Among them, the first loss is determined by the first gap between the prediction result (such as t′) obtained by processing the perturbed sample target information (such as x′ = x + z0) after adding the predicted perturbation (such as z0) to the sample target information (such as x) through the first prediction model and the target prediction result t, and the second loss is determined by the second gap between the perturbed sample target information (such as x′) and the sample target information (such as x). It should be noted that the "first" and "second" in the first loss and the second loss are for distinguishing the loss terms and do not represent ranking the losses themselves and their importance or making other substantial limitations. The "first" and "second" in the first gap and the second gap are for corresponding to the first loss and the second loss and do not constitute other substantial limitations.

[0041] Generally, during the model training process, in a single iteration, the model loss can be determined by the generative model processing the sample target information of the current batch (one or more). In iterations other than the last one, in addition to adjusting the undetermined parameters in the generative model with the goal of minimizing the model loss, the current perturbation z can also be updated with the predicted perturbation z0, so as to use a perturbation z0 closer to the training goal as the input data of the generative model for prediction in the next update iteration. In an alternative embodiment, the sample target information of a single batch can be used for multiple rounds of updates until an end condition is met. The end condition here can include, for example, model parameter convergence, the loss function approaching 0, the gradient approaching 0, and so on.

[0042] In an alternative implementation, the sample target information (such as x′) after adding the predicted perturbation is the sum of the sample target information x and the predicted perturbation z0, and the first loss can be positively correlated with the first gap. It can be understood that: when the prediction result t′ and the target prediction result t are in numerical form, the first gap can be the difference between the prediction result t′ and the target prediction result t; when the prediction result t′ and the target prediction result t are in vector form, the first gap can be the vector distance between the prediction result and the target prediction result, and the vector distance here is measured by at least one of cosine similarity, cross entropy, Jaccard coefficient, Euclidean distance, Chebyshev distance, Mahalanobis distance, and variance. Generally, the vector distance is negatively correlated with cosine similarity, cross entropy, etc., and positively correlated with Euclidean distance, variance, etc.

[0043] The second gap can be used to describe the distance by which the sample target information is offset after adding the predicted perturbation, that is, the predicted perturbation itself. For example, it is the 2-norm between the tensor of the perturbed sample target information and the tensor of the sample target information, or the 2-norm between the predicted perturbation and the reference perturbation. Here, the reference perturbation can be the minimum perturbation to the sample target information, for example, all elements are 0. When all elements of the reference perturbation are 0, the 2-norm between the tensor of the perturbed sample target information and the tensor of the sample target information, or the 2-norm between the predicted perturbation and the reference perturbation is also equal to the 2-norm of the predicted perturbation. The second loss can be positively correlated with the second gap, that is, the larger the second gap, the larger the second loss, and vice versa, the smaller the second gap, the smaller the second loss. In this way, it can be constrained that the adversarial perturbation does not cause an obvious change in the observation result of the target information.

[0044] Since the target prediction result and the correct prediction result of the target information may be in vector form or in the form of characters (numerical values or symbols, etc.), the processing of the target prediction result and the correct prediction result of the target information by the generation model can be different according to their forms. In one embodiment, if the target prediction result and the correct prediction result of the target information are in vector form, they can be fused by one of the methods of addition, averaging, subtraction, and concatenation to obtain a fused tensor as the input of the generation model, and the generation model performs related processing. In another embodiment, if the target prediction result and the correct prediction result of the target information are in the form of character identifiers (such as class identifiers), the target prediction result and the correct prediction result of the target information can be embedded (embedding, mapping discrete data to a vector space of a predetermined dimension) to obtain corresponding target embedding tensors and correct embedding tensors, and the target embedding tensors and the correct embedding tensors are fused by one of the methods of addition, averaging, subtraction, and concatenation to obtain a fused tensor as the input of the generation model for related processing. In other embodiments, the target prediction result and the correct prediction result of the target information can also be used as the input of the generation model for related processing in other ways (such as using the original character identifiers, embedding the vector form as a two-dimensional tensor, etc.), which will not be elaborated here.

[0045] To further clarify the adjustment process of the undetermined parameters of the generation model during the pre-training process, Figure 3 Taking the diffusion model with a U-Net structure as the generation model and the target information as picture information as an example, the parameter update principle of a single iteration update round in the pre-training process of the generation model is described in detail.

[0046] Figure 3 In the example of, it is assumed that the sample target information is a sample picture, which can be the recognition object of the target model, such as privacy pictures like face pictures and fingerprint pictures. The sample picture x is any picture used in the pre-training process of the generation model.

[0047] Here, the label y of the sample image x is, for example, the classification category of the sample image (such as the person to whom the face or fingerprint belongs, etc.). Since usually, the sample image x and its corresponding label y form a training sample, x and y can be jointly denoted as the sample c. It can be understood that the label y corresponds to the correct prediction result of the sample image x, and the target prediction result t can correspond to the incorrect prediction result expected to be obtained during the adversarial attack process, such as identifying the image of person A as the target - person B or non - person A, identifying the image of a puppy as the target - a kitten, and so on. The perturbations z and z0 can be noise maps with the same shape as the sample image x. For example, if the sample image x has a shape formed by 512×960 pixels, then both the perturbations z and z0 are numerical lattices of 512×960 (which can also be regarded as noise maps, etc.), and the value corresponding to a single point is the perturbation value for the corresponding pixel on the sample image x. In practice, the sample image x and the current perturbation z can be respectively input into the generation model, or they can be fused and then input into the generation model. The sum of the sample image x and the current perturbation z is the noise image after adding the current perturbation z, denoted as x z , and x z = x + z.

[0048] As Figure 3 shown, the current perturbation is denoted as z, x and z can be input into the generation model in the form of images, and t and y can be input into the generation model in the form of tensors. It can be understood that the initial tensor corresponding to the current perturbation z can be a tensor composed of predetermined values (such as all elements being 0.5), or a noise tensor sampled based on a predetermined distribution (such as the standard normal distribution).

[0049] In one embodiment, if t and y are class identifiers representing classification categories (such as the number 5, person A, etc.), they can be converted into tensors through an embedding network and used as the input of the generation model. Among them, the embedding network can be implemented through any reasonable conventional embedding network layer, which will not be elaborated here.

[0050] In another embodiment, if t and y are vectors representing classification categories, for example, vectors with 1 in the dimension corresponding to the respective classification and 0 in other dimensions, then t and y can be directly used as the input of the generation model, or they can be respectively embedded for each dimension to obtain the corresponding embedding vectors, and the two - dimensional tensor composed of the embedding vectors of each dimension is used as the input of the generation model.

[0051] In an alternative implementation, the tensors corresponding to t and y (the vectors themselves or the tensors obtained by embedding) can also be fused through a predetermined fusion method, and the fusion result (such as denoted as the fusion tensor) is used as the input of the generation model. The predetermined fusion method is, for example: addition, averaging, subtraction, concatenation, and so on.

[0052] In more embodiments, t and y can also be used as the input of the generation model in other ways, which will not be elaborated here.

[0053] The generative model can predict the noise z0 by processing x, y, t, and z. If the generative model is denoted as F and the undetermined parameters in the generative model are denoted as w, then the prediction process of the generative model can be denoted as: F(x, y, t, z, w) = z0. As can be seen from the foregoing, the noise z0 can be an image with the same shape and size as the sample image x, including feature points with the same number of pixels as the sample image x, and a single feature point corresponds to feature values on one or more image channels. In order to adjust the undetermined parameters in the generative model, the generated perturbation can be supervised. In this specification, supervising the generated perturbation includes two aspects: the effectiveness of the adversarial perturbation (e.g., ensuring misclassification); the concealment of the perturbation (e.g., the perturbation is as small as possible).

[0054] On the one hand, in order to ensure the effectiveness of the adversarial perturbation, it is necessary to make the target model process the sample image x after adding the predicted noise z0, and obtain a target prediction result t that is as different as possible from the label y. That is to say, the larger the gap between the prediction result t' obtained by the target model processing the sample image x after adding the predicted noise z0 (denoted as the perturbed sample image x') and the target prediction result t, the greater the model loss. Assuming that the target model is denoted as Y, the prediction result obtained by the target model processing the sample image x after adding the predicted noise z0 can be denoted as Y(x + z0), then the gap between Y(x + z0) and the target prediction result t can be measured by the function L1 = j(x, z0). When Y(x + z0) and t are classification category identifiers (e.g., represented by numbers or characters), j(x, z0) can be positively correlated with |Y(x + z0) - t|. When Y(x + z0) and t are tensors, j(x, z0) can be negatively correlated with the similarity between Y(x + z0) and t, and the similarity between Y(x + z0) and t can be measured by a number that is positively correlated with one of cosine similarity, cross entropy, Jaccard coefficient, etc., or by a number that is negatively correlated with one of Euclidean distance, Chebyshev distance, Mahalanobis distance, variance, etc. When Y(x + z0) and t are multi-dimensional tensors, they can be flattened into vectors for related calculations.

[0055] On the other hand, in order to ensure the concealment of the perturbation, the predicted noise z0 can be made as close as possible to the predetermined perturbation z pr , when the predetermined perturbation z pr is a perturbation with all element values being 0, the predicted noise z0 being close to the predetermined perturbation z pr means that the perturbation is as small as possible. That is, by using the predetermined perturbation z pr as a partial supervision signal of the predicted noise z0, the second loss in the model loss is determined. The second loss can be related to the predicted noise z0 and the predetermined perturbation z pris positively correlated with the gap therebetween. Among them, the difference between the predicted noise z0 and the predetermined perturbation z pr The difference therebetween can be measured by the similarity or distance between the two. The similarity is proportional to at least one of the following: cosine similarity, cross entropy, Jaccard coefficient, etc. The distance is proportional to at least one of the following: Euclidean distance, Chebyshev distance, Mahalanobis distance, variance, etc. As a specific example, the second loss can be denoted as: L2 = ||z pr - z0|| 2 . The constraint L2 regularized by the 2-norm can avoid generating too large a perturbation, which is beneficial to maintaining the visual similarity between the adversarial sample after adding the perturbation noise and the original sample image, and avoiding attracting the attention of human observers.

[0056] In an alternative embodiment, the model loss of the generative model predicting the predicted noise z0 under the currently undetermined parameters may further include other loss terms, which will not be elaborated here. It can be understood that the model loss may include the fusion result of the first loss and the second loss. The model loss is at least one of the sum, weighted sum, and average of the first loss and the second loss. For example, it is the weighted sum of the first loss and the second loss, such as: L = L2 + βL1. Wherein, β is a preset hyperparameter.

[0057] After determining the model loss, the undetermined parameters of the generative model can be adjusted with the aim of reducing the model loss. Generally, gradient methods such as the gradient descent method and the Newton method can be used to adjust the undetermined parameters. These methods often require determining the gradient of the undetermined parameters. It can be understood that the first loss contains the prediction result of the target model. During the attack process of the attacker using the target model, it may be impossible to obtain the internal structure and the true prediction result of the target model, so the true gradient cannot be determined during the gradient transmission process. For example, during the process of calculating the gradient in the example given above, let z0 = F, and it is necessary to determine the part contained therein is transmitted from and cannot be directly determined.

[0058] In the pre-training stage of the generative model (i.e., the U-Net in Figure 3 ), a known first prediction model Y′ can be used to replace the target model, and is used to replace In this way, the gradient value can be normally backpropagated, so as to adjust each undetermined parameter in the generative model in the direction of reducing the model loss.

[0059] In the pre-training stage, a single update cycle can process one or more sample images simultaneously. For a single sample image, the generated noise can be iterated. During the iteration process, the current noise z can be updated with the currently predicted noise z0, and the above-mentioned process of updating the undetermined parameters of the generation model can be repeated until the predetermined conditions for pre-training are met. The predetermined conditions here can include, for example: the loss function L approaches 0 (less than a predetermined value for consecutive multiple cycles), the preset number of cycles (such as 10) is reached, the gradient of the undetermined parameters approaches 0 (less than a predetermined value for consecutive multiple cycles), the undetermined parameters converge (steadily tend to a certain value), and so on.

[0060] The pre-trained generation model can be further fine-tuned by the target model used by the attacker to adapt to the target model. Among them, when the possible directions of the target information to be attacked in resisting the attack are determined (for example, if the picture is a face picture and the attacked direction is to identify the person or user corresponding to the face), the generation model can be pre-trained for the attack direction in advance and quickly fine-tuned according to the target model when encountering an attacker. When the attacked direction of the picture cannot be determined, this pre-training process can be carried out when encountering an attacker using the target model for an attack, and no limitation is made here.

[0061] Next, in step 202, the undetermined parameters in the generation model are fine-tuned by querying the target model.

[0062] According to the description of the pre-training process of the generation model above, the model loss includes the prediction results of the target model for the sample target information with the predicted noise of the generation model added. Therefore, the current predicted noise of the generation model can be added to the sample target information, and the perturbed sample target information with the added noise (such as picture x′) can be provided to the attacker's target model to obtain the output result of the target model. However, in order to update the generation model, it is also necessary to determine the gradient (i.e., derivative) of the model loss with respect to each undetermined parameter in the generation model. The gradient of each undetermined parameter involves the part of the gradient transmitted by the target model during the transmission process. When the structure of the target model is unknown, gradient estimation methods can be used to determine it.

[0063] Gradient estimation is a method of estimating the gradient by sampling when the gradient cannot be directly calculated. Gradient estimation is mainly divided into two categories: Derivatives of Measure and Derivative of Paths. Examples of Derivatives of Measure-based derivative estimation include the Monte Carlo Gradient Estimation method (MCGE), etc., and examples of Derivative of Paths-based derivative estimation include the Path Integral Gradient Estimation method (PGE), etc. Taking the Monte Carlo Gradient Estimation method as an example, it is a sampling-based gradient estimation technique that approximates the calculation of the gradient of a function through Monte Carlo sampling.

[0064] As a specific example, sampling can be performed around the input data points, and the gradient direction can be estimated by evaluating the changes in the model output, so as to find a better gradient direction and update the parameters along the corresponding gradient direction. Here, the input data points can be the target information x. Sampling near the target information x can be understood as sampling in multiple directions within a predetermined distance range of the sample target information, that is, adding perturbations dx in multiple directions to the target information x. When adding multiple perturbations dx i (i = 0, 1, 2...) to x, multiple x + dx can be obtained i corresponding to multiple prediction results of the target model (such as denoted as Y0, Y1, Y2...). Each prediction result is compared with the target prediction result t respectively. The closer the prediction result is to t, the closer it is to the training objective of the generative model. If the prediction result closest to the target prediction result is the optimal prediction result, then the corresponding perturbation dx j can be used as the direction for adjusting the prediction perturbation z0 of the generative model, that is, the estimated gradient direction. Thus, by backpropagating the estimated gradient direction of the prediction perturbation z0 to the generative model, the gradient directions of the various undetermined parameters of the generative model can be determined and adjusted accordingly.

[0065] In this way, through the estimation of the gradient, the undetermined parameters in the generative model can be adjusted in one step. When reaching an end condition such as after a predetermined number of update rounds or detecting that the change in the undetermined parameters approaches 0 (the change in consecutive multiple cycles is less than a predetermined value, such as 0.01), the generative model can be made more adaptable to the target model, that is, the generated noise can more effectively counter the recognition attack of the target model. Usually, after a few rounds of adjustment, the corresponding generative model can be quickly trained, that is, the fine-tuning of the generative model.

[0066] The trained generative model can accept any unseen target information of the same type (if images are used during training, it accepts images; if audio is used during training, it accepts audio), and generate the perturbation noise added to the target information to counter the attack on the target model by processing the initial noise, target information, correct prediction result of the target information, and target prediction result, without additional model queries.

[0067] In a specific embodiment, the adversarial attack process provided in this specification includes:

[0068] In response to detecting an identification attack of the target model on the first target information, after T iterations using the generative model trained in the manner shown by Figure 2 generate a first adversarial perturbation for the first target information;

[0069] Add the first adversarial perturbation to the first target information to counter the identification attack of the target model.

[0070] Here, the first target information can be any target information that the target model may attack. The first adversarial perturbation can be the predicted perturbation output by the first model, such as Figure 3 z0 shown.

[0071] In a possible design, the process of generating perturbation noise for a target can be an iterative process. As Figure 4 shown, in the example where the target information is a picture, the process of generating noise through T cycles of iteration is shown. Figure 4 The generation model shown in is implemented by U-Net, and in practice, it can also be other generation models, such as other diffusion models, variational autoencoders, GANs, etc.

[0072] In Figure 4 the shown example, it is assumed that the initially added noise Z T is a noise image that satisfies the standard normal distribution N(0, 1). The picture targeted by the adversarial attack is a picture of a puppy, and the puppy picture and the correct recognition result are denoted as c, and the target prediction result is t (such as a kitten). Z T , c, and t are provided as input data to the generation model, and the generation model outputs the predicted noise Z T-1 . Then, Z T-1 is used as the current noise, and together with c and t, it is provided as input data to the generation model, and the generation model outputs the predicted noise Z T-2 . And so on, until the T-th cycle, Z1 is used as the current noise, and together with c and t, it is provided as input data to the generation model, and the generation model outputs the predicted noise Z0.

[0073] In Figure 4 the shown example, the generation of perturbation is simulated by gradually reducing the noise. By controlling the gradual reduction of the noise based on the initial noise, the quality of the generated image (as different from the original picture as little as possible) and the adversarial attack effect (effectively guiding misclassification) can be guaranteed.

[0074] Reviewing the above process, in the case where the target model used by the attacker to conduct an identification attack on the target information is a black-box model, in order to train a generation model that can effectively generate adversarial perturbations, first, a known first prediction model is used to pre-train the generation model, and then the generation model is fine-tuned through a small number of queries to the target model. During the training process, two aspects of the predicted noise output by the generation model are restricted: as little as possible to produce observable changes to the target information; as much as possible to make the target model mispredict towards the target prediction result. In this way, the effectiveness of the adversarial attack can be improved.

[0075] According to an embodiment of another aspect, there is also provided a training device for an adversarial perturbation generation model. The device can be provided in any computer, device or server with a certain computing power, for example, a trusted device of the attacked party. Among them, the generation model is used to process the following data and output a predicted perturbation: target information, correct prediction result of the target information, target prediction result, and current perturbation. Here, the predicted perturbation is an adversarial perturbation for countering the recognition attack of the target model for predicting the target information, and the target model generates an attack by recognizing the target information. The target information here can be at least one of a picture, text, and audio.

[0076] Figure 5 FIG. 500 shows a training device for an adversarial perturbation generation model according to an embodiment of the present specification. As Figure 5 shown, the training device 500 for the adversarial perturbation generation model may include:

[0077] A pre-training unit 501, configured to pre-train the generation model based on a first prediction model. Among them, in a single iteration of pre-training, the model loss includes a first loss and a second loss. The first loss is determined by the first gap between the prediction result obtained by processing the sample target information after adding the predicted perturbation through the first prediction model and the target prediction result. The second loss is determined by the second gap between the perturbed sample target information after adding the predicted perturbation and the sample target information. The first prediction model is a known model with the same prediction task as the target model;

[0078] A fine-tuning unit 502, configured to fine-tune the undetermined parameters in the generation model based on the target model. During the fine-tuning process, the prediction result in the first loss is determined by the target model.

[0079] It should be noted that Figure 5 the device 500 shown corresponds to Figure 2 the method described, Figure 2 and the corresponding descriptions in the method embodiments shown also apply to the device 500 and will not be repeated here.

[0080] According to an embodiment of another aspect, there is also provided a computer-readable storage medium, on which a computer program is stored. When the computer program is executed in a computer, it causes the computer to execute the method described in conjunction with Figure 2 etc.

[0081] According to an embodiment of still another aspect, there is also provided a computing device, including a memory and a processor. An executable code is stored in the memory. When the processor executes the executable code, the method described in conjunction with Figure 2 etc. is implemented.

[0082] Those skilled in the art should be able to realize that in one or more of the above examples, the functions described in the embodiments of this specification can be implemented by hardware, software, firmware, or any combination thereof. When implemented using software, these functions can be stored in a computer-readable medium or transmitted as one or more instructions or codes on a computer-readable medium.

[0083] The specific implementation manners described above further elaborate on the purpose, technical solutions, and beneficial effects of the technical concept of this specification. It should be understood that the above description is only the specific implementation manners of the technical concept of this specification and is not used to limit the protection scope of the technical concept of this specification. Any modifications, equivalent replacements, improvements, etc. made on the basis of the technical solutions in the embodiments of this specification should be included within the protection scope of the technical concept of this specification.

Claims

1. A training method for an adversarial perturbation generation model, where the generation model is used to process the following data and output a predicted perturbation: target information, correct prediction results of the target information, target prediction results, and current perturbation; the predicted perturbation is an adversarial perturbation for an identification attack against a target model predicted for the target information, and the target model generates an attack by identifying the target information. The method includes: Pre-training the generation model based on a first prediction model. Wherein, in a single iteration of pre-training, the model loss includes a first loss and a second loss. The first loss is determined by a first gap between a prediction result obtained by processing a perturbed sample target information, which is the sample target information added with the predicted perturbation, through the first prediction model and the target prediction result. The second loss is determined by a second gap between the perturbed sample target information and the sample target information. The first prediction model is a known model with the same prediction task as the target model. Fine-tuning the undetermined parameters in the generation model by querying the target model.

2. The method according to claim 1, wherein The target information includes at least one of pictures, audio, and text.

3. The method according to claim 1, wherein, In a single iteration of pre-training, the model loss is determined by processing the sample target information of the current batch through the generation model. The pre-training of the generation model based on the first prediction model further includes: Adjusting the undetermined parameters in the generation model with the goal of minimizing the model loss; Updating the current perturbation with the predicted perturbation.

4. The method according to claim 1, wherein, The sample target information added with the predicted perturbation is the sum of the sample target information and the predicted perturbation, and the first loss is positively correlated with the first gap. When the prediction result and the target prediction result are in numerical form, the first gap is the difference between the prediction result and the target prediction result. When the prediction result and the target prediction result are in vector form, the first gap is the distance between the prediction result and the target prediction result, and the distance is measured by at least one of cosine similarity, cross entropy, Jaccard coefficient, Euclidean distance, Chebyshev distance, Mahalanobis distance, and variance.

5. The method according to claim 1, wherein, The second gap is the 2-norm between the tensor of the perturbed sample target information and the tensor of the sample target information, and the second loss is positively correlated with the second gap.

6. The method according to claim 1, wherein, The processing of the target prediction result and the correct prediction result of the target information by the generation model is the processing of the following fusion tensor: Embedding the target prediction result and the correct prediction result of the target information to obtain corresponding target embedding tensors and correct embedding tensors; Fusing the target embedding tensor and the correct embedding tensor by one of addition, averaging, subtraction, and concatenation to obtain the fusion tensor.

7. The method according to claim 1, wherein, The fine-tuning of the undetermined parameters in the generation model based on the target model includes: Sampling in multiple directions within a predetermined distance range of the sample target information and querying the target model for each prediction result output for each sampling result; Comparing each prediction result with the target prediction result respectively, and determining the prediction result closest to the target prediction result as the optimal prediction result. Determine the estimated gradient direction of the prediction perturbation of the generation model according to the sampling direction corresponding to the optimal prediction result; Backpropagate the estimated gradient direction to the generation model, so as to adjust the undetermined parameters of the generation model with the goal of minimizing the model loss.

8. A method for adversarial attack, comprising: In response to detecting an identification attack of a target model on first target information, use the generation model trained by the method of claim 1 to generate a first adversarial perturbation for the first target information after T iterations; Add the first adversarial perturbation to the first target information to counter the identification attack of the target model.

9. A training device for an adversarial perturbation generation model, the generation model is used to process the following data and output a prediction perturbation: target information, correct prediction result of the target information, target prediction result, current perturbation; the prediction perturbation is an adversarial perturbation predicted for the target information to counter the identification attack of the target model, and the target model generates an attack by identifying the target information; The device includes: A pre-training unit configured to pre-train the generation model based on a first prediction model. Wherein, in a single iteration of pre-training, the model loss includes a first loss and a second loss. The first loss is determined by the first gap between the prediction result obtained by processing the sample target information after adding the prediction perturbation through the first prediction model and the target prediction result. The second loss is determined by the second gap between the perturbed sample target information after adding the prediction perturbation and the sample target information. The first prediction model is a known model with the same prediction task as the target model; A fine-tuning unit configured to fine-tune the undetermined parameters in the generation model based on the target model. During the fine-tuning process, the prediction result in the first loss is determined by the target model.

10. A computer-readable storage medium, on which a computer program is stored. When the computer program is executed on a computer, the computer is made to execute the method according to any one of claims 1-8.

11. A computing device, comprising a memory and a processor, characterized in that, Executable code is stored in the memory. When the processor executes the executable code, the method according to any one of claims 1-8 is implemented.