Grey box back door defense method of runtime AI model
By dividing the AI model into the front model and the back model, and using the variational autoencoder to perform backdoor detection and fine-tuning parameters to eliminate the backdoor, the problem of backdoor detection and elimination of the AI model when there is unknown whether the backdoor exists and the access permissions are restricted is solved, and the backdoor detection and elimination of the AI model is realized under the gray box access permission is realized, which is suitable for complex inference scenarios.
Patent Information
- Application Number
- CN202510141612.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-08
- Publication Date
- 2025-06-06
Smart Images

Figure CN120105417A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of cyberspace security, and in particular to a gray box backdoor defense method for a runtime AI model. Background Art
[0002] In today's intelligent society, the security of AI models is crucial to protecting model users and the safety of enterprise production processes. Enterprises, governments and individuals are facing huge challenges. How to protect the normal operation of models and prevent models from being attacked by backdoors during model operation is still a challenge. Especially when model service providers use third-party dishonest MLaaS platforms for model training, the model security problem is further highlighted. The traditional runtime backdoor input detection method detects the backdoor of input data by modifying the model input data multiple times. However, with the increasing complexity of AI models, the diversification of model deployment and the gradual strengthening of user privacy protection awareness, this way of modifying user data will inevitably seriously affect the operating performance of the model and be restricted by privacy permissions. In addition, after detecting the existence of a backdoor in the model, most current backdoor elimination schemes need to obtain the structure, parameters, gradients and other information required for model training of the entire model. When the model owner does not fully disclose this information, most methods will be difficult to apply. In order to solve the problem of backdoor input detection, some scholars have proposed that the backdoor function of a given backdoor model can be extracted into a backdoor expert model, and the backdoor input detection of the model at runtime can be implemented based on the backdoor expert model, but this scheme needs to assume that the model is known to have a backdoor. In order to mitigate the backdoor in the model, some scholars have proposed that a new layer can be inserted into the backdoor model and only specific fine-tuning training can be performed on this layer, so that this layer can filter the information of the backdoor trigger to mitigate the model backdoor attack. However, this solution only mitigates the backdoor attack and does not eliminate the backdoor in the original model. Therefore, in order to solve the problem of whether the model is unknown or not, and when the model access rights are limited, it is urgent to propose a method that can realize runtime backdoor detection and backdoor elimination to ensure the security of the model runtime.
[0003] An existing method for model detection and backdoor removal was proposed by Sun et al. [1]. Specifically, the scheme proposed semantic backdoor detection and mitigation. The key idea is to perform lightweight causal analysis, identify potential semantic backdoors based on the contribution of hidden neurons to prediction, and detect and eliminate backdoors by optimizing and adjusting the contribution of responsible neurons to correct prediction.
[0004] However, although the existing backdoor security protection scheme can detect the input data at runtime and delete the backdoor in the model, and achieve the security goal to a certain extent, it still has some shortcomings:
[0005] 1) When detecting backdoor input, the existing technology does not take into account the scenario of model split deployment and split reasoning, and the changes in defender capabilities caused by this scenario;
[0006] 2) When backdoor defense is performed during the model operation phase, the existing technology focuses on backdoor elimination under the condition of complete white-box access to the model, but does not achieve backdoor elimination of the runtime model under gray-box access rights;
[0007] 3) When the model structure, parameters, and gradient information are restricted by intellectual property rights, access rights, or MLaaS platforms, current backdoor elimination solutions are difficult to implement;
[0008] 4) The existing technology has not yet achieved a more comprehensive defense system, especially it has not yet combined the backdoor detection for the gray box model input data with the backdoor elimination for the gray box model. Summary of the invention
[0009] In order to solve the above problems existing in the prior art, the present invention provides a gray box backdoor defense method for a runtime AI model, which specifically includes:
[0010] In a first aspect, the present invention provides a gray box backdoor defense method for a runtime AI model, comprising:
[0011] The position of the intermediate feature of the AI model in operation is obtained, and the AI model is divided into a front model and a back model with the obtained position as the tangent point, the access permission of the front model is black box, and the access permission of the back model is white box;
[0012] After any sample is input into the AI model, the inference result corresponding to the input sample is obtained through inference by the AI model; the inference label and reconstruction distance extreme value corresponding to the input sample are obtained based on the trained variational autoencoder and the subsequent model; and according to the inference result, inference label, reconstruction distance extreme value corresponding to the input sample and the reconstruction distance threshold of the normal sample determined by the trained variational autoencoder, it is determined whether the input sample is a backdoor sample; the trained variational autoencoder is obtained by training with trusted samples, and the backdoor sample is a sample with a backdoor trigger;
[0013] If a backdoor sample is detected in the sample input to the AI model, the subsequent model is copied to obtain the model to be optimized, and the parameters of the subsequent model are used to initialize the parameters of the model to be optimized;
[0014] Obtain a trusted sample set, and for any first trusted sample in the trusted sample set: obtain the subsequent intermediate features corresponding to the first trusted sample through the subsequent model, obtain the new features corresponding to the first trusted sample through the model to be optimized, determine the distance loss corresponding to the first trusted sample according to the subsequent intermediate features and the new features corresponding to the first trusted sample, determine the task loss corresponding to the first trusted sample according to the task performed by the AI model, and fine-tune the parameters of the model to be optimized according to the distance loss and task loss corresponding to the first trusted sample;
[0015] The model to be optimized that has been fine-tuned with all samples in the trusted sample set is used to replace the subsequent model in the AI model to obtain an AI model that eliminates the backdoor.
[0016] In a second aspect, the present invention further provides an electronic device, comprising a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other via the communication bus;
[0017] Memory, used to store computer programs;
[0018] The processor is used to implement any method provided in the first aspect when executing a program stored in the memory.
[0019] Beneficial effects of the present invention:
[0020] The gray-box backdoor defense method of the runtime AI model provided by the present invention is to obtain the position of the intermediate feature of the AI model in operation, determine the obtained position as the tangent point, and divide the AI model into a front model and a back model, the access right of the front model is black box, and the access right of the back model is white box; after any sample is input into the AI model, the inference result corresponding to the input sample is obtained through the AI model inference, and the inference label and the reconstruction distance extreme value corresponding to the input sample are obtained based on the trained variational autoencoder and the back model, and determine whether the input sample is a backdoor sample according to the inference result, inference label, reconstruction distance extreme value and the reconstruction distance threshold of the normal sample determined by the trained variational autoencoder corresponding to the input sample, the trained variational autoencoder is obtained by training with trusted samples, and the backdoor sample is a sample with a backdoor trigger; if it is detected that there is a backdoor sample in the sample input into the AI model, the back model is copied to obtain the model to be optimized, and the parameters of the model to be optimized are initialized with the parameters of the back model; a trusted sample set is obtained, for any first trusted sample in the trusted sample set: by obtaining the back intermediate corresponding to the first trusted sample in the back model Features, obtain new features corresponding to the first trusted sample through the model to be optimized, determine the distance loss corresponding to the first trusted sample according to the subsequent intermediate features and the new features corresponding to the first trusted sample, determine the task loss corresponding to the first trusted sample according to the tasks performed by the AI model, and fine-tune the parameters of the model to be optimized according to the distance loss and task loss corresponding to the first trusted sample; use the model to be optimized that has been fine-tuned by all samples in the trusted sample set to replace the subsequent model in the AI model to obtain an AI model that eliminates the backdoor, thereby linking the backdoor detection of the input sample with the backdoor elimination of the model, and realizing the backdoor detection and elimination in the case of gray box access to the AI model; based on this method, the service provider does not need to contact the user data during the model operation phase and only needs to fine-tune part of the model parameters to realize the backdoor detection of the user input data and eliminate the backdoor of the model; compared with the existing solution of full white box access to the entire model M, the detection and elimination of the access to the previous model in the present invention is completely black box, which reduces the access rights of the backdoor defender to the model, and is more suitable for the end-cloud joint reasoning scenario and MLaaS segmentation training scenario in actual applications.
[0021] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 A schematic diagram of an application scenario provided by the present invention;
[0023] Figure 2 A flowchart of a gray box backdoor defense method for a runtime AI model provided by the present invention;
[0024] Figure 3 A schematic diagram of the architecture of an AI model provided by the present invention;
[0025] Figure 4 A schematic diagram of a backdoor consistency principle provided by the present invention;
[0026] Figure 5 A schematic diagram of the backdoor detection effect of input samples under different attack modes provided by the present invention;
[0027] Figure 6 This is a schematic diagram of the model backdoor elimination effect under different attack modes provided by the present invention. DETAILED DESCRIPTION
[0028] The present invention is further described in detail below with reference to specific embodiments, but the embodiments of the present invention are not limited thereto.
[0029] The method provided by the present invention can be applied to Figure 1 In the system shown, Figure 1 In the system shown, the service provider designs the model, trains the model, and publishes the model as a service. The model user uses this service. However, the service provider lacks the data and computing power to train the model, and adopts a third-party MLaaS to jointly implement model training and reasoning. The MLaaS platform can provide data sets and computing power for model training, and provide model deployment and online reasoning. However, backdoor attackers may implant backdoors into the model by poisoning the public data set, or MLaaS and attackers may collude to implant backdoors into the model. The defender is the role of the service provider at runtime, which is used to monitor the state of the model and perform backdoor input detection and backdoor elimination. In addition, for some specific tasks, such as segmentation learning, cascade learning, and end-cloud joint reasoning, when the model is released, some models may be deployed to the non-server side. This makes the model service provider's access rights to the part of the model reduced to black box during the model operation stage, the access rights to other models still deployed on the service provider remain white box, and the access rights to the overall AI model are reduced to gray box. Even when the model deployed to other end devices is the front part of the AI model, the defender loses the right to modify the input data. Defenders need to use clean datasets for runtime backdoor defense under this gray-box access.
[0030] In order to solve the problems existing in the prior art, the present invention provides a gray box backdoor defense method for a runtime AI model, such as Figure 2 As shown, the method includes:
[0031] S201. Obtain the position of the intermediate feature of the AI model in operation, determine the obtained position as a tangent point, and divide the AI model into a front model and a back model.
[0032] Among them, the access rights of the former model are black box, and the access rights of the latter model are white box. It can be seen that the access rights to some contents of the AI model are black box, and the access rights to some contents are white box. Therefore, the access rights to the AI model are gray box.
[0033] like Figure 3 As shown, the cut point is a mark, and the defender only has black box access to the AI model before the cut point, and maintains white box access to the AI model after the cut point.
[0034] For example, for the AI model M to be detected, we call the location where the intermediate features can be obtained the cut point, and formally divide the model into two parts through the cut point: the model before the cut point is denoted as M 1 , the model after the cut point is represented as the main M 2 When using the MLaaS platform for segmentation training, M 1 Often deployed in MLaaS platforms, M 2 Often deployed in service provider equipment.
[0035] S202. After any sample is input into the AI model, the AI model is used to infer the inference result corresponding to the input sample. The inference label and reconstruction distance extreme value corresponding to the input sample are obtained based on the trained variational autoencoder and the subsequent model. According to the inference result, inference label, reconstruction distance extreme value corresponding to the input sample and the reconstruction distance threshold of the normal sample determined by the trained variational autoencoder, it is determined whether the input sample is a backdoor sample. The trained variational autoencoder is obtained by training with trusted samples, and the backdoor sample is a sample with a backdoor trigger.
[0036] like Figure 4 As shown in the figure, the backdoor implanted in the AI model has consistency corresponding to the backdoor trigger on the input, and this consistency is called backdoor consistency. That is to say, in two AI models implanted with unrelated backdoors, two unrelated backdoor triggers cannot activate each other's backdoors. For example, for the backdoor model subjected to attack method 1, the output extracted from the backdoor input of attack method 2 is a benign output that approaches normal, while the output extracted from the backdoor input of attack method 1 is a backdoor output.
[0037] The defender inputs the clean sample into the AI model before the cut point (i.e., the previous model) to obtain the intermediate features, adjusts the shape of the intermediate features, and uses the intermediate features to train the Variational Auto-Encoder (VAE), and uses the reconstruction extreme values of the variational autoencoder and the AI model after the cut point (the subsequent model) to detect the difference in the inference results of the intermediate features reconstructed by the VAE to achieve backdoor detection. Based on the above backdoor consistency, if a backdoor is detected in the input, it can be determined that there is at least one backdoor in the model. Therefore, it is possible to determine whether there is a backdoor in the AI model by detecting whether there are backdoor samples in the samples input to the AI model. Specifically, the VAE model can be used to reconstruct the benign features based on the difference in the intermediate features between the backdoor input and the benign input, and then the backdoor input can be further detected by the difference in the target model inference results and the distribution of the VAE reconstruction distance.
[0038] Before this, it is necessary to first train a variational autoencoder. In one possible implementation, training a variational autoencoder includes: obtaining a trusted dataset And obtain multiple benign intermediate features based on the trusted data set, and obtain multiple benign three-dimensional feature maps based on the multiple benign intermediate features Based on multiple benign 3D feature maps Constructing a training set for the variational autoencoder and test set Among them, N 1 +N 2 =N; Based on the constructed variational autoencoder loss function and model parameter optimization function, through the training set The initial variational autoencoder is trained to obtain a trained variational autoencoder.
[0039] The loss function of the variational autoencoder consists of two parts: KL (Kullback-Leibler) loss and reconstruction loss. Specifically, the loss function of the variational autoencoder is expressed as:
[0040]
[0041] Among them, θ vae represents the weight parameter of the variational autoencoder, J represents the dimension of the latent variable z, j =μ i +ε j ·σ i , z j represents the value of the hidden variable z of the variational autoencoder in the jth dimension, μ j represents the value of the mean vector in the jth dimension, ε j represents the value of the normal random vector in the jth dimension, σi represents the value of the standard deviation vector in the jth dimension, where the mean vector μ and the standard deviation vector σ are the outputs of the encoder in the variational autoencoder. Represents the training set The three-dimensional feature map corresponding to any intermediate feature of represents the variational lower bound of KL loss, MSE represents the mean squared error loss, and α represents the trade-off coefficient between KL loss and MSE loss.
[0042] Furthermore, the parameters of the variational autoencoder are optimized through the model parameter optimization function. Specifically, the model parameter optimization function is expressed as:
[0043]
[0044] Among them, θ vae represents the model parameters of the variational autoencoder, represents the model parameters of the optimized variational autoencoder, N 1 represents the total number of training samples in the training set, Represents the three-dimensional feature map corresponding to the rth training sample.
[0045] Specific, trusted datasets Only a small number of samples are included.
[0046] Further, after the variational autoencoder is fully optimized to convergence, the trained variational autoencoder determines the reconstruction distance threshold of the normal sample. Specifically, the reconstruction distance threshold of the normal sample determined by the trained variational autoencoder includes: And the trained variational autoencoder, the reconstruction distance threshold of the normal sample is obtained, which is expressed as:
[0047]
[0048] Among them, τ represents the reconstruction distance threshold of normal samples, r represents the distribution of the extreme value of the reconstruction distance of normal samples, and p is the preset signal threshold. Represents interval statistical probability.
[0049] It can be seen that the reconstruction distance threshold τ of normal samples depends on the preset trust probability level p and satisfies:
[0050] Further, optionally, determining the intermediate features and the three-dimensional feature map corresponding to the input sample includes: after any input sample is inferred by the previous model and the subsequent model, the intermediate features and the inference results of the AI model are obtained, which are expressed as:
[0051] m=M 1 (x),
[0052]
[0053] Among them, m represents the intermediate features output by the previous model, M 1 Corresponding to the reasoning process of the previous model, x represents the input sample, Represents the inference result of the AI model, M 2 Corresponding to the reasoning process of the subsequent model.
[0054] Reshape the intermediate result into a three-dimensional feature map, expressed as:
[0055]
[0056] Among them, reshape(·) represents the operation of reshaping the intermediate features output by the previous model into a three-dimensional feature map.
[0057] For example, m is reshaped into a three-dimensional feature map The size is similar to 1×w×h, where w and h represent the width and height of the variational autoencoder input, and the variational autoencoder input and output channels have only one.
[0058] Optionally, the inference label and reconstruction distance extreme value corresponding to the input sample are obtained based on the trained variational autoencoder and the post-model, expressed as:
[0059]
[0060] Among them, M 1 (x ⊿ ) represents the input sample x ⊿ The intermediate feature, L 0 Represents the input sample x ⊿ The inference result obtained by AI model inference. dereshape(·) is the inverse operation of reshape(·). Represents the input sample x ⊿ The corresponding three-dimensional feature map, VAE strain classification autoencoder, L 1 Represents the input sample x ⊿ The corresponding inference label, r ⊿ Represents the input sample x ⊿ The corresponding reconstruction distance extreme values, |·| means taking the absolute value of all values, Max(·) and Min(·) mean taking the maximum and minimum values of all values.
[0061] Optionally, determining whether the input sample is a backdoor sample based on the inference result, inference label, reconstruction distance extreme value corresponding to the input sample and the reconstruction distance threshold of the normal sample determined by the trained variational autoencoder, including: based on backdoor consistency, if the inference result corresponding to the input sample is different from the inference label corresponding to the input sample, and the reconstruction distance extreme value corresponding to the input sample is greater than the reconstruction distance threshold of the normal sample, then determining that the input sample is a backdoor sample.
[0062] Specifically, whether the input sample is a backdoor sample can be determined based on L 0 ≠L 1 &r ⊿ >τ condition. Based on backdoor consistency, when a backdoor is detected in the input, it is considered that there is a backdoor in the AI model and it needs to be eliminated.
[0063] The method provided by the present invention can realize backdoor input detection at runtime. It can realize backdoor detection of input samples under the authority of full gray box access to the model by reconstructing the intermediate features of the input data through a variational autoencoder and analyzing the differences in the model inference results. In the model running stage, by analyzing the features corresponding to each input sample, backdoor input detection can be realized in one inference, without performing multiple inferences on the input.
[0064] S203. If a backdoor sample is detected in the sample input to the AI model, the subsequent model is copied to obtain the model to be optimized, and the parameters of the model to be optimized are initialized using the parameters of the subsequent model.
[0065] Exemplarily, before backdoor elimination, copy the back model M 2 Get a new model to be optimized And use the following model M 2 Model parameters To initialize the model to be optimized Parameters of the model to be optimized
[0066] S204. Obtain a trusted sample set, and for any first trusted sample in the trusted sample set: obtain the subsequent intermediate features corresponding to the first trusted sample through the subsequent model, obtain the new features corresponding to the first trusted sample through the model to be optimized, determine the distance loss corresponding to the first trusted sample according to the subsequent intermediate features and the new features corresponding to the first trusted sample, determine the task loss corresponding to the first trusted sample according to the task performed by the AI model, and fine-tune the parameters of the model to be optimized according to the distance loss and task loss corresponding to the first trusted sample.
[0067] In the present invention, the features inferred by the previous model in the AI model are intermediate features (i.e., features output from the tangent point of the AI model), and the features inferred by the subsequent model are subsequent intermediate features.
[0068] Specifically, the trusted sample set may be a trusted sample set used to train a variational autoencoder.
[0069] Optionally, the distance loss corresponding to the first trusted sample is expressed as:
[0070]
[0071] Among them, m b represents the subsequent intermediate feature corresponding to the first trustworthy sample x inferred by the subsequent model, m new represents the new feature corresponding to the first trustworthy sample x inferred by the model to be optimized, It is used to represent the distance loss calculated between the subsequent model and the model to be optimized based on the first trusted sample x. represents the model parameters of the model to be optimized, x represents the first trustworthy sample, Represents the model parameters of the subsequent model.
[0072] Optionally, the task loss corresponding to the first trusted sample is expressed as:
[0073]
[0074] in, represents the task loss of the model, and its specific representation is determined by the loss function L(·) of the model’s own task. y represents the label of the first trusted sample x, which is used to identify the first trusted sample x as a trusted sample.
[0075] Optionally, fine-tuning the parameters of the model to be optimized according to the distance loss and the task loss corresponding to the first trusted sample includes: fine-tuning the parameters of the model to be optimized according to the distance loss corresponding to the first trusted sample, the task loss corresponding to the first trusted sample, and a pre-constructed overall loss function, wherein the overall loss function is expressed as:
[0076]
[0077] Among them, β represents the trade-off coefficient between task loss and distance loss.
[0078] S205. Use the model to be optimized that has been fine-tuned with all samples in the trusted sample set to replace the subsequent model in the AI model, so as to obtain an AI model with the backdoor eliminated.
[0079] The gray-box backdoor defense method of the runtime AI model provided by the present invention is to obtain the position of the intermediate feature of the AI model in operation, determine the obtained position as the tangent point, and divide the AI model into a front model and a back model, the access right of the front model is black box, and the access right of the back model is white box; after any sample is input into the AI model, the inference result corresponding to the input sample is obtained through the AI model inference, and the inference label and the reconstruction distance extreme value corresponding to the input sample are obtained based on the trained variational autoencoder and the back model, and determine whether the input sample is a backdoor sample according to the inference result, inference label, reconstruction distance extreme value and the reconstruction distance threshold of the normal sample determined by the trained variational autoencoder corresponding to the input sample, the trained variational autoencoder is obtained by training with trusted samples, and the backdoor sample is a sample with a backdoor trigger; if it is detected that there is a backdoor sample in the sample input into the AI model, the back model is copied to obtain the model to be optimized, and the parameters of the model to be optimized are initialized with the parameters of the back model; a trusted sample set is obtained, for any first trusted sample in the trusted sample set: by obtaining the back intermediate corresponding to the first trusted sample in the back model Features, obtain new features corresponding to the first trusted sample through the model to be optimized, determine the distance loss corresponding to the first trusted sample according to the subsequent intermediate features and the new features corresponding to the first trusted sample, determine the task loss corresponding to the first trusted sample according to the tasks performed by the AI model, and fine-tune the parameters of the model to be optimized according to the distance loss and task loss corresponding to the first trusted sample; use the model to be optimized that has been fine-tuned by all samples in the trusted sample set to replace the subsequent model in the AI model to obtain an AI model that eliminates the backdoor, thereby linking the backdoor detection of the input sample with the backdoor elimination of the model, and realizing the backdoor detection and elimination in the case of gray box access to the AI model; based on this method, the service provider does not need to contact the user data during the model operation phase and only needs to fine-tune part of the model parameters to realize the backdoor detection of the user input data and eliminate the backdoor of the model; compared with the existing solution of full white box access to the entire model M, the detection and elimination of the access to the previous model in the present invention is completely black box, which reduces the access rights of the backdoor defender to the model, and is more suitable for the end-cloud joint reasoning scenario and MLaaS segmentation training scenario in actual applications.
[0080] The present invention also provides a structure of an electronic device, including a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other via the communication bus.
[0081] Memory, used to store computer programs;
[0082] The processor is used to implement the steps provided in the above method embodiment when executing the program stored in the memory.
[0083] The communication interface is used for communication between the above electronic device and other devices.
[0084] The method provided in the embodiment of the present invention can be applied to electronic devices. Specifically, the electronic device can be: a desktop computer, a portable computer, an intelligent mobile terminal, a server, etc. This is not limited here, and any electronic device that can implement the present invention belongs to the protection scope of the present invention.
[0085] As for the electronic device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the specific contents and beneficial effects and other related matters can be referred to the partial description of the method embodiment.
[0086] In order to further prove the beneficial effects of the present invention, the present invention also provides a set of experimental data, as follows:
[0087] Using the ResNet34 model and the Tiny-ImageNet dataset, this paper compares the detection effect of backdoor samples with different attack methods, where the intermediate features of the model come from the output of the third residual block of ResNet34. Figure 5 As shown, the present invention can achieve Area Under the Receiver Operating Characteristic (AUROC) scores of 98.6%, 97.9%, 97.8% and 92.7% under BadNets, AdvDoor, Blend and WaNet attacks respectively. The present invention can detect input backdoors of different attack modes very well, and can distinguish backdoor samples from normal samples very well.
[0088] Using the ResNet34 model and the Tiny-ImageNet dataset, this paper compares the model accuracy (ACC) and the changes in the backdoor attack accuracy (ASR) under different attack methods during the backdoor elimination process. Figure 6 As shown in the figure, in the backdoor model under BadNets, AdvDoor, Blend and WaNet attacks, the present invention can completely eliminate the backdoor in the model and ensure the accuracy of the model itself. Specifically, through the backdoor elimination solution of the present invention, the ASR under the four backdoor attack methods is reduced by more than 97% on average and the accuracy of the model itself is only reduced by 2.8% on average. Figure 6 shown.
[0089] The terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the features. In the description of the present invention, the meaning of "plurality" is two or more, unless otherwise clearly and specifically defined.
[0090] The above contents are further detailed descriptions of the present invention in combination with specific preferred embodiments, and it cannot be determined that the specific implementation of the present invention is limited to these descriptions. For ordinary technicians in the technical field to which the present invention belongs, several simple deductions or substitutions can be made without departing from the concept of the present invention, which should be regarded as falling within the protection scope of the present invention.
Claims
1. A gray box backdoor defense method for a runtime AI model, characterized in that: include: Obtaining the position of an intermediate feature of the AI model in operation, and dividing the AI model into a front model and a back model with the obtained position as a tangent point, wherein the access permission of the front model is a black box, and the access permission of the back model is a white box; After any sample is input into the AI model, the AI model is used to infer the inference result corresponding to the input sample; based on the trained variational autoencoder and the subsequent model, the inference label and reconstruction distance extreme value corresponding to the input sample are obtained; and according to the inference result, inference label, reconstruction distance extreme value corresponding to the input sample and the reconstruction distance threshold of the normal sample determined by the trained variational autoencoder, it is determined whether the input sample is a backdoor sample, wherein the trained variational autoencoder is obtained by training with trusted samples, and the backdoor sample is a sample with a backdoor trigger; If it is detected that there is a backdoor sample in the sample input to the AI model, the subsequent model is copied to obtain the model to be optimized, and the parameters of the model to be optimized are initialized using the parameters of the subsequent model; Obtain a trusted sample set, and for any first trusted sample in the trusted sample set: obtain a subsequent intermediate feature corresponding to the first trusted sample through the subsequent model, obtain a new feature corresponding to the first trusted sample through the model to be optimized, determine a distance loss corresponding to the first trusted sample according to the subsequent intermediate feature and the new feature corresponding to the first trusted sample, determine a task loss corresponding to the first trusted sample according to the task performed by the AI model, and fine-tune the parameters of the model to be optimized according to the distance loss and task loss corresponding to the first trusted sample; The model to be optimized that has been fine-tuned by all samples in the trusted sample set is used to replace the subsequent model in the AI model to obtain an AI model that eliminates the backdoor.
2. The method according to claim 1, characterized in that Training the variational autoencoder comprises: Obtaining a trusted dataset After the samples in the trusted data set are inferred by the previous model, a plurality of benign intermediate features are obtained; A plurality of benign three-dimensional feature maps are obtained according to the plurality of benign intermediate features. According to the plurality of benign three-dimensional feature maps Constructing a training set for the variational autoencoder and test set Among them, N1+N2=N; Based on the constructed variational autoencoder loss function and model parameter optimization function, through the training set The initial variational autoencoder is trained to obtain a trained variational autoencoder, and the loss function of the variational autoencoder is expressed as: Among them, θ vae represents the weight parameter of the variational autoencoder, J represents the dimension of the latent variable z, j =μ i +ε j ·σ i , z j represents the value of the hidden variable z of the variational autoencoder in the jth dimension, μ j represents the value of the mean vector in the jth dimension, ε j represents the value of the normal random vector in the jth dimension, σ i represents the value of the standard deviation vector in the jth dimension, where the mean vector μ and the standard deviation vector σ are the outputs of the encoder in the variational autoencoder. Represents the training set The three-dimensional feature map corresponding to any intermediate feature of represents the variational lower bound of KL loss, MSE represents the mean square error loss, and α represents the trade-off coefficient between KL loss and MSE loss; The model parameter optimization function is expressed as: Among them, θ vae represents the model parameters of the variational autoencoder, represents the model parameters of the optimized variational autoencoder, N1 represents the total number of training samples in the training set, Represents the three-dimensional feature map corresponding to the rth training sample.
3. The method according to claim 2, characterized in that The reconstruction distance threshold of the normal sample determined by the trained variational autoencoder includes: According to the test set And the trained variational autoencoder, the reconstruction distance threshold of the normal sample is obtained, which is expressed as: Among them, τ represents the reconstruction distance threshold of normal samples, r represents the distribution of the extreme value of the reconstruction distance of normal samples, and p is the preset signal threshold. Represents interval statistical probability.
4. The method according to claim 3, characterized in that Determine the intermediate features and three-dimensional feature maps corresponding to the input samples, including: After any input sample is inferred by the previous model and the subsequent model, the intermediate features and the inference result of the AI model are obtained, which can be expressed as: m=M1(x), Among them, m represents the intermediate features output by the previous model, M1 corresponds to the reasoning process of the previous model, x represents the input sample, It represents the inference result of the AI model, and M2 corresponds to the inference process of the subsequent model; The intermediate result is reshaped into a three-dimensional feature map, represented as: Among them, reshape(·) represents the operation of reshaping the intermediate features output by the previous model into a three-dimensional feature map.
5. The method according to claim 4, characterized in that Based on the trained variational autoencoder and the subsequent model, the inference label and reconstruction distance extreme value corresponding to the input sample are obtained, which are expressed as: Among them, M1(x ⊿ ) represents the input sample x ⊿ The intermediate feature of L0 represents the input sample x ⊿ The inference result obtained by AI model inference, dereshape(·) is the inverse operation of reshape(·), m ⊿ Represents the input sample x ⊿ The corresponding three-dimensional feature map, VAE corresponds to the strain classification self-encoder, L1 represents the input sample x ⊿ The corresponding inference label, r ⊿ Represents the input sample x ⊿ The corresponding reconstruction distance extreme values, |·| means taking the absolute value of all values, Max(·) and Min(·) mean taking the maximum and minimum values of all values.
6. The method according to claim 5, characterized in that The determining whether the input sample is a backdoor sample according to the inference result, inference label, reconstruction distance extreme value corresponding to the input sample and the reconstruction distance threshold of the normal sample determined by the trained variational autoencoder includes: Based on backdoor consistency, if the inference result corresponding to the input sample is different from the inference label corresponding to the input sample, and the reconstruction distance extreme value corresponding to the input sample is greater than the reconstruction distance threshold of the normal sample, the input sample is determined to be a backdoor sample.
7. The method according to claim 6, characterized in that The distance loss corresponding to the first trusted sample is expressed as: Among them, m b represents the subsequent intermediate feature corresponding to the first trustworthy sample x inferred by the subsequent model, m new represents the new feature corresponding to the first trustworthy sample x inferred by the model to be optimized, It is used to represent the distance loss calculated between the subsequent model and the model to be optimized based on the first trusted sample x. represents the model parameters of the model to be optimized, x represents the first trustworthy sample, Represents the model parameters of the subsequent model.
8. The method according to claim 7, characterized in that The task loss corresponding to the first trusted sample is expressed as: in, represents the task loss of the model, and y represents the label of the first trusted sample x, which is used to identify the first trusted sample x as a trusted sample.
9. The method according to claim 7, characterized in that: The step of fine-tuning the parameters of the model to be optimized according to the distance loss and the task loss corresponding to the first trusted sample includes: According to the distance loss corresponding to the first trusted sample, the task loss corresponding to the first trusted sample and the pre-constructed overall loss function, the parameters of the model to be optimized are fine-tuned, and the overall loss function is expressed as: Among them, β represents the trade-off coefficient between task loss and distance loss.
10. An electronic device, characterized in that: It includes a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other through the communication bus; Memory, used to store computer programs; A processor, for implementing any of the methods described in claims 1-9 when executing a program stored in a memory.
Citation Information
Cited By
Back door trigger detection method and system based on disturbance separation
CN122333462A