Diffusion model-oriented backdoor defense method based on trigger synthesis and backdoor parameter fine tuning
The backdoor in the diffusion model is detected by synthesizing the trigger vector and backdoor indication indicators, and the backdoor is removed through parameter fine-tuning, which solves the problem of poor backdoor detection and removal effects in the prior art, achieving high-precision backdoor defense and model generation performance maintenance.
Patent Information
- Application Number
- CN202510158309.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-13
- Publication Date
- 2025-05-13
AI Technical Summary
The existing backdoor defense technology is difficult to effectively detect and remove backdoors in diffusion models, and it is difficult to improve the backdoor detection accuracy and the effect of removing backdoors without damaging the model generation performance.
By synthesizing, the trigger vector of the backdoor in the model can be activated, and the new backdoor indication indicators can be calculated for backdoor detection; the backdoor related parameters in the model can be identified and fine-tuned, the backdoor in the model can be removed, while retaining the model's generation ability.
It improves the accuracy and reliability of backdoor detection, can effectively remove backdoors in the model, while maintaining the model generation performance, with good generalization and resource efficiency.
Smart Images

Figure CN119990251A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence security technology, and more specifically to a backdoor defense method oriented to a diffusion model based on trigger synthesis and backdoor parameter fine-tuning. Background Art
[0002] Diffusion Model is one of the cutting-edge technologies in the field of Generative Artificial Intelligence (GAI) in recent years. Compared with traditional generative models (such as Generative Adversarial Networks (GAN) and Variational Autoencoders (VAE), diffusion models have shown excellent performance in image, video, audio and other generation tasks, and have also been widely used in downstream tasks such as image restoration and image super-resolution. Diffusion models are also used to synthesize high-quality training data for tasks such as autonomous driving and medical image analysis, promoting technological progress in these fields.
[0003] However, as diffusion models are widely used in various industries, the security risks associated with them are becoming increasingly prominent, especially the backdoor attack problem. A backdoor attack refers to an attacker implanting a hidden backdoor in a model by maliciously modifying the training process or training data. In this attack, when the model receives input containing a specific trigger, the backdoor is activated, causing the model to generate malicious content preset by the attacker; while for normal input that does not contain a trigger, the model will maintain normal behavior. The threats posed by backdoor attacks on diffusion models can be mainly divided into two aspects: on the one hand, attackers can use backdoors to manipulate model outputs, leading to the spread of harmful content and causing reputational and economic losses to model owners; on the other hand, attackers can use backdoor attacks to inject bias into diffusion models, thereby affecting downstream models trained using data synthesized by the diffusion model. The training process of diffusion models usually requires a large amount of high-quality data and a large amount of computing resources. Therefore, many companies and research institutions outsource model training to third parties or directly use open source models, which gives attackers more opportunities to conduct backdoor attacks.
[0004] However, most of the existing backdoor defense techniques are not designed for diffusion models. Since the diffusion model uses noise as the prediction target and needs to use the back-diffusion process to iteratively remove the noise in the data, the defense strategies in traditional neural networks cannot be directly applied. The existing backdoor defense methods for diffusion models can be mainly divided into two categories: input-level defense and model-level defense. Input-level defense technology prevents backdoor attacks by detecting whether there are triggers in the input data. Common detection methods include analyzing the distribution difference between input data and clean input samples, evaluating the ability of input data to resist adversarial perturbations, etc. There are obvious limitations to this type of method: first, this type of method cannot effectively detect backdoors before the attacker uploads malicious inputs; second, detecting each input sample will significantly increase the computational overhead and application cost. Therefore, the input-level defense method is difficult to become an ideal defense solution. In contrast, model-level defense technology aims to directly detect and remove backdoors in the model, which is a more feasible defense strategy. Although some model-level defense technologies have proposed preliminary solutions for diffusion models, these methods still face challenges in defending against advanced backdoor attacks in practical applications. In particular, there are still many technical difficulties in effectively detecting model backdoors and removing them while maintaining the original generation capabilities of the model.
[0005] Therefore, how to improve the accuracy of the backdoor detection algorithm for the diffusion model and design an algorithm that can effectively remove the backdoor without compromising the model generation performance has become a key issue in current research and a technical challenge that technical personnel in this field urgently need to solve. Summary of the invention
[0006] In view of this, the present invention provides a backdoor defense method for diffusion models based on trigger synthesis and backdoor parameter fine-tuning. By synthesizing a trigger vector that can activate the backdoor in the model and calculating a novel backdoor indicator, backdoor detection is realized, which can effectively prevent backdoor attacks; the backdoor in the model is removed by identifying and fine-tuning the backdoor related parameters in the model, while retaining the generation ability of the model. Compared with the existing methods, the trigger synthesis method proposed in the present invention better utilizes the basic properties of backdoor attacks. The synthesized trigger vector can effectively activate the backdoor in the diffusion model, so it can effectively reveal the existence of the backdoor and has better interpretability and reliability; the proposed backdoor detection indicator can more accurately distinguish between clean models and backdoor models, thereby improving the backdoor detection accuracy; the proposed parameter importance identification method can accurately locate the backdoor related parameters, ensuring the effectiveness and stability of the subsequent fine-tuning method; the proposed fine-tuning loss can effectively remove the backdoor while retaining the generation ability of the model to the greatest extent, and has good practicality. The present invention has good generalization and is suitable for a variety of diffusion models and samplers; and the running time is relatively short and the resource consumption is low.
[0007] In order to achieve the above object, the present invention adopts the following technical solution:
[0008] Step S1: Generate a batch of clean samples that are not affected by the backdoor using the diffusion model to be detected to form a clean data set;
[0009] Step S2: Design a target loss function and iteratively optimize the synthetic trigger; use the synthetic trigger to calculate the backdoor indicator to determine whether there is a backdoor in the model;
[0010] Step S3: If a backdoor is detected, the importance scores of each parameter in the diffusion model for the clean sampling and backdoor sampling processes are calculated;
[0011] Step S4: Identify parameters related to the backdoor based on the importance score; freeze model parameters unrelated to the backdoor, initialize and fine-tune parameters related to the backdoor; after fine-tuning, the backdoor in the diffusion model is removed.
[0012] In step S1, a clean data set is generated using the diffusion model to be tested. The generation of this data set specifically includes:
[0013] Step S1.1: Determine the structure, specific parameters, sampler, input data dimension, and clean training loss of the diffusion model M0 to be tested Determine the value and range of the time step t during model training and inference, and the range is recorded as [T min ,T max ]; where the clean training loss is the training loss function of the diffusion model when it is not affected by the backdoor attack; The t in it represents the time step, and c represents that the loss is a clean loss;
[0014] Step S1.2: Use a random number generator to generate a clean input that meets the input data dimension of M0; perform a sampling process to convert the clean input into a clean sample; wherein the clean input is input data without triggers, and the clean sample is noise-free data obtained by the diffusion model using the clean input to perform a sampling process; in this step and subsequent steps, a standard normal distribution random generator is usually used to generate random vectors, and the random vectors can be translated and scaled according to different diffusion model types;
[0015] Step S1.3: Repeat step S1.2 to obtain multiple non-repetitive clean samples to form a clean data set.
[0016] In step S2, the synthetic trigger is updated by an optimization method, a backdoor sample is generated by using the synthetic trigger, and a backdoor indicator is calculated to detect the backdoor in the model. This backdoor detection method specifically includes:
[0017] Step S2.1: Generate a clean input that conforms to the input data dimension of M0, with a shape of [bs, dim1]; where bs is the number of samples used in this step and dim1 is the dimension of a single input sample of M0;
[0018] Step S2.2: Generate a random vector of shape [1, dim1] as the initial value of the synthetic trigger;
[0019] Step S2.3: Add the synthetic trigger to each clean input sample element by element to obtain the backdoor input; send the backdoor input to the diffusion model, perform the sampling process, and obtain the backdoor sample; wherein the backdoor input is the input data containing the synthetic trigger, and the backdoor sample is the noise-free data obtained by the diffusion model using the backdoor input to perform the sampling process;
[0020] Step S2.4: Use the pre-trained feature encoder to map the backdoor sample to the embedding space to obtain an embedding vector Z of shape [bs, dim2]; where dim2 is the dimension of a single sample in the embedding space; the pre-trained feature encoder (such as CLIP-VIT-B32) can effectively extract features from the sample, so that samples containing more similar concepts are closer in the embedding space;
[0021] Step S2.5: Use Z to calculate the backdoor indicator, the maximum similar cluster ratio MSCR (Max Similar Cluster Ratio):
[0022]
[0023] Among them, S(Z i ,Z j ) is used to calculate Z i and Z j The cosine similarity of the two vectors; |·| in the numerator and denominator is the number of elements in the calculation set, and |Z|=bs; κ is the similarity threshold, which is selected according to the type of M0, and the empirical range is 0.90~0.95; if the value of MSCR is greater than the preset backdoor indication threshold, it is determined that there is a backdoor in M0 and enters step S3, otherwise execute step S2.6; the empirical range of the backdoor indication threshold is 0.4~0.6;
[0024] Step S2.6: If the MSCR does not exceed the backdoor indication threshold, the similarity loss loss is calculated using the embedded vector and the backdoor sample respectively. s and entropy loss e :
[0025]
[0026] Among them, S(Z i ,Zj ) is used to calculate Z i and Z j The cosine similarity of these two vectors; |·| is the number of elements in the calculation set, and |Z| = bs; is the entropy of the i-th backdoor sample; κ + and κ - is a predefined threshold whose empirical range is κ + ∈[8.1,11.3],κ - ∈[0.1,2.8]; max(·,·) is to calculate the maximum value between two values.
[0027] Step S2.7: Define the optimization target loss o Loss s and loss e The specific formula of this objective is as follows:
[0028] min loss o =loss s +λloss e
[0029] Among them, λ is the dynamically adjusted weight coefficient, which is usually set to loss s and λloss e Have the same magnitude; calculate loss by back propagation o The gradient of the synthetic trigger is calculated and the Adam optimizer is used to update the synthetic trigger. The optimization goal is to make the loss o Minimum; The parameter requirements of the Adam optimizer are not strict. In practice, lr=1E-3, β1=0.9, β2=0.999;
[0030] Step S2.8: Repeat steps S2.3 to S2.7 until the MSCR is greater than the backdoor indication threshold or the number of iterations reaches a predefined upper limit.
[0031] In step S3, the importance scores of the parameters in the diffusion model are calculated using the Taylor expansion of the clean training loss and the backdoor training loss. The importance score calculation method specifically includes:
[0032] Step S3.1: Get a batch of clean samples from the clean dataset, whose shape is [bs, dim1]; where bs is the number of samples used in this step, and dim1 is the dimension of a single input sample of M0;
[0033] Step S3.2: Generate a clean input that conforms to the input data dimension of M0, with a shape of [bs, dim1]; add the synthetic trigger to each clean input sample element by element to obtain the backdoor input; send the backdoor input to the diffusion model, perform the sampling process, and obtain the backdoor sample;
[0034] Step S3.3: Calculate the clean training loss using clean samples and backdoor samples respectively and backdoor training loss Among them, t represents the time step, c and b represent the loss is clean training loss and backdoor training loss respectively; backdoor training loss is the specific training loss used by the attacker to implant a backdoor into the diffusion model. There are two cases for its calculation formula: when the defender has an understanding of the potential backdoor attack method, it can be modified appropriately To get formula and use the backdoor sample; if you have no knowledge of the potential backdoor attack method, you can directly use The formula only replaces the input data from clean samples to backdoor samples;
[0035] Step S3.4: Calculate the Taylor importance score of each parameter in M0 respectively; regard all parameters θ of M0 as a two-dimensional matrix, where each row vector θ i =[θ i1 ,θ i2 ,…,θ iK ] represents a parameter containing K scalars; the value range of i is determined according to the parameter value of M0; θ i The Taylor importance score I(θ i ) can be calculated by the following formula:
[0036]
[0037] Among them, |·| is to calculate the absolute value of its internal vector; θ ik is θ i is the kth scalar parameter in ; t is the time step in the training process; represent For θ ik Derivation; is the training loss at time step t, When, I(θ i ) is θ i Importance score for clean sampling process I ci ; When, I(θ i ) is θ i Importance score for the backdoor sampling process I bi ; Calculate I ci and I bi .
[0038] In step S4, the backdoor related parameters in the diffusion model are identified based on the importance score, and these parameters are reset and fine-tuned to remove the backdoor in the model while maintaining the model's generation capability. This backdoor removal method specifically includes:
[0039] Step S4.1: Calculate the differential importance score of each parameter in θ, θ i The differential importance score of di =I bi -I ci ;
[0040] Step S4.2: Select θ in |I di |The largest part of the parameters is taken as the backdoor related parameters, and the remaining parameters are taken as the backdoor irrelevant parameters; where |·| is the absolute value of the internal vector; the proportion of backdoor related parameters is selected according to the model structure and parameter quantity of m0, usually 1% to 5%;
[0041] Step S4.3: Create a backup M1 of m0, whose parameters are recorded as θ'; freeze all parameters in θ and backdoor-independent parameters in θ', and initialize backdoor-related parameters in θ'; the initialization method may use random initialization, zero initialization, etc.;
[0042] Step S4.4: Define reference loss
[0043]
[0044] in, is the clean input of the diffusion model at time step t, ∈ θ (·) and ∈ θ' (·) represent the output of models M0 and M1 respectively, t∈[T min ,T max ] is the time step in the training process, ‖·‖2 is the 2-norm of its internal vector;
[0045] Step S4.5: Define fine-tuning loss It is the clean training loss and reference loss The sum of:
[0046]
[0047] Step S4.6: Using the clean dataset and Fine-tune the backdoor related parameters in θ'; this process gradually removes the backdoor in the model while maintaining its original generative ability.
[0048] The beneficial effects brought by the technical solution provided by the present invention are:
[0049] By proposing a trigger synthesis technology based on similarity loss and entropy loss, we can synthesize triggers that can more effectively activate backdoors than existing methods, thereby improving the interpretability and reliability of backdoor detection results; the proposed backdoor detection indicators can more effectively distinguish between clean models and backdoor models than existing indicators, thereby improving the accuracy of backdoor detection results; by proposing a backdoor removal method based on importance score and fine-tuning loss function, we can completely remove the backdoor in the diffusion model while maintaining the model's generation ability. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying creative work.
[0051] Figure 1 A flow chart of a backdoor defense method for a diffusion model based on trigger synthesis and backdoor parameter fine-tuning provided by the present invention.
[0052] Figure 2 The present invention provides a backdoor defense method framework based on trigger synthesis and backdoor parameter fine-tuning.
[0053] Figure 3 Schematic diagram of sampling results of the poisoning diffusion model in an embodiment of the present invention before and after the backdoor is removed. DETAILED DESCRIPTION
[0054] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0055] like Figure 1 and Figure 2 The embodiment of the present invention discloses a backdoor defense method for a diffusion model based on trigger synthesis and backdoor parameter fine-tuning, comprising the following steps:
[0056] Step S1: Generate a batch of clean samples that are not affected by the backdoor using the diffusion model to be detected to form a clean data set.
[0057] Step S1.1: Determine that the diffusion model M0 to be detected is a DDPM model trained on the CIFAR10 dataset, and the sampler used is DDIM. Determine the model structure and parameters of M0. The single input data of M0 is random Gaussian noise with a dimension of 3*32*32, and the generated clean sample is a picture with a resolution of 3*32*32 that conforms to the distribution of the CIFAR10 dataset. The time step in the training process is t=1,2,…,T, a total of T positive integers. Clean training loss The formula is:
[0058]
[0059] in, represents the clean input of the diffusion model at time step t. ∈ represents The true value of the noise contained in is generated by a random number generator. The random vectors generated in this step and subsequent steps use a standard normal distribution random generator. ∈ θ (·) is the noise predicted by M0, and ||·||2 is the 2-norm of its intrinsic vector.
[0060] Step S1.2: Use a random number generator to generate a clean input that meets the input data dimension of M0, with a dimension of 128*3*32*32. Use the DDIM sampler to perform the sampling process and convert the clean input into 128 clean samples. Among them, the clean input is the input data without triggers, and the clean sample is the noise-free data obtained by the diffusion model using the clean input to perform the sampling process.
[0061] Step S1.3: Repeat step S1.2 to obtain 6,000 non-repetitive clean samples to form a clean data set.
[0062] Step S2: Design a target loss function and iteratively optimize the synthetic trigger; use the synthetic trigger to calculate the backdoor indicator to determine whether there is a backdoor in the model.
[0063] Step S2.1: Generate a clean input that conforms to the input data dimension of M0, which has a shape of [16, 3, 32, 32].
[0064] Step S2.2: Generate a random vector of shape [1,3,32,32] as the initial value of the synthetic trigger. Figure 2 In order to more clearly show the superposition of triggers and clean inputs, a picture of "Hello Kitty" is used as an example of a synthetic trigger.
[0065] Step S2.3: Add the synthetic trigger to each clean input sample element by element to obtain a backdoor input of shape [16, 3, 32, 32]. Send the backdoor input to the diffusion model and use the DDIM sampler to perform the sampling process to obtain the backdoor sample. The backdoor input is the input data containing the synthetic trigger, and the backdoor sample is the noise-free data obtained by the diffusion model using the backdoor input to perform the sampling process.
[0066] Step S2.4: Use the pre-trained feature encoder CLIP-VIT-B32 to map the backdoor sample to the embedding space and obtain an embedding vector Z with a shape of [16,512]. CLIP-VIT-B32 can effectively extract features from the image, so that images containing more similar concepts are closer in the embedding space.
[0067] Step S2.5: Use Z to calculate the backdoor indicator, the maximum similar cluster ratio MSCR (Max Similar Cluster Ratio):
[0068]
[0069] Among them, S(Z i ,Z j ) is used to calculate Z i and Z j The cosine similarity of the two vectors. The |·| in the numerator and denominator is the number of elements in the calculation set, and |Z|=bs. κ is the similarity threshold, which is set to 0.95. If the value of MSCR is greater than the preset backdoor indication threshold, it is determined that there is a backdoor in M0 and the process goes to step S3, otherwise, it goes to step S2.6. The backdoor indication threshold is set to 0.6.
[0070] Step S2.6: If the MSCR does not exceed the backdoor indication threshold, the similarity loss loss is calculated using the embedded vector and the backdoor sample respectively. s and entropy loss e :
[0071]
[0072] Among them, S(Z i ,Z j ) is used to calculate Z i and Z j The cosine similarity of these two vectors; |·| is the number of elements in the calculation set, and |Z| = bs; is the entropy of the i-th backdoor sample. + and κ - The value of is κ + =8.5,κ - = 0.1. max(·,·) calculates the maximum value of two values.
[0073] Step S2.7: Define the optimization target loss o loss s and loss e The specific formula of this objective is as follows:
[0074] min loss o =loss s +λloss e
[0075] Among them, λ is the dynamically adjusted weight coefficient, which is set in the code to let loss s and λloss e Have values of the same magnitude. Loss is calculated by back propagation o The gradient of the synthetic trigger is calculated and the synthetic trigger is updated using the Adam optimizer. The parameters of the Adam optimizer are lr = 1e-3, β1 = 0.9, β2 = 0.999.
[0076] Step S2.8: Repeat steps S2.3 to S2.7 until the MSCR is greater than the backdoor indication threshold or the number of iterations reaches a predefined upper limit. Figure 2 As shown, the trigger vector synthesized by this method can activate the backdoor in the diffusion model, causing the model to generate the target image set by the attacker. In this embodiment, "Mickey" is used as the target image.
[0077] Step S3: The presence of a backdoor in the model is detected, and the importance scores of each parameter in the diffusion model for the clean sampling and backdoor sampling processes are calculated.
[0078] Step S3.1: Get a batch of clean samples from the clean dataset, whose shape is [16, 3, 32, 32].
[0079] Step S3.2: Generate a clean input that conforms to the input data dimension of M0, with a shape of [16, 3, 32, 32]. Add the synthetic trigger to each clean input sample element-wise to obtain a backgate input with a shape of [16, 3, 32, 32]. Feed the backgate input into the diffusion model and perform the sampling process using the DDIM sampler to obtain a backgate sample with a shape of [16, 3, 32, 32].
[0080] Step S3.3: Calculate the clean training loss using clean samples and backdoor samples respectively and backdoor training loss The formulas are:
[0081]
[0082] in, represents a clean sample; ∈ θ (·) is the output of M0, i.e., the noise predicted by the model; ∈ represents the true value of the noise in the model input at time step t; ||·||2 is the 2-norm of its internal vector. represents the clean input to the diffusion model at time step t, represents the backdoor input of the diffusion model at time step t. g is the value of the synthetic trigger optimized in step S2. α t ,β t ,ρ t It is a sequence set according to the type of M0 and has nothing to do with the setting of the backdoor attack. t} is a monotonically decreasing sequence linearly interpolated from 1 to 0, {β t} and {ρ t} is a monotonically increasing sequence of linear interpolation from 0 to 1, with an interpolation step of
[0083] Step S3.4: Calculate the Taylor importance score of each parameter in M0. Consider all the parameters θ of M0 as a two-dimensional matrix, where each row vector θ i =[θ i1 ,θ i2 ,…,θ iK ] represents a parameter containing K scalars. θ i The Taylor importance score I(θ i ) can be calculated by the following formula:
[0084]
[0085] Among them, |·| is to calculate the absolute value of its internal vector; θ ik is θ i is the kth scalar parameter in ; t is the time step in the training process; represent For θ ik Derivative. is the training loss at time step t, When, I(θ i ) is θ i Importance score for clean sampling process I ci ; When, I(θ i ) is θ i Importance score for the backdoor sampling process I bi . Calculate I ci and I bi .
[0086] Step S4: Identify backdoor-related parameters in the diffusion model based on the importance scores. Freeze model parameters not related to the backdoor, and initialize and fine-tune parameters related to the backdoor. After fine-tuning, the backdoor in the diffusion model is removed.
[0087] Step S4.1: Calculate the differential importance score of each parameter in θ, θ i The differential importance score of di =I bi -I ci ;
[0088] Step S4.2: Select θ in |I di The largest 1% parameters are regarded as backdoor related parameters, and the remaining parameters are regarded as backdoor irrelevant parameters. Among them, |·| is the absolute value of the internal vector.
[0089] Step S4.3: Create a backup M1 of M0, whose parameters are recorded as θ'; freeze all parameters in θ and backdoor-independent parameters in θ', and initialize backdoor-related parameters in θ'; use random initialization as the initialization method.
[0090] Step S4.4: Define reference loss
[0091]
[0092] in, is the clean input of the diffusion model at time step t, ∈ θ (·) and ∈ θ' (·) represents the output of models M0 and M1, t is the time step in the training process, and ||·||2 is the 2-norm of its internal vector;
[0093] Step S4.5: Define fine-tuning loss It is the clean training loss and reference loss The sum of:
[0094]
[0095] Step S4.6: Using the clean dataset and Fine-tune the backdoor related parameters in θ'. This process will gradually remove the backdoor in the model while maintaining its original generative ability. Figure 3 As shown in (ac), the original diffusion model generates clean samples when it receives clean input, and generates backdoor samples (pictures of "Mickey") when it receives backdoor input containing real triggers or synthetic triggers. The real trigger is the trigger vector actually used by the attacker to inject the backdoor into the diffusion model, and the synthetic trigger is the trigger vector optimized in step S2 of this method. Figure 3As shown in (df), the model purified by this method cannot generate the target image when receiving backdoor input containing real triggers or synthetic triggers, while the quality of the generated image is almost the same as that of the original model when receiving clean input.
[0096] The above description of the disclosed embodiments enables one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A backdoor defense method for a diffusion model based on trigger synthesis and backdoor parameter fine-tuning, characterized in that: The method is applicable to non-conditional generation tasks and text graph tasks based on diffusion models; the backdoor defense method comprises the following steps: Step S1: Generate a batch of clean samples that are not affected by the backdoor using the diffusion model to be detected to form a clean data set; Step S2: Design a target loss function and iteratively optimize the synthetic trigger; use the synthetic trigger to calculate the backdoor indicator to determine whether there is a backdoor in the model; Step S3: If a backdoor is detected, the importance scores of each parameter in the diffusion model for the clean sampling and backdoor sampling processes are calculated; Step S4: Identify parameters related to the backdoor based on the importance score; freeze model parameters unrelated to the backdoor, initialize and fine-tune parameters related to the backdoor; after fine-tuning, the backdoor in the diffusion model is removed.
2. The backdoor defense method for diffusion model according to claim 1 is characterized in that In step S1, the specific steps are: Step S1.1: Determine the structure, specific parameters, sampler, input data dimension, and clean training loss of the diffusion model M0 to be tested Determine the value and range of the time step t during model training and inference, and the range is recorded as [T min ,T max ]; where the clean training loss is the training loss function of the diffusion model when it is not affected by the backdoor attack; The t in it represents the time step, and c represents that the loss is a clean loss; Step S1.2: Use a random number generator to generate a clean input that meets the input data dimension of M0; perform a sampling process to convert the clean input into a clean sample; wherein the clean input is input data without triggers, and the clean sample is noise-free data obtained by performing a sampling process using the clean input by the diffusion model; Step S1.3: Repeat step S1.2 to obtain multiple non-repetitive clean samples to form a clean data set.
3. The backdoor defense method for diffusion model according to claim 1 is characterized in that In step S2, the specific steps are: Step S2.1: Generate a clean input that conforms to the input data dimension of M0, with a shape of [bs, dim1]; where bs is the number of samples used in this step and dim1 is the dimension of a single input sample of M0; Step S2.2: Generate a random vector of shape [1, dim1] as the initial value of the synthetic trigger; Step S2.3: Add the synthetic trigger to each clean input sample element by element to obtain the backdoor input; send the backdoor input to the diffusion model, perform the sampling process, and obtain the backdoor sample; wherein the backdoor input is the input data containing the synthetic trigger, and the backdoor sample is the noise-free data obtained by the diffusion model using the backdoor input to perform the sampling process; Step S2.4: Use the pre-trained feature encoder to map the backdoor sample to the embedding space to obtain an embedding vector Z of shape [bs, dim2]; where dim2 is the dimension of a single sample in the embedding space; Step S2.5: Use Z to calculate the backdoor indicator, the maximum similar cluster ratio MSCR (Max Similar Cluster Ratio): Among them, S(Z i ,Z j ) is used to calculate Z i and Z j The cosine similarity of the two vectors; |·| in the numerator and denominator is the number of elements in the calculation set, and |Z|=bs; κ is the similarity threshold, which is selected according to the type of M0; if the value of MSCR is greater than the preset backdoor indication threshold, it is determined that there is a backdoor in M0 and the process goes to step S3, otherwise, step S2.6 is executed; Step S2.6: If the MSCR does not exceed the backdoor indication threshold, the similarity loss loss is calculated using the embedded vector and the backdoor sample respectively. s and entropy loss e : Among them, S(Z i ,Z j ) is used to calculate Z i and Z j The cosine similarity of these two vectors; |·| is the number of elements in the calculation set, and |Z| = bs; is the entropy of the ith backdoor sample, κ + and κ - is a predefined threshold; max(·,·) is the maximum value between two values; Step S2.7: Define the optimization target loss o loss s and loss e The specific formula of this objective is as follows: min loss o =loss s +λloss e Among them, λ is the dynamically adjusted weight coefficient; loss is calculated by back propagation o The gradient of the synthetic trigger is calculated and the Adam optimizer is used to update the synthetic trigger. The optimization goal is to make the loss o Minimum; Step S2.8: Repeat steps S2.3 to S2.7 until the MSCR is greater than the backdoor indication threshold or the number of iterations reaches a predefined upper limit.
4. The backdoor defense method for diffusion model according to claim 1 is characterized in that In step S3, the specific steps are: Step S3.1: Get a batch of clean samples from the clean dataset, whose shape is [bs, dim1]; where bs is the number of samples used in this step, and dim1 is the dimension of a single input sample of M0; Step S3.2: Generate a clean input that conforms to the input data dimension of M0, with a shape of [bs, dim1]; add the synthetic trigger to each clean input sample element by element to obtain the backdoor input; send the backdoor input to the diffusion model, perform the sampling process, and obtain the backdoor sample; Step S3.3: Calculate the clean training loss using clean samples and backdoor samples respectively and backdoor training loss Among them, t represents the time step, c and b represent the loss is clean training loss and backdoor training loss respectively; backdoor training loss is the specific training loss used by the attacker to implant a backdoor into the diffusion model. There are two cases for its calculation formula: when the defender has an understanding of the potential backdoor attack method, it can be modified appropriately To get formula and use the backdoor sample; if you have no knowledge of the potential backdoor attack method, you can directly use The formula only replaces the input data from clean samples to backdoor samples; Step S3.4: Calculate the Taylor importance score of each parameter in M0 respectively; regard all parameters θ of M0 as a two-dimensional matrix, where each row vector θ = [θ i1 ,θ i2 ,…,θ iK ] represents a parameter containing K scalars; the value range of i is determined according to the parameter value of M0; θ i The Taylor importance score I(θ i ) can be calculated by the following formula: Among them, |·| is to calculate the absolute value of its internal vector; θ ik is θ i is the kth scalar parameter in ; t is the time step in the training process; represent For θ ik Derivation; is the training loss at time step t, When, I(θ i ) is θ i Importance score for clean sampling process I ci ; When, I(θ i ) is θ i Importance score for the backdoor sampling process I bi ; Calculate I ci and I bi .
5. The backdoor defense method for diffusion model according to claim 1 is characterized in that In step S4, the specific steps are: Step S4.1: Calculate the differential importance score of each parameter in θ, θ i The differential importance score of di =I bi -I ci ; Step S4.2: Select θ in |I di |The largest part of parameters is taken as backdoor related parameters, and the remaining parameters are taken as backdoor irrelevant parameters; where |·| is the absolute value of the internal vector; the proportion of backdoor related parameters is selected according to the model structure and parameter quantity of M0; Step S4.3: Create a backup M1 of M0, whose parameters are recorded as θ'; freeze all parameters in θ and backdoor-independent parameters in θ', and initialize backdoor-related parameters in θ'; Step S4.4: Define reference loss in, is the clean input of the diffusion model at time step t, ∈ θ (·) and ∈ θ' (·) represent the output of models M0 and M1 respectively, t∈[T min ,T max ] is the time step in the training process, ||·||2 is the 2-norm of its internal vector; Step S4.5: Define fine-tuning loss It is the clean training loss and reference loss The sum of: Step S4.6: Using the clean dataset and Fine-tune the backdoor related parameters in θ'; this process gradually removes the backdoor in the model while maintaining its original generative ability.