Method and device for evaluating robustness of diffusion-based purification model
By employing the DiffHammer selective attack method, the EM algorithm is used to identify vulnerable purification processes and reduce gradient complexity. Combined with N evaluations, the gradient dilemma and insufficient resubmission risk assessment of the diffusion purification model are resolved, achieving efficient robustness assessment.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- THE HONG KONG UNIV OF SCI & TECH
- Filing Date
- 2025-10-09
- Publication Date
- 2026-05-08
AI Technical Summary
Existing diffusion-based cleanup models suffer from gradient dilemmas and insufficient assessment of resubmission risk when evaluating robustness, resulting in poor robustness assessment performance.
A selective attack method, DiffHammer, is proposed. It identifies vulnerable purification processes through the expectation-maximization (EM) algorithm, reduces backpropagation complexity through gradient grafting, and comprehensively evaluates the robustness of the model by combining an N-times evaluation protocol.
It achieved near-optimal attack performance within 10-30 iterations, overcame the gradient dilemma, accurately assessed the risk of resubmission, and improved the robustness assessment efficiency and accuracy of the diffusion purification model.
Smart Images

Figure CN121999264A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of this disclosure relate to the field of computer technology, and more particularly to methods and apparatus for evaluating the robustness of diffusion-based cleanup models. Background Technology
[0002] Diffusion-based purification, as an adversarial defense method, has demonstrated impressive robustness. However, concerns have arisen regarding whether this robustness stems from insufficient evaluation. Existing research indicates that attacks based on Expectation of Transformation (EOT) suffer from a gradient dilemma due to global gradient averaging, leading to poor evaluation performance. Furthermore, a single evaluation (1-evaluation) underestimates the resubmission risk in stochastic defenses. To address these issues, this application proposes an efficient attack method—DiffHammer. This method bypasses the gradient dilemma by selectively attacking vulnerable purification processes, incorporating N-evaluations into the loop, and utilizing gradient grafting to achieve comprehensive and efficient evaluation. Experiments demonstrate that DiffHammer achieves ideal results within 10-30 iterations, outperforming other methods. This raises questions about the reliability of diffusion-based purification after mitigating the gradient dilemma and carefully examining its resubmission risk. Summary of the Invention
[0003] Embodiments of this disclosure present methods and apparatus for evaluating the robustness of diffusion-based cleanup models.
[0004] In a first aspect, embodiments of this disclosure provide a method for evaluating the robustness of a diffusion-based purification model, comprising: performing a first number of random purifications on a noisy image through the purification model to obtain a first number of purification results, wherein the noisy image is obtained by superimposing adversarial noise onto a sample image; determining a first number of classification results by passing the first number of purification results through a classifier; determining the loss value and gradient in each random purification process based on the difference between each classification result and the actual category of the sample image; determining a representative random purification process based on the loss value and gradient in each random purification process; performing backpropagation based on the representative random purification process to determine the target gradient; and attacking the purification model based on the target gradient to determine the robustness of the purification model.
[0005] Secondly, embodiments of this disclosure provide an apparatus for evaluating the robustness of a diffusion-based purification model, comprising: a purification unit configured to perform a first number of random purifications on a noisy image through the purification model to obtain a first number of purification results, wherein the noisy image is obtained by superimposing adversarial noise onto a sample image; a classification unit configured to determine a first number of classification results by passing the first number of purification results through a classifier; a calculation unit configured to determine a loss value and gradient in each random purification process based on the difference between each classification result and the actual category of the sample image; a selection unit configured to determine a representative random purification process based on the loss value and gradient in each random purification process; a propagation unit configured to perform backgrad propagation based on the representative random purification process to determine a target gradient; and an attack unit configured to attack the purification model based on the target gradient to determine the robustness of the purification model.
[0006] Thirdly, embodiments of this disclosure provide an electronic device, including: one or more processors; and a storage device having one or more computer programs stored thereon, which, when executed by the one or more processors, cause the one or more processors to perform the method as described in any one of the first or second aspects.
[0007] Fourthly, embodiments of this disclosure provide a computer-readable medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the method as described in any one of the first or second aspects.
[0008] Fifthly, embodiments of this disclosure provide a computer program product including a computer program that, when executed by a processor, implements the method as described in any one of the first or second aspects.
[0009] The embodiments of this disclosure provide a method and apparatus for evaluating the robustness of diffusion-based purification models, and propose a selective attack method, DiffHammer, based on the expectation-maximization (EM) algorithm. This algorithm bypasses the gradient dilemma by identifying vulnerable purification processes in the E-step iterations and aggregating their gradients in the M-step. This method focuses on shared vulnerabilities, analogous to striking weak points with a hammer rather than the entire structure. Furthermore, by grafting gradients, the backpropagation complexity of the diffusion process is reduced from O(N) to O(1), improving attack efficiency.
[0010] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0011] Other features, objects, and advantages of this disclosure will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings: Figure 1 This is an exemplary system architecture diagram to which one embodiment of this disclosure can be applied; Figure 2 This is a flowchart of an embodiment of a method for evaluating the robustness of a diffusion-based cleanup model according to this disclosure; Figures 3.1-3.5 This is a schematic diagram illustrating an application scenario of the robustness method for evaluating a diffusion-based cleanup model according to this disclosure; Figure 4 This is a schematic diagram of a structure of an embodiment of a robustness device for evaluating a diffusion-based purification model according to the present disclosure; Figure 5 This is a schematic diagram of the structure of a computer system suitable for implementing embodiments of the present disclosure. Detailed Implementation
[0012] The present disclosure will now be described in further detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and not intended to limit it. Furthermore, it should be noted that, for ease of description, only the parts relevant to the invention are shown in the accompanying drawings.
[0013] It should be noted that, unless otherwise specified, the embodiments and features described in this disclosure can be combined with each other. This disclosure will now be described in detail with reference to the accompanying drawings and embodiments.
[0014] Figure 1 An exemplary system architecture 100 is shown, representing an embodiment of a method or apparatus for evaluating the robustness of a diffusion-based cleanup model to which the present disclosure may be applied.
[0015] like Figure 1 As shown, system architecture 100 may include terminal devices 101, 102, and 103, a network 104, and a server 105. Network 104 serves as the medium for providing communication links between terminal devices 101, 102, and 103 and server 105. Network 104 may include various connection types, such as wired or wireless communication links, or fiber optic cables, etc.
[0016] Users can use terminal devices 101, 102, and 103 to interact with server 105 via network 104 to receive or send messages, etc. Various communication client applications can be installed on terminal devices 101, 102, and 103, such as facial recognition applications, autonomous vehicle control programs, 3D video players, web browser applications, shopping applications, search applications, instant messaging tools, email clients, social media platform software, etc.
[0017] Server 105 can be a server providing various services, such as a backend classification server for classifying images provided by terminal devices 101, 102, and 103. The backend classification server has an image classification model installed, which includes a diffusion-based cleanup model and a classifier. The cleanup model is used to clean the images provided by terminal devices 101, 102, and 103, and then the classifier classifies the images. Server 105 can also evaluate the robustness of the diffusion-based cleanup model. This allows for effective assessment of defense risks and the deployment of suitable image classification models for different application scenarios.
[0018] It should be noted that the method for evaluating the robustness of a diffusion-based cleanup model provided in the embodiments of this disclosure is generally executed by server 105, and correspondingly, the apparatus for evaluating the robustness of a diffusion-based cleanup model is generally located in server 105.
[0019] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.
[0020] Continue to refer to Figure 2 The diagram illustrates a flow 200 of an embodiment of a method for evaluating the robustness of a diffusion-based cleanup model according to the present disclosure. This method for evaluating the robustness of a diffusion-based cleanup model includes the following steps: Step 201: The noisy image is randomly purified a first number of times using the purification model to obtain a first number of purification results.
[0021] In this embodiment, the implementer of the method for evaluating the robustness of the diffusion-based purification model (e.g., Figure 1 The server shown can acquire sample images and superimpose anti-noise onto the sample images to obtain a noisy image. Initially, random noise can be added.
[0022] For example, select one sample image x (actual class) from the CIFAR10 test set. y =“cat”), initialize adversarial noise r (0) To obtain an initial noisy image x+, the values are set to random to ensure that the noise is imperceptible.r (0) N random purifications: x+ r (0) Input the DiffPure purification model and perform 10 independent random purification runs (due to the randomness of the diffusion model, each purification process...). (Differences exist in noise sampling and denoising step size for i=1,2,...,10), resulting in 10 purification results. .
[0023] Step 202: Determine the first number of classification results by using a classifier for the first number of purification results.
[0024] In this embodiment, the 10 purification results obtained in step 201 are input one by one into a classifier (e.g., WideResNet-70-16), and the predicted category of each purification result is output: Classifier prediction logic: ,in Let represent the logit probability of the k-th class.
[0025] Example results: Of the 10 classification results, 6 were predicted as "dog", 3 as "fox", and 1 as "cat" (i.e., consistent with the actual category y), denoted as . = {dog, dog, fox, cat, dog, fox, dog, fox, dog, dog}.
[0026] Step 203: Based on the difference between each classification result and the actual category of the sample image, determine the loss value and gradient in each random purification process.
[0027] In this embodiment, for each purification process Based on classification results Calculate the loss value and gradient based on the difference from the actual class y: 1. Loss calculation: The cross-entropy loss function is used to measure the classifier's confidence in the "correct class y" - the larger the loss value, the lower the classifier's confidence in the correct class, and the easier it is for an attack to succeed.
[0028] 2. Gradient calculation: Approximate gradient: calculated using the BPDA method. The purification process Treated as an identity mapping, directly affecting the purification result Differentiation has low computational cost.
[0029] Full gradient: Theoretically, it needs to be calculated However, direct calculation is costly, so we will not calculate it directly for now, but will optimize it through gradient grafting later.
[0030] Ultimately, 10 pairs of "loss value - approximate gradient" were obtained.
[0031] Step 204: Determine a representative random cleanup process based on the loss value and gradient in each random cleanup process.
[0032] In this embodiment, based on the loss value and gradient from step 203, the "sanitization set S1 of shared vulnerabilities" is identified through the E-step of the EM algorithm, and a representative sanitization process is selected from it. : 1. Step E: Identify S1: Optimize virtual adversarial noise increment The goal is to maximize the probability of attacking more purification processes: ,in For the Sigmoid function (which maps the loss to the attack success probability in [0,1]).
[0033] Calculate each purification process The probability of belonging to S1 ,q i The closer to 1, the better. It is more susceptible to the same adversarial noise attack (belonging to S1).
[0034] Example:
[0035] 2. Select representative : Choose the option that "maximizes the attack on other purification processes in S1". As :
[0036] in of =0.99 and can maximize the attack probability against S1, therefore = .
[0037] Step 205: Perform backgrad propagation based on a representative stochastic purification process to determine the target gradient.
[0038] In this embodiment, the gradients of S1 are aggregated through gradient grafting, and based on... Perform backpropagation to obtain the target gradient g(t): 1. Gradient aggregation: Calculate the weighted sum of the approximate gradients in S1 (with weights qi).
[0039] 2. Gradient grafting and backpropagation: use The derivative of the purification process Approximate all in S1 of Then the approximation of the target gradient is:
[0040] 3. Stepwise-EM Gradient Update: To stabilize the gradient direction, the current gradient is merged with historical gradients: , where α=0.5 and t is the number of iterations.
[0041] Gradient grafting reduces the gradient computation complexity from O(N) to O(1) (requiring only O(N) to O(1)). Perform one full gradient backpropagation (while retaining the shared vulnerability information of S1).
[0042] Step 206: Attack the purification model based on the target gradient to determine the robustness of the purification model.
[0043] In this embodiment, the target gradient g(t) is input into the attack algorithm (e.g., AA) to update the adversarial noise.
[0044] Repeat steps 201-205 for multiple iterations (e.g., 17 times) to obtain the optimized adversarial noise.
[0045] Perform 10 random cleanups based on the optimized adversarial noise, and calculate the average robustness and worst-case robustness (measures the worst case where a sample is successfully attacked in 10 resubmissions).
[0046] The method provided in the above embodiments of this disclosure, through the core technologies of the DiffHammer framework (N-times randomized purification, EM screening for shared vulnerability, gradient grafting, and N-evaluation), successfully quantifies the true robustness of the DiffPure purification model. 1. The “randomness” of diffusion-based purification is not the true source of robustness, but rather a “gradient dilemma” that causes EOT-type attacks to fail. This dilemma can be overcome by selectively attacking S1. 2. Traditional 1-evaluation overestimates robustness; N-evaluation (N ≥ number of resubmissions M) is a more reliable evaluation protocol. 3. DiffHammer can achieve near-optimal attack results within 10-30 iterations, providing an efficient solution for rapid robustness testing of diffusion-based purification models.
[0047] In some optional implementations of this embodiment, determining a representative random purification process based on the loss value and gradient in each random purification process includes: determining the shared adversarial noise that can influence the random purification process the most based on the loss value and gradient in each random purification process; and determining a representative random purification process based on the shared adversarial noise.
[0048] 1. Define the optimization variables and objective function: set up To combat the noise increment, the objective function is to "maximize the average probability of 10 random purification attacks", that is:
[0049] Among them, The Sigmoid function maps the "loss value + gradient increment contribution" to the [0,1] interval. The closer the value is to 1, the better the purification process is. The higher the probability of a successful attack; This indicates an expectation for 10 purification processes, with the goal of making... For as many as possible Both can increase the probability of attack.
[0050] 2. Solving the optimization problem: Because The optimization relies only on approximate gradients (without needing to call the cleanup model for recalculation), and can be solved quickly using the gradient ascent method.
[0051] 3. Determine shared adversarial noise: Shared adversarial noise r This is the superposition of the initial noise and the optimal increment. Through verification, r... It can ensure that 8 out of 10 purification attacks succeed (the classification result is inconsistent with y), and only 2 attacks fail, proving its characteristic of "affecting the most random purification process".
[0052] 4. Representative randomized purification process It must satisfy the condition that "it can represent the vulnerabilities in set S1, and its gradient can approximate the gradient of other cleanup processes in S1". The specific selection logic is as follows: For each purification process in S1 Substitute the loss value and gradient calculation The gradient's contribution to the attack on other purification processes in S1. A higher contribution indicates a greater attack contribution. The better the gradient direction matches the vulnerabilities of other purification processes in S1, the more suitable it is as a representative purification process.
[0053] In some optional implementations of this embodiment, determining a representative random cleanup process based on shared adversarial noise includes: estimating the loss caused by shared adversarial noise to each random cleanup process; mapping the loss caused by each random cleanup process to the probability that a shared vulnerability exists in the random cleanup process; and determining a representative random cleanup process by maximizing the probability.
[0054] The cross-entropy loss function is used to estimate the loss caused by shared adversarial noise in each random cleanup process: ,in, This represents the classifier's confidence in the "actual class y of the sample"—the higher the loss value, the lower the classifier's confidence in the correct class. Because r The more obvious the vulnerabilities are.
[0055] Since the loss value is positively correlated with the existence of shared vulnerabilities (the greater the loss, the more likely the vulnerability exists), the loss value needs to be mapped to the [0,1] interval through a monotonically increasing function as the probability qi of the shared vulnerability.
[0056] A representative randomized cleanup process must satisfy two core conditions: (1) the probability qi of having a shared vulnerability is as high as possible; and (2) its gradient can "drive" other vulnerabilities. The probability of shared vulnerabilities is maximized. This is specifically achieved through "maximizing the probability-weighted score".
[0057] In some optional implementations of this embodiment, backpropagation based on representative random cleanup processes is used to determine the target gradient, including: determining the probability distribution of misclassification by maximizing the loss value in each random cleanup process; determining the aggregated gradient by weighted summation of the gradients in each random cleanup process based on the probability distribution; and determining the target gradient by backpropagation based on representative random cleanup processes and aggregated gradients.
[0058] Maximizing the loss value is equivalent to maximizing the misclassification probability (the higher the loss value, the lower the classifier's confidence in the correct category, and the higher the misclassification probability). Maximizing the loss value in each cleanup involves mapping the loss value to the probability distribution of misclassification using a monotonically increasing function (such as the sigmoid function).
[0059] The gradient of each purification is calculated by differentiating the purification result, and then summed based on a probability distribution, where the weights are the misclassification probabilities.
[0060] Low-complexity backpropagation is achieved by approximating all derivatives of S1 using the derivatives of a representative purification process (i.e., gradient grafting).
[0061] In some optional implementations of this embodiment, backpropagation based on a representative random cleanup process and aggregated gradients is used to determine the target gradient, including: cleaning the noisy image through a representative random cleanup process to obtain a cleanup result; and performing backpropagation on the inner product of the cleanup result and the aggregated gradient to determine the target gradient.
[0062] A representative cleansing process is used to process noisy images, obtaining a "representative cleansing result" for backpropagation. The aggregated gradient information is injected through "inner product calculation → backpropagation". The gradient path enables low-complexity objective gradient solving.
[0063] In some optional implementations of this embodiment, attacking the purification model based on the target gradient to determine the robustness of the purification model includes: attacking the purification model based on the target gradient to determine the robustness of the current attack and the updated adversarial noise; repeating the following attack steps until the number of repetitions reaches a second number of times: superimposing the updated adversarial noise on the sample image to obtain an updated noisy image; performing a first number of random purifications on the updated noisy image through the purification model, and then determining a first number of classification results through a classifier; determining the loss value and gradient in each random purification process based on the difference between each classification result and the actual category of the sample image; determining a representative random purification process based on the loss value and gradient in each random purification process; performing backpropagation based on the representative random purification process to determine the target gradient; attacking the purification model based on the target gradient to determine the robustness of the current attack and the updated adversarial noise; and determining the worst value of the obtained robustness as the robustness of the purification model.
[0064] The initial attack lays the foundation for subsequent iterations, with the core being the transition from "target gradient calculation" to "robustness evaluation." Repeated attacks are performed, each starting with the "updated adversarial noise," repeating the process of "generating a noisy image → N cleansing iterations → calculating loss and gradient → selecting representative cleansing → calculating the target gradient → attacking → calculating robustness." All robustness results from all attacks are collected, and the minimum value is taken as the final robustness.
[0065] In some optional implementations of this embodiment, backpropagation based on a representative stochastic purification process is used to determine the target gradient, including: backpropagation based on a representative stochastic purification process to determine intermediate gradients; and weighting the intermediate gradients based on the number of repetitions to determine the target gradient.
[0066] intermediate gradient g mid This is the product of the derivative and the aggregate gradient (result of backpropagation). The target gradient g targetThe "current intermediate gradient" and the "historical target gradient" need to be integrated, with the weights determined by the number of attack repetitions t, following the Stepwise-EM gradient update logic:
[0067] The weighting coefficient is , where w t The weights of the gradient in the current round, 1-w t The weights of the historical target gradient are w; the smaller t is, the greater w is. t The larger the value of w, the more it depends on the current gradient; the larger the value of t, the more it depends on the current gradient. t The smaller the value, the more dependent it is on historical gradients.
[0068] The vulnerability of deep neural networks (DNNs) to adversarial attacks hinders their application in security-critical fields—attackers can deceive DNNs by introducing carefully crafted noise. To mitigate this problem, researchers have proposed various defense strategies to improve their robustness, among which diffusion-based sanitization methods have emerged as a promising approach. The diffusion model aims to construct a stochastic process from a noisy data distribution to a clean data distribution. Therefore, small adversarial noises are masked by larger noises, and sanitization can then be achieved through iterative denoising of the diffusion model or optimization of the diffusion process. The iterative algorithm and stochasticity of the diffusion process empirically improve sanitization capabilities, but also present challenges to robustness evaluation. Its effectiveness remains questionable: does its effectiveness stem from inherent robustness, or is it due to insufficient evaluation?
[0069] The Expectation Transformation (EOT) method averages the gradients of the sampled cleanup process, allowing attacks to adapt to the random cleanup process. This method assumes that most cleanup processes share a vulnerability (i.e., are susceptible to the same adversarial noise attack), thus maximizing the attack success rate. However, diffusion-based cleanup processes are highly random, making this assumption difficult to hold, resulting in poor EOT-based attack performance. A cleanup process with a non-shared vulnerability (denoted as...) ,like Figure 3.1 As shown, this can generate inconsistent gradients during an attack, leading to a gradient dilemma. Furthermore, obtaining the gradients for the purification process is time-consuming, and this gradient dilemma further increases computational overhead.
[0070] To address this problem, we propose a selective attack method, DiffHammer, based on the Expectation-Maximization (EM) algorithm. For example... Figure 3.1As shown, the algorithm bypasses the gradient dilemma by iteratively identifying vulnerable cleanup processes in the E-step and aggregating their gradients in the M-step. This approach focuses on shared vulnerabilities, similar to striking weak points with a hammer rather than the entire structure. Furthermore, we improve attack efficiency by reducing the backpropagation complexity of the diffusion process from O(N) to O(1) through gradient grafting.
[0071] Another challenge with diffusion-based cleanup is insufficient assessment. For attackers, if a single attack is sufficient to achieve their objective (such as login verification), they can increase their success rate by resubmitting adversarial examples to a randomized defense system. A single evaluation (1-evaluation) designed for deterministic defenses struggles to accurately assess this resubmission risk, especially when the diffusion process is highly random. To address this, we upgrade the single evaluation to an N-evaluation and seamlessly integrate it into the DiffHammer framework, achieving a more comprehensive risk assessment without incurring additional time costs.
[0072] Figure 3.1 The DiffHammer principle diagram from left to right: (a) has a non-shared vulnerability ( ) and shared vulnerabilities ( (a) A collection of purification processes; (b) and (c) For ∪ Maximizing the loss function may result in a weaker attack effect than simply targeting... To launch an attack; (d)(e) Selective attack To avoid the gradient dilemma—identifying in the E-step The purification process in the process aggregates the gradients of these purification processes in the M step.
[0073] We re-examine the robustness of diffusion-based cleanup within the DiffHammer framework. DiffHammer achieves near-optimal attack performance within 10-30 iterations by circumventing the gradient dilemma, outperforming other state-of-the-art attack methods. Under the combined effects of the N-times evaluation protocol and selective attacks, we find that the risk of diffusion-based cleanup is significantly underestimated. Mainstream cleanup methods even struggle against 10 resubmission attacks and may even degrade the performance of robust models, indicating that their robustness potential has not been fully realized. We hope that DiffHammer's insights into evaluating the robustness of diffusion-based cleanup will drive the development of more robust defense methods. The main contributions of this paper are as follows: - It points out the limitations of EOT-based attacks and proposes an efficient evaluation method (i.e., DiffHammer) that can comprehensively diagnose adversarial risks of diffusion-based purification. - A selective attack strategy is proposed, which avoids the gradient dilemma by targeting shared vulnerabilities and accelerates the attack process by using gradient grafting; - The effectiveness of DiffHammer was verified through extensive experiments on mainstream purification methods, and the underestimated risks were revealed through standardized N evaluations.
[0074] Given an image x∈ and its label y∈[K], with preprocessor : → Classifier f: → It will classify the image into f[ (x)]=argma ,…,K [ (x)], where K is the number of classes. The attacker's goal is to add imperceptible adversarial noise r to the image, causing the classifier to misclassify it, i.e., f[ [x+r)]≠y. Adversarial noise is typically addressed by maximizing the loss function L and applying... Norm constraints are used to ensure imperceptibility. For example, in Projected Gradient Descent (PGD), the adversarial example update formula satisfying the constraint ‖r‖∞≤ε is:
[0075] Where α is the step size and Π is the projection operator.
[0076] Preprocessors used to defend against adversarial noise can lead to vanishing gradients, exploding gradients, or inaccurate gradients. Attackers can approximate gradients using Backward Pass Differentiable Approximation (BPDA), where the preprocessor is typically treated as an identity mapping during backpropagation.
[0077] For randomized preprocessors, the EOT method can be used to estimate gradients, and its effectiveness is based on the assumption that "different preprocessors share vulnerabilities". Given N sampled preprocessors... The estimated gradient is:
[0078] The diffusion model establishes a clean data distribution through forward (noise addition) and reverse (noise removal) processes. (x) and the distribution of noisy data The relationship between (x). Forward process This can be represented as a stochastic differential equation (SDE) where t ranges from 0 to 1:
[0079] in, (x), ∈ For the standard Wiener process, f: × → Let g be the drift coefficient, and R be the diffusion coefficient (designed to make...). (x) follows a standard Gaussian distribution N(0, I) d The deep diffusion probabilistic model (DDPM) can be considered a special case of this SDE, where f(x, t) := -β(t)x / 2, g(t) := , where β(t) is the noise scheduler, which is usually set to be a linear function of t.
[0080] The evolution of the reverse-time SDE (t from 1 to 0) corresponds to sample generation:
[0081] Where dt is the infinitesimal negative time step. For the standard reverse-time Wiener process, ( ) is a time-dependent score function, which is usually estimated by a neural network.
[0082] A well-trained diffusion model can accurately model the score function and denoise noisy samples, thus achieving adversarial cleanup. Considering the small amplitude of adversarial noise, DiffPure uses a time range from t=0 to t0. The forward process masks adversarial noise (while preserving semantic information), and then purification is achieved through a denoising process. GDMP (Guided Diffusion Model Patching) uses the distance from the initial sample as a guide to further preserve semantic information. Using a diffusion model can more accurately approximate x. This property of the score function of (x) allows Likelihood Maximization (LM) to optimize adversarial examples to minimize the estimation error, making its distribution similar to... (x) Alignment.
[0083] To address the issues of ineffective attack phases and insufficient evaluation phases, we propose DiffHammer for evaluating the robustness of diffusion-based cleanup. First, DiffHammer overcomes the gradient dilemma by designing a selective attack based on the EM algorithm. Second, it seamlessly integrates N evaluations into the algorithm to accurately diagnose resubmission risks and provides approximate gradient-assisted attacks. The selective attack comprises two parts: the EM algorithm and gradient grafting.
[0084] The EM algorithm is as follows: Symbols and Objectives: Given a sample x∈ Given its label y∈[K], we aim to design adversarial noise r∈ To mislead about random purification processes : → The classifier f. Let A represent " "The goal of this misclassification event is to maximize the misclassification probability, that is..." Where P represents the random purification process. The relevant probability density, p is a given specific The log probability density of the classifier f over time. We will use the loss function. gradient with respect to x This is called the complete gradient or full gradient. gradient This is called the approximate gradient. In unambiguous cases, the loss function is sometimes simply written as... The gradient is abbreviated as .
[0085] Assumption: Assume that the diffusion-based cleanup process has a non-shared vulnerability, which could be... They are divided into two categories. (Note:) The largest with shared vulnerabilities Set, that is =argmax P({ : f[ (x+r )]≠y (for the same r) In terms of) The complement of is affected by inconsistent adversarial noise r. Set. Assume that for... In Optimization attempts can encounter gradient hindrance, leading to poor attack effectiveness. Therefore, our task is to identify... and for In Design r.
[0086] Represented by z=1 or z=0 belong or q(z) represents the given The estimated distribution of z. According to Jensen's inequality, we maximize the lower bound of the objective function:
[0087] Here, Q(q,r) is the lower bound of evidence, and KL(q‖p(z|A,r)) is the KL divergence between the posterior distributions of q and z. Since the variables r and z are coupled during the optimization process, we use the EM algorithm as the solver, optimizing r in the M-step (maximizing Q(q,r)) and estimating z in the E-step (minimizing the KL divergence), alternating between the two steps.
[0088] M-step: In the M-step, q(z) is fixed, and the lower bound of the objective function is increased by maximizing Q(q,r). Ignoring terms independent of r (i.e., q(z)ln q(z) and prior terms ln p(z=0), ln p(z=1)), the objective can be simplified to:
[0089] Based on the assumption, the second term can be further ignored in the optimization.
[0090] In adversarial attacks, the objective function Typically replaced by maximizing the loss function. This can be achieved using a method based on average gradients. Therefore, the M-step update can be integrated as a plugin into existing attack algorithms such as PGD and AA, the difference being that each... The gradient needs to be weighted according to q (z=1). We further adopt the stepwise EM algorithm, using weights... Online updates are achieved by linearly interpolating the current gradient with the previous gradient, where t is the number of iterations and α is a hyperparameter. This method intuitively guides adversarial noise optimization towards... To avoid shared vulnerabilities. The gradient dilemma in [the context of the text].
[0091] E-step: In the E-step, for a given r, q(z=1) is updated to p(z=1|A,r) to eliminate the gap between the objective function and its lower bound (i.e., KL divergence). Note that... Defined as susceptible to the same adversarial noise r The maximum attack Set, and usually r ≠r. We can utilize the byproducts from the observations of misclassified event A. and By approximating r Then, estimate p(z=1|A,r). Let r be... The difference between r and r is Δr. Since Δr is not large (r and r... All satisfy || ≤ε constraint), loss function Regarding r The linear approximation is:
[0092] The greater the loss, the higher the probability of misclassification. Therefore, a monotonically increasing function σ:R→[0,1] can be used to... Mapped to p(A|r) ).according to Definition of r It should be able to influence as many as possible. Therefore, r is solved by the following optimization. :
[0093] Determine r back, belong The probability q (z=1) can be estimated as .
[0094] Note: In each iteration, we process N... Example ( The expectation is estimated by the empirical average of (i=1,…,N). In the E-step, r The optimization does not involve the model and has a low cost. In the M-step, through For each The gradients are weighted and then used in existing attack algorithms (such as PGD) to update the adversarial noise r. The computational cost mainly stems from the complexity of the cleanup process—computing the gradients. This requires a significant amount of time. We will later explain how to reduce the computational complexity from O(N) to O(1) through gradient grafting.
[0095] Algorithm 1: DiffHammer
[0096] Input: Data (x, y), with a stochastic purification model Given a classifier f, the number of resubmissions M, and existing attack algorithms.
[0097] Output: Rob's robustness after M resubmissions
[0098] 1. Initialize r (0) ;
[0099] 2 for t ← 1 to T do
[0100] 3 / / t-1 iterations are used for evaluation, and t iterations are used for input;
[0101] 4 Sampling i , i = 1, ..., N;
[0102] 5 Rob(t 1) = Eval (r(t 1), M) / / Used for evaluation;
[0103] 6 Records / / Used for attacks
[0104] 7 / / Step E: Identify sets with shared vulnerabilities;
[0105] 8
[0106] 9 / / Probability of membership;
[0107] 10 / / M-step: Estimating the aggregation of all gradients;
[0108] 11 / / exist Phase aggregation;
[0109] 12. Select according to formula 11 / / A representative randomized purification process;
[0110] 13 / / Gradient grafting;
[0111] 14 / / Updated gradually;
[0112] 15 r(t) = AttackAlgorithm (r(t) 1), g(t)); / / Attack
[0113] 16 end
[0114] 17 Rob(T) = Eval (r(T), M) with ;
[0115] 18. Return Rob = min(Rob(t), t = 1, ..., T)
[0116] Gradient Grafting: To address the efficiency issue in diffusion-based robustness assessment of cleanup, we aim to estimate weighted gradient aggregation with minimal computational cost. Our method aggregates low-cost approximate gradients in the early stages. To estimate the weighted complete gradient, the expression is:
[0117] Where, assuming / x can be derived from / x is approximated. For example, Figure 3.2 As shown, we first calculate the weighted approximate gradient. Then "graft" it onto Backpropagation is performed on the upper layer to estimate the gradient expectation, thereby shifting x← The backpropagation complexity of (x) is reduced to O(1). The BPDA method approximates the identity matrix I. / While x is computationally simple, oversimplification can lead to performance degradation. Our gradient grafting improves estimation accuracy through an additional backpropagation, achieving a better balance between efficiency and performance. The relevant design details will be discussed further below.
[0118] Phase E-step: Although r in E-step (Formula 3) Involving the full gradient But we mainly focus on Therefore, it is possible Optimize in stage x rather than stage x to avoid relying on the full gradient. stage Represented as Then "attack as many as possible" This optimization objective can be rewritten as: This formula depends only on the approximate gradient. Finally, through the optimized... Derivation .
[0119] Selection: We hope to select a vulnerability that is "representative". During backpropagation, its adversarial noise can also affect other... .for Single-step attack (adversarial noise is) ,like Under ∞ constraint ,from (i=1, Choose from (N) The strategy is:
[0120] This "attack as many as possible" The optimization goal ensured that from Consistency of shared vulnerabilities from phase x to phase x. The choice involves discrete optimization within a finite set, which can be solved by traversing when N is not too large.
[0121] Our selective attacks (targeting) It is superior to EOT-based attacks in both effectiveness and efficiency (targeting...). ∪ ): 1. When all When a common vulnerability exists, attacks based on EOT become a special case of our algorithm—in this case... For an empty set, q(z=1)=1, selective attacks degenerate into EOT-based attacks with no additional side effects; 2. Identification Only approximate gradients are needed, with near-zero cost; gradient grafting further improves the computational efficiency of gradient expectations. 3. Selective attacks achieve efficient attacks by circumventing the gradient dilemma, which may cause EOT-based attacks to fail even in simple binary classification scenarios.
[0122] Theorem 1 (Attack Failure Modes Based on EOT): Assumptions The sets can be divided into two categories, A and B, whose loss functions with respect to r are defined as follows: , And are strongly concave functions mA and mB respectively. If the distance between their optimal solutions rA and rB satisfies ||rA|||rB ... rB ²≥ (where m:=min{P(A)mA, P(B)mB}), then the attack effect based on EOT is weaker than a simple attack targeting only A or B.
[0123] When rA and rB are significantly different It is often one of category A or B, while the other becomes It also provides a counteracting gradient. Experimental results show that the gradient dilemma is prevalent in diffusion-based purification, highlighting the necessity of selective attacks.
[0124] N-times evaluation within the loop: Traditional single-time evaluation (1-evaluation) is insufficient to accurately assess the robustness of diffusion-based cleanup, especially in the face of resubmission attacks. As a stochastic defense system, even with the same input query, the model output may be inconsistent, allowing attackers to achieve their goals through resubmission. In scenarios where attack costs are controllable and a single successful attack yields a benefit (such as login verification), defenders need to focus on the model's robustness under M resubmissions. Therefore, we propose using N-times evaluation as the robustness assessment protocol for the following reasons: 1. As stated in Theorem 2, a single assessment is significantly biased in estimating the risk of M resubmissions. This is due to the highly random nature of the diffusion process. Non-shared vulnerabilities are quite common (e.g.) Figure 1 (As shown). A single assessment only records "success" or "failure," failing to capture the crucial probabilistic information needed for resubmission risk estimation. For example, in DiffPure, the attack success rate (ASR) P(…) was 46.3% of the samples. The fact that )∈(0,1) leads to a 17.9% overestimation of the robustness of a single evaluation for 10 resubmissions (Rob_ML) =41.7%, while the actual Rob = 59.6%.
[0125] Theorem 2 (Resubmitting Risk Estimates) Suppose the sample is subjected to an M-th resubmission attack ( , The robustness in A_M) is A consistent minimum variance unbiased estimator (UMVUE) for Rob exists if and only if the number of evaluations N ≥ M; when N ≤ M, the maximum likelihood estimator (MLE) tends to overestimate Rob in expectation, i.e. .
[0126] 2. N evaluations can be seamlessly integrated into the attack loop without additional cost. N evaluations Example ( i=1, The ,N) can be used for the evaluation of the previous iteration as well as as the input of the current iteration, and each evaluation uses unseen instances, so there is no information leakage.
[0127] The DiffHammer framework integrates selective attacks with N-fold evaluation, as detailed in Algorithm 1. Notably, our proposed N-fold resubmission robustness metric extends the traditional single-submission metric to risk assessment more closely aligned with real-world deployment scenarios, and is independent of the attack algorithm design. Therefore, DiffHammer not only provides a more comprehensive risk assessment but also improves attack effectiveness across multiple metrics.
[0128] Application scenarios
[0129] Our method enables rapid robustness evaluation of deep learning models with integrated random purifiers in security-critical domains. Deep neural networks often leverage such purifiers to enhance their robustness against adversarial attacks. Before model deployment, our algorithm can evaluate model robustness through simulation under various attack scenarios. Its core advantages are as follows: (1) achieving more accurate robustness evaluation through selective attacks; (2) comprehensively quantifying resubmission risk through N-evaluation; and (3) providing time efficiency advantages in the evaluation process. Effective robustness evaluation is crucial, as it can prevent the model from suffering attacks of unexpected strength, minimize potential losses, and promote rapid algorithm iteration.
[0130] Here are two application examples: 1. Facial Recognition in Security: In scenarios where facial recognition is used for identity verification, such as laboratories and offices, users typically have a maximum of 5 attempts. To prevent attackers from gaining unauthorized access through adversarial attacks, administrators deploy random purifiers. Our method efficiently assesses defense risks by quantifying the "probability of an attacker breaching the system within 5 attempts," providing support for practical deployment.
[0131] 2. Autonomous Driving: Image recognition algorithms in autonomous vehicles are vulnerable to adversarial attacks, potentially leading to traffic accidents. While random purifiers can serve as a defense, our algorithm can simulate attacks in a low-cost simulation environment, accurately assessing the effectiveness of defensive measures.
[0132] In summary, our algorithm provides an efficient approach to adversarial defense assessment, a capability that is crucial in an era where security is of paramount importance.
[0133] The content protected by the patent: Our proposed DiffHammer framework enables a robust evaluation of diffusion-based purification models, integrating selective attacks with an N-evaluation protocol.
[0134] The selective attack identifies vulnerable purification sets using the expectation-maximization (EM) algorithm, assigns weights to gradients, aggregates approximate gradients using gradient grafting, and performs backpropagation using representative samples.
[0135] In the EM algorithm, the expectation step (E-step) determines the probability that each purification process belongs to the vulnerable set by optimizing "adversarial noise that can attack as many purification processes as possible" and mapping the loss to probability; the maximization step (M-step) then uses these probabilities as weights for the gradient.
[0136] Gradient grafting selects representative samples with shared vulnerabilities and replaces their original gradients with the aggregated approximate gradients for backpropagation.
[0137] N-evaluation involves multiple evaluations using a maximum likelihood estimator (MLE) to assess the model's robustness against resubmission attacks, while also providing approximate gradients to accelerate the next evaluation stage.
[0138] Further reference Figure 4 As an implementation of the methods shown in the above figures, this disclosure provides an embodiment of a device for evaluating the robustness of a diffusion-based purification model, which is similar to... Figure 2 Corresponding to the method embodiments shown, this device can be specifically applied to various electronic devices.
[0139] like Figure 4 As shown, the robustness assessment device 400 for a diffusion-based purification model in this embodiment includes: a purification unit 401, a classification unit 402, a calculation unit 403, a selection unit 404, a propagation unit 405, and an attack unit 406. The purification unit 401 is configured to perform a first number of random purifications on a noisy image using the purification model, obtaining a first number of purification results, wherein the noisy image is obtained by superimposing adversarial noise onto a sample image; the classification unit 402 is configured to use a classifier to determine a first number of classification results from the first number of purification results; the calculation unit 403 is configured to determine the loss value and gradient in each random purification process based on the difference between each classification result and the actual category of the sample image; the selection unit 404 is configured to determine a representative random purification process based on the loss value and gradient in each random purification process; the propagation unit 405 is configured to perform backgrad propagation based on the representative random purification process to determine a target gradient; and the attack unit 406 is configured to attack the purification model based on the target gradient to determine the robustness of the purification model.
[0140] In this embodiment, the specific processing of the purification unit 401, classification unit 402, calculation unit 403, selection unit 404, propagation unit 405, and attack unit 406 of the robustness assessment device 400 based on the diffusion-based purification model can be referred to Figure 2 Steps 201-206 in the corresponding embodiment.
[0141] In some optional implementations of this embodiment, the selection unit 404 is further configured to: determine the shared adversarial noise that can have the greatest impact on the random purification process based on the loss value and gradient in each random purification process; and determine the representative random purification process based on the shared adversarial noise.
[0142] In some optional implementations of this embodiment, the selection unit 404 is further configured to: estimate the loss caused by shared adversarial noise to each random cleanup process; map the loss caused by each random cleanup process to the probability that a shared vulnerability exists in the random cleanup process; and determine a representative random cleanup process by maximizing the probability.
[0143] In some optional implementations of this embodiment, the propagation unit 405 is further configured to: determine the probability distribution of misclassification by maximizing the loss value in each random purification process; determine the aggregate gradient by weighted summation of the gradients in each random purification process based on the probability distribution; and determine the target gradient by backpropagation based on the representative random purification process and the aggregate gradient.
[0144] In some optional implementations of this embodiment, the propagation unit 405 is further configured to: clean the noisy image through a representative random cleanup process to obtain a cleanup result; and perform backgrad propagation of the cleanup result and the inner product of the aggregated gradient to determine the target gradient.
[0145] In some optional implementations of this embodiment, the attack unit 406 is further configured to: attack the purification model based on the target gradient, determine the robustness of the attack and the updated adversarial noise; repeat the following attack steps until the number of repetitions reaches a second number of times: superimpose the updated adversarial noise onto the sample image to obtain an updated noise image; perform a first number of random purifications on the updated noise image through the purification model, and then determine a first number of classification results through a classifier; determine the loss value and gradient in each random purification process based on the difference between each classification result and the actual category of the sample image; determine a representative random purification process based on the loss value and gradient in each random purification process; perform backpropagation based on the representative random purification process to determine the target gradient; attack the purification model based on the target gradient to determine the robustness of the attack and the updated adversarial noise; and determine the worst value of the obtained robustness as the robustness of the purification model.
[0146] In some optional implementations of this embodiment, the propagation unit 405 is further configured to: perform back gradient propagation based on a representative random purification process to determine an intermediate gradient; and weight the intermediate gradient based on the number of repetitions to determine the target gradient.
[0147] According to embodiments of this disclosure, this disclosure also provides an electronic device and a readable storage medium.
[0148] An electronic device includes: one or more processors; and a storage device having one or more computer programs stored thereon, wherein when the one or more computer programs are executed by the one or more processors, the one or more processors implement the method described in process 200.
[0149] A computer-readable medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the method described in process 200.
[0150] Figure 5 A schematic block diagram of an example electronic device 500 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0151] like Figure 5 As shown, device 500 includes a computing unit 501, which can perform various appropriate actions and processes based on a computer program stored in read-only memory (ROM) 502 or a computer program loaded from storage unit 508 into random access memory (RAM) 503. RAM 503 may also store various programs and data required for the operation of device 500. The computing unit 501, ROM 502, and RAM 503 are interconnected via bus 504. Input / output (I / O) interface 505 is also connected to bus 504.
[0152] Multiple components in device 500 are connected to I / O interface 505, including: input unit 506, such as keyboard, mouse, etc.; output unit 507, such as various types of monitors, speakers, etc.; storage unit 508, such as disk, optical disk, etc.; and communication unit 509, such as network card, modem, wireless transceiver, etc. Communication unit 509 allows device 500 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0153] The computing unit 501 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 501 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 501 performs the various methods and processes described above, such as road planning methods. For example, in some embodiments, the road planning method may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 508. In some embodiments, part or all of the computer program may be loaded and / or installed on device 500 via ROM 502 and / or communication unit 509. When the computer program is loaded into RAM 503 and executed by the computing unit 501, one or more steps of the road planning method described above may be performed. Alternatively, in other embodiments, the computing unit 501 may be configured to perform the road planning method by any other suitable means (e.g., by means of firmware).
[0154] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0155] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0156] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0157] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0158] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0159] Computer systems can include clients and servers. Clients and servers are generally geographically separated and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be servers in distributed systems or servers incorporating blockchain technology. Servers can also be cloud servers, or intelligent cloud computing servers or intelligent cloud hosts with artificial intelligence technology.
[0160] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.
[0161] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. A method for evaluating the robustness of a diffusion-based cleanup model, comprising: The noisy image is randomly purified a first number of times using the purification model to obtain a first number of purification results, wherein the noisy image is obtained by superimposing anti-noise onto the sample image; The first number of purification results are used to determine the first number of classification results by a classifier; The loss value and gradient for each random purification process are determined based on the difference between each classification result and the actual category of the sample image. A representative stochastic cleanup process is determined based on the loss value and gradient in each stochastic cleanup process; Backward gradient propagation is performed based on the representative stochastic purification process to determine the target gradient; The robustness of the purification model is determined by attacking the purification model based on the target gradient.
2. The method according to claim 1, wherein, The process of determining a representative random cleanup process based on the loss value and gradient in each random cleanup process includes: Based on the loss value and gradient in each random cleanup process, the shared adversarial noise that can have the greatest impact on the random cleanup process is identified. A representative random cleanup process is determined based on the shared adversarial noise.
3. The method according to claim 2, wherein, The random cleanup process for determining representativeness based on the shared adversarial noise includes: Estimate the losses caused by the shared adversarial noise in each random cleanup process; The loss caused by each random cleanup process is mapped to the probability that a shared vulnerability exists in the random cleanup process; A representative random purification process is determined by maximizing the probability.
4. The method according to claim 1, wherein, The backpropagation of the target gradient based on the representative stochastic purification process includes: The probability distribution of misclassification is determined by maximizing the loss value in each random cleanup process; The gradients in each random purification process are weighted and summed based on the probability distribution to determine the aggregate gradient; Based on the representative stochastic purification process and the aggregated gradient, backgradient propagation is performed to determine the target gradient.
5. The method according to claim 4, wherein, The backpropagation of the target gradient based on the representative stochastic purification process and the aggregated gradient includes: The noisy image is purified through a representative random purification process to obtain the purification result; The inner product of the purification result and the aggregation gradient is backpropagated to determine the target gradient.
6. The method according to claim 1, wherein, The attack on the purification model based on the target gradient to determine the robustness of the purification model includes: The purification model is attacked based on the target gradient to determine the robustness of the attack and the updated adversarial noise. Repeat the following attack steps until the number of repetitions reaches a second number: superimpose the updated adversarial noise onto the sample image to obtain an updated noisy image; perform a first number of random purifications on the updated noisy image using the purification model, and then determine a first number of classification results using a classifier; determine the loss value and gradient in each random purification process based on the difference between each classification result and the actual category of the sample image; determine a representative random purification process based on the loss value and gradient in each random purification process; perform backpropagation based on the representative random purification process to determine the target gradient; attack the purification model based on the target gradient to determine the robustness of this attack and the updated adversarial noise; The worst value of the obtained robustness is determined as the robustness of the purification model.
7. The method according to claim 6, wherein, The backpropagation of the target gradient based on the representative stochastic purification process includes: Backward gradient propagation is performed based on the representative random purification process to determine the intermediate gradient; The intermediate gradient is weighted based on the number of repetitions to determine the target gradient.
8. An apparatus for evaluating the robustness of a diffusion-based cleanup model, comprising: The purification unit is configured to perform a first number of random purifications on the noisy image through the purification model to obtain a first number of purification results, wherein the noisy image is obtained by superimposing anti-noise onto the sample image; The classification unit is configured to determine a first number of classification results by passing the first number of purification results through a classifier; The computation unit is configured to determine the loss value and gradient for each random purification process based on the difference between each classification result and the actual category of the sample image. The selection unit is configured to determine a representative random cleanup process based on the loss value and gradient in each random cleanup process; The propagation unit is configured to perform backgrad propagation based on the representative random purification process to determine the target gradient; An attack unit is configured to attack the purification model based on the target gradient to determine the robustness of the purification model.
9. An electronic device, comprising: One or more processors; Storage device, on which one or more computer programs are stored, When the one or more computer programs are executed by the one or more processors, the one or more processors implement the method as described in any one of claims 1-7.
10. A computer-readable medium having a computer program stored thereon, wherein, When the computer program is executed by a processor, it implements the method as described in any one of claims 1-7.