Image enhancement method and system for hidden backdoor attack based on de-noising model

By constructing a denoising diffusion model in the frequency domain and adaptively adjusting the parameters, the problem of misjudgment of artifacts in medical images by the diffusion model is solved, achieving higher accuracy in disease diagnosis and model security.

CN120634890APending Publication Date: 2025-09-12NAT UNIV OF DEFENSE TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510777533.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-11
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

Existing diffusion models in fields such as medical imaging are susceptible to the problem of unnatural artifacts introduced by spatial triggers, leading to misdiagnosis of disease conditions.

Method used

By conducting attacks in the frequency domain, building a denoising diffusion model, defining triggers using frequency transformation and perturbation functions, and implanting them into the diffusion model, the attack hyperparameters are optimized through adaptive noise scheduling to achieve a hidden backdoor attack.

Benefits of technology

The accuracy of patient condition judgment is improved, the trade-off between generation quality and concealment is optimized, the occurrence of artifacts is reduced, and the safety of the model is ensured.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120634890A_ABST
    Figure CN120634890A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of security, and particularly relates to an image enhancement method and system for hidden backdoor attack based on a denoising model, and the method comprises the steps: obtaining an image frequency feature according to a training noise image, and obtaining a disturbance feature through a disturbance function; obtaining a spatial domain signal according to the disturbance characteristic, and defining a trigger according to the spatial domain signal and the image frequency characteristic; implanting a trigger into the diffusion model to obtain a modified diffusion model, a forward process and a posterior process; obtaining clean loss and poisoning loss according to the forward process, the posterior process and the loss function, and obtaining total loss according to the clean loss and the poisoning loss; adjusting parameters of the disturbance function according to the total loss to obtain an adjustment diffusion model; and inputting the actual noise image into the adjustment diffusion model to obtain an enhanced image. From the perspective of a frequency domain, a diffusion model can output a desired feature image through a backdoor attack mode, texture features of required features are enhanced directionally, and the accuracy of judging the condition of a patient is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of security technology, and in particular relates to an image enhancement method and system for hidden backdoor attacks based on a denoising model. Background Art

[0002] In recent years, diffusion models have rapidly become one of the core technologies in the field of image synthesis and restoration due to their powerful generative capabilities and stable training process. The core idea is to gradually generate high-quality images by simulating the diffusion process (forward denoising) and the inverse process (reverse denoising) of data distribution over time. Compared with traditional generative adversarial networks (GANs), diffusion models perform better in detail fidelity, pattern coverage, and training stability, and have been widely used in scenarios such as image restoration, super-resolution reconstruction, and style transfer. However, with the deepening of technical applications, the "black box" characteristics of diffusion models and their limitations in refined control and domain adaptability have gradually become apparent, especially in medical imaging, where the reliability and interpretability of the results are extremely high.

[0003] In the related art, the connection between diffusion models and autoregressive processes in the frequency domain has been revealed, emphasizing the importance of spectral analysis for understanding how different frequency components emerge in the generative process. Although diffusion models have achieved remarkable success, they also inherit the known security vulnerabilities of discriminative models, especially backdoor attacks. In traditional backdoor attacks, deep networks are trained to respond to a secret trigger (usually a small patch or pattern in the input) and produce outputs chosen by the attacker, while behaving normally on clean inputs. Most existing backdoor attacks embed such triggers in the spatial domain of the image (e.g., printing a pixel patch on the image) [7]. However, these spatial triggers often introduce unnatural artifacts.

[0004] Regarding the above-mentioned related technologies, spatial triggers often introduce unnatural artifacts, which may be misjudged as real objects or features, or even misjudged as lesions in the medical field, thereby making incorrect judgments about the patient's condition. Summary of the Invention

[0005] The technical problem to be solved by the present invention is to provide an image enhancement method and system for covert backdoor attacks based on a denoising model. The method attacks from the frequency domain, effectively deceives the diffusion model during training and inference, and has higher concealment, providing a theoretical reference for developing a dual-domain collaborative defense mechanism to ensure model security.

[0006] An image enhancement method for hidden backdoor attacks based on a denoising model, comprising:

[0007] Constructing a diffusion model, wherein the diffusion model includes a denoised diffusion probability model or a denoised diffusion implicit model;

[0008] Acquire a training noise image, where the training noise image is a lesion image;

[0009] Performing frequency transformation on the training noise image to obtain image frequency features;

[0010] Set the perturbation function;

[0011] Obtaining a disturbance feature according to the image frequency feature and the disturbance function;

[0012] Obtaining a spatial domain signal according to the disturbance characteristics, and defining a trigger according to the spatial domain signal and an image frequency characteristic;

[0013] implanting a trigger into the diffusion model to obtain a modified diffusion model;

[0014] According to the modified diffusion model, a forward process of processing a training noisy image into a clean image and a posterior process of processing a clean image into a training noisy image are obtained;

[0015] According to the forward process, the posterior process and the loss function, a clean loss and a poisoning loss are obtained respectively, and a total loss is obtained according to the clean loss and the poisoning loss;

[0016] Adjusting parameters of the disturbance function according to the total loss to obtain an adjusted diffusion model;

[0017] The actual noise image is input into the adjusted diffusion model to obtain an enhanced image.

[0018] Optionally, performing frequency transformation on the training noise image to obtain image frequency features includes:

[0019] Performing Fourier transform on the training noise image to obtain image frequency features;

[0020] The Fourier transform process is:

[0021]

[0022] Where M and N are the image sizes, x and y are the pixel coordinates of the image in the spatial domain, u and v are the coordinates in the frequency domain, corresponding to the frequency components in the horizontal x direction and vertical y direction, respectively, and f represents the pixel value of the corresponding coordinate of the image;

[0023] or performing a two-dimensional discrete cosine transform on the training noise image to obtain image frequency features;

[0024] The two-dimensional discrete cosine transform process is:

[0025]

[0026] Optionally, the perturbation function is expressed as:

[0027]

[0028] in, is the indicator function of frequency support, Ω={(u,v):u1≤u≤u2,v1≤v≤v2}, w is the coverage strength, M(u,v) is the normalization operation, and φ(u,v) is the perturbation function.

[0029] Optionally, obtaining a disturbance feature according to the image frequency feature and the disturbance function includes:

[0030] Multiplying the image frequency feature and the disturbance function to obtain a disturbance feature;

[0031] The disturbance characteristic is expressed as:

[0032]

[0033] in, represents the frequency characteristics of the training noise image, and φ(u,v) is the perturbation function.

[0034] Optionally, obtaining a spatial domain signal according to the disturbance feature, and defining a trigger according to the spatial domain signal and an image frequency feature includes:

[0035] Obtain spatial domain signal transformation formula;

[0036] Inputting the disturbance feature into the spatial domain signal conversion formula to obtain a spatial domain signal;

[0037] The spatial domain signal conversion formula is:

[0038]

[0039] in, represents a two-dimensional frequency transform, represents its inverse transform, is the disturbance characteristic;

[0040] The trigger is represented as:

[0041]

[0042] in, is the spatial domain signal of the disturbance feature, x T is the spatial domain signal of the original feature.

[0043] Optionally, obtaining, according to the modified diffusion model, a forward process of processing a training noisy image into a clean image and a posterior process of processing a training clean image into a noisy image includes:

[0044] The forward process is expressed as:

[0045]

[0046] in, β t is variance scheduling, is the spatial domain signal;

[0047] The posterior process is:

[0048]

[0049] Among them, x0 is a clean image, Noisy image, is the noise image at the previous moment, t is the time, I is the unit matrix, The mean vector of the posterior process.

[0050] Optionally, obtaining a clean loss and a poisoned loss respectively according to the forward process, the posterior process, and the loss function, and obtaining a total loss according to the clean loss and the poisoned loss includes:

[0051] According to the posterior process, obtaining a posterior mean;

[0052] The posterior mean is expressed as:

[0053]

[0054] in, g ω is a trigger, x0 is a clean image, is the noise image, α t =1-β t , β t is variance scheduling,

[0055]

[0056] The poisoning loss is expressed as:

[0057]

[0058] in, is the expected value, ∈ is the Gaussian noise actually added, ∈ θ is the noise predicted by the model, Losses due to poisoning;

[0059] The clean loss is expressed as:

[0060]

[0061] in, For clean loss;

[0062] The total loss is expressed as:

[0063]

[0064] Among them, D p For poisoned samples, D c is a clean sample and 1 is the indicator function.

[0065] A hidden backdoor attack system based on a denoising model, comprising:

[0066] A construction module, configured to construct a diffusion model, wherein the diffusion model includes a denoised diffusion probability model or a denoised diffusion implicit model;

[0067] A first acquisition module is used to acquire a training noise image;

[0068] a conversion module, configured to perform frequency conversion on the training noise image to obtain image frequency features;

[0069] A setting module is used to set a perturbation function, wherein the training noise image is a lesion image;

[0070] A disturbance module, configured to obtain a disturbance feature according to the image frequency feature and the disturbance function;

[0071] a definition module, configured to obtain a spatial domain signal according to the disturbance characteristics, and define a trigger according to the spatial domain signal and an image frequency characteristic;

[0072] An implantation module, configured to implant a trigger into the diffusion model to obtain a modified diffusion model;

[0073] a modification module, configured to obtain, based on the modified diffusion model, a forward process for processing a training noisy image into a clean image and a posterior process for processing a clean image into a training noisy image;

[0074] A calculation module, configured to obtain a clean loss and a poisoned loss according to the forward process, the posterior process, and the loss function, and obtain a total loss according to the clean loss and the poisoned loss;

[0075] an adjustment module, adjusting parameters of a disturbance function according to the total loss to obtain an adjusted diffusion model;

[0076] The output module is used to input the actual noise image into the adjusted diffusion model to obtain an enhanced image.

[0077] The content of the present invention includes

[0078] A terminal device includes a memory and a processor. The memory stores a computer program that can be run on the processor. When the processor loads and executes the computer program, an image enhancement method for hidden backdoor attacks based on a denoising model is adopted.

[0079] A computer-readable storage medium stores a computer program. When the computer program is loaded and executed by a processor, an image enhancement method for a hidden backdoor attack based on a denoising model is adopted.

[0080] The beneficial effects of the present invention are:

[0081] 1. The training noise image is transformed into a frequency domain to obtain image frequency features. A perturbation function is then defined to perturb the image frequency features into perturbation features. Based on the perturbation features, a spatial domain signal is obtained. Triggers are defined based on the spatial domain signal and the image frequency features. These triggers are then embedded into a diffusion model to obtain a modified diffusion model. Based on the modified diffusion model, a forward process for converting the training noise image into a clean image and a posterior process for converting the clean image into a training noise image are obtained. Based on the forward and posterior processes and a loss function, a clean loss and a poisoning loss are obtained, respectively. A total loss is then derived from the clean and poisoning losses. The parameters of the perturbation function are adjusted based on the total loss to obtain an adjusted diffusion model. The actual noise image is then input into the adjusted diffusion model to obtain an enhanced image. By using a backdoor attack from a frequency domain perspective, the diffusion model is enabled to output the desired feature image, specifically, the texture features required for targeted enhancement, thereby improving the accuracy of patient condition diagnosis.

[0082] 2. We propose an adaptive noise scheduling method that dynamically fine-tunes attack hyperparameters based on the characteristics of the generative model. This method optimizes the trade-off between generation quality and stealth, achieving better performance than traditional spatial triggers. BRIEF DESCRIPTION OF THE DRAWINGS

[0083] Figure 1 Comparison between the diffusion process of the present invention and the proposed backdoor attack strategy;

[0084] Figure 2 The specific performance of different attack methods of the present invention on the CelebA / CIFAR dataset is verified;

[0085] Figure 3 is the attack success rate of different frequency methods of the present invention;

[0086] Figure 4This is a performance comparison of the present invention on the CIFAR-10 and CelebA datasets under different attack scenarios;

[0087] Figure 5 This is a schematic diagram of the effect of the adaptive noise scheduling of the present invention;

[0088] Figure 6 Schematic diagram of the trade-off between attack success rate (ASR) and concealment of the present invention;

[0089] Figure 7 Schematic diagram of the trade-off between attack success rate, peak signal-to-noise ratio, structural similarity index, and Fréchet inception distance (FID) for different frequency attack methods of the present invention. DETAILED DESCRIPTION

[0090] An image enhancement method for hidden backdoor attacks based on a denoising model, the present invention comprises:

[0091] S1. Construct a diffusion model, where the diffusion model includes a denoising diffusion probability model or a denoising diffusion implicit model.

[0092] Specifically, obtaining and using diffusion models involves three key steps.

[0093] (1) Diffusion process: Define a diffusion process that gradually transforms the data distribution q(x) into a known distribution r(x) over T time steps. The data distribution refers to the natural probability distribution of the original data, and the essence of the diffusion process is to transform this complex distribution into a known simple prior distribution (such as a Gaussian distribution) by gradually adding noise.

[0094] (2) Training: Optimize the parameter θ so that the learned generation (reverse) process approximates the true reverse diffusion process, i.e.

[0095]

[0096] Among them, x t is the noise image, x t-1 is the noise image at the previous moment, represents the normal distribution, μθ is the core parameter of the reverse process of the diffusion model, which is responsible for guiding the denoising direction of each step, and p θ It is the core of the reverse process in the diffusion model, which simulates the process of data recovery through a parameterized probability distribution (learned by a neural network).

[0097] (3) Sampling: Through the reverse process from learning The image is generated by iterative sampling from t=T to t=1.

[0098] Denoising Diffusion Probabilistic Model (DDPM): DDPM sets the target distribution as And the forward diffusion process is modeled as a Markov chain:

[0099]

[0100] Among them, α t =1-β t and is a predefined variance schedule. Definition Then, it can be proved that given an initial sample x0~q(x), a time step t uniformly drawn from {1,…,T}, and The forward process can be expressed as:

[0101]

[0102] By minimizing losses

[0103]

[0104] Among them, ∈ θ is the predicted noise, ∈ is the real noise.

[0105] The trained model produces the reverse (generation) process

[0106]

[0107] in,

[0108]

[0109] and,

[0110] Given Model from Where t=T is iterative sampling until t=1, and finally x0 is obtained, and I is the unit matrix.

[0111] Denoising Diffusion Implicit Model (DDIM): DDIM uses the same target distribution as the Denoising Diffusion Probabilistic Model (DDPM) and forward diffusion process, but adopts an improved reverse process. Under the equivalent training objective, the generation process in DDIM is defined as:

[0112]

[0113] where the inverse mean is given by

[0114]

[0115] in, is the reverse mean.

[0116] And the variance is parameterized as

[0117]

[0118] Compared with DDPM, DDIM adopts stride sampling scheduling to speed up the reverse process.

[0119] S2. Acquire a training noise image, where the training noise image is a lesion image.

[0120] S3. Perform frequency transformation on the training noise image to obtain image frequency features.

[0121] Perform frequency transformation on the training noise image to obtain the image frequency features including:

[0122] Specifically, frequency domain transform converts the image from the spatial domain into a representation that can clearly show its frequency components, including Fourier transform and two-dimensional discrete cosine transform.

[0123] Perform Fourier transform on the training noise image to obtain the image frequency features.

[0124] The Fourier transform process is:

[0125]

[0126] Where M and N are the image sizes, x and y are the pixel coordinates of the image in the spatial domain, u and v are the coordinates in the frequency domain, corresponding to the horizontal (x-direction) and vertical (y-direction) frequency components, respectively, and f represents the pixel value of the corresponding coordinate of the image, u = 0, 1, ..., M-1, v = 0, 1, ..., N-1.

[0127] Alternatively, the training noise image can be subjected to a two-dimensional discrete cosine transform to obtain the image frequency features.

[0128] The two-dimensional discrete cosine transform process is:

[0129]

[0130] For u = 0, ..., M-1 and v = 0, ..., N-1. The two-dimensional discrete wavelet transform (2DDWT) decomposes f(x, y) into multiple subbands through wavelet filtering and downsampling, typically producing a low-frequency approximation subband (LL) and three high-frequency detail subbands (LH, HL, and HH). Each transform provides a unique perspective: the fast Fourier transform (FFT) emphasizes global frequency components, the discrete cosine transform (DCT) provides energy compression, and the discrete wavelet transform (DWT) captures local multi-resolution details.

[0131] S4. Set the perturbation function.

[0132] Specifically, the backdoor trigger is injected into the initial noise of the diffusion model through adaptive frequency domain perturbation. represents a two-dimensional frequency transform (e.g., a two-dimensional fast Fourier transform (2DFFT), a discrete cosine transform (DCT), or a discrete wavelet transform (DWT)), Denotes its inverse transform. Given a frequency representation The initial noise image Define the perturbation function.

[0133] The perturbation function is expressed as:

[0134]

[0135] in, is the indicator function of frequency support, Ω={(u,v):u1≤u≤u2,v1≤v≤v2}, w is the coverage strength, M(u,v) is the normalization operation, φ(u,v) is the perturbation function, and Ω is the support set.

[0136] S5. Obtain disturbance features according to the image frequency features and the disturbance function.

[0137] According to the image frequency characteristics and the perturbation function, the perturbation features include:

[0138] The perturbation feature is obtained by multiplying the image frequency feature and the perturbation function.

[0139] The disturbance characteristics are expressed as:

[0140]

[0141] in, represents the frequency characteristics of the training noise image, and φ(u,v) is the perturbation function.

[0142] S6. Obtain a spatial domain signal according to the disturbance characteristics, and define a trigger according to the spatial domain signal and the image frequency characteristics.

[0143] According to the disturbance characteristics, the spatial domain signal is obtained, and the trigger is defined according to the spatial domain signal and the image frequency characteristics, including:

[0144] Get the spatial domain signal transformation formula.

[0145] The disturbance characteristics are input into the spatial domain signal conversion formula to obtain the spatial domain signal.

[0146] The spatial domain signal conversion formula is:

[0147]

[0148] in, represents a two-dimensional frequency transform, represents its inverse transform, is the disturbance characteristic.

[0149] A trigger is represented as:

[0150]

[0151] in, is the spatial domain signal of the disturbance feature, x T is the spatial domain signal of the original feature.

[0152] if and are independent of each other, then the final noise distribution is:

[0153]

[0154] S7. Implant a trigger into the diffusion model to obtain a modified diffusion model.

[0155] S8. According to the modified diffusion model, a forward process of processing the training noisy image into a clean image and a posterior process of processing the clean image into a training noisy image are obtained.

[0156] According to the modified diffusion model, the forward process of processing the training noisy image into a clean image and the posterior process of processing the training clean image into a noisy image are obtained, which include:

[0157] The forward process is expressed as:

[0158]

[0159] in, α t =1-β t , β t is variance scheduling, is the spatial domain signal.

[0160] The posterior process is:

[0161]

[0162] Among them, x0 is a clean image, Noisy image, is the noise image at the previous moment, t is the time, I is the unit matrix, The mean vector of the posterior process.

[0163] S9. According to the forward process, the posterior process and the loss function, the clean loss and the poisoned loss are obtained respectively, and the total loss is obtained based on the clean loss and the poisoned loss.

[0164] According to the forward process, the posterior process and the loss function, the clean loss and the poisoning loss are obtained respectively, and the total loss is obtained according to the clean loss and the poisoning loss, including:

[0165] According to the posterior process, the posterior mean is obtained.

[0166] The posterior mean is expressed as:

[0167]

[0168] in, g ω is a trigger, x0 is a clean image, is the noise image, α t =1-β t , β t is variance scheduling,

[0169] Poisoning loss is expressed as:

[0170]

[0171] in, is the expected value, ∈ is the Gaussian noise actually added, ∈ θ is the noise predicted by the model, Poisoning loss.

[0172] The clean loss is expressed as:

[0173]

[0174] in, For clean loss.

[0175] The total loss is expressed as:

[0176]

[0177] Among them, D p For poisoned samples, D c is a clean sample, and 1 is an indicator function, which selectively activates the corresponding loss term according to the type of sample x (clean sample or poisoned sample).

[0178] S10. Adjust the parameters of the disturbance function according to the total loss to obtain an adjusted diffusion model.

[0179] Specifically, dynamic hyperparameter adjustment involves dynamically adjusting the trigger size (i.e., support set Ω) and coverage strength w during training based on the calculated Fréchet Inception Distance (FID) score. Specifically, we implement an adaptive FID-based optimization mechanism: the FID is periodically evaluated on the validation set. If the score deviates from a predefined threshold, Ω and w are adjusted to maintain the optimal balance between attack stealth and image quality. This ensures that the backdoor perturbations are both visually and statistically indistinguishable while maintaining the generative performance of the diffusion model. After backpropagating the diffusion model parameters through the loss function, the diffusion model is adjusted based on FID evaluation. The stopping condition is when the attack accuracy reaches a set value, i.e., when the number of clean images output in a dataset reaches a set value. The adaptive mechanism integrates a feedback loop in which FID fluctuations guide the optimization of the spectral trigger position. If the FID increases significantly, indicating a significant degradation in image quality, the trigger strength w is reduced and Ω is repositioned to a less sensitive frequency band. Conversely, if the backdoor attack success rate (ASR) drops below a set threshold while the FID remains low, w is gradually increased to enhance the texture of the lesion without producing noticeable artifacts, thereby improving the accuracy of the patient's condition. This self-regulatory method dynamically fine-tunes the perturbation parameters throughout the training process, ensuring that the model achieves both high ASR and low detectability.

[0180] S11. Input the actual noise image into the adjusted diffusion model to obtain an enhanced image.

[0181] Specifically, after the enhanced image passes the backdoor attack, the enhanced image can enhance the texture contrast of certain types of lesions and avoid misjudgment due to artifacts.

[0182] Specific examples:

[0183] Datasets, Models, and Implementation Details. We evaluate the performance of DIFF on two benchmark vision datasets: CIFAR-10 (32×32) and CelebA (64×64). Following the standard implementation of the Denoising Diffusion Probabilistic Model (DDPM) and the Denoising Diffusion Implicit Model (DDIM), we train the diffusion model from scratch using a batch size of 128 for 200 epochs. We use the Adam optimizer (β1=0.9,β2=0.999) and a learning rate of 5×10 -4 The model is optimized and linear noise scheduling is adopted, where β start =0.0001, β end = 0.02 over 1000 time steps. For evaluation, we generate 50k benign and 10k poisoned samples using η = 0.0 and S = 100 acceleration steps in the DDIM configuration.

[0184] Attack configuration: For attack scenarios, the target of the in-domain device-to-device (In-D2D) attack is the generation of a specific category, redirecting the output of the CIFAR-10 dataset to the horse category through frequency domain perturbations, and converting CelebA samples into faces with heavy makeup and smiling attributes. The inter-domain device-to-device (Out-D2D) attack maps the input to a cross-domain target, such as MNIST handwritten digits (such as the number 8), through coordinated spatial-spectral triggers. The device-to-image (D2I) attack directly hijacks the generation process to synthesize predefined images (such as cartoon characters such as Mickey Mouse) by combining adaptive noise scheduling and spectral operations. Figure 1 All attacks prioritize stealth by restricting the perturbations to critical spectral bands while maintaining the generative fidelity of the model.

[0185] Evaluation Metrics: We select four widely used metrics to evaluate the attack performance, namely attack success rate (ASR, the proportion of generated images that the classification model identifies as the target category), peak signal-to-noise ratio (PSNR)

[25] , structural similarity index (SSIM)

[14] , and Fréchet-Inception distance (FID). ASR is used to measure the accuracy of generated images in terms of the target category. We use PSNR and SSIM to quantify the stealth of poisoned images in the spatial domain. PSNR measures the pixel-level fidelity between generated and clean images through logarithmic error analysis (higher values ​​indicate less perceptible), while SSIM evaluates perceptual similarity by comparing brightness, contrast, and structural patterns (closer to 1 indicates stronger stealth). Although PSNR ignores the sensitivity of human vision to structural distortion, it provides a complementary perspective when combined with SSIM: PSNR ensures that pixel deviation is minimized and SSIM ensures that semantic features are preserved, jointly verifying the undetectability of the attack in both numerical and perceptual dimensions. In addition, we introduce FID to evaluate the overall quality of generated images by measuring the distribution difference between generated and clean samples in the feature space of the pre-trained Inception network. Lower FID values ​​indicate higher fidelity and realism, ensuring that the backdoor model produces high-quality output while maintaining the stealth of the attack. Through the methods of this application, the attack was successfully implemented, enabling the diffusion model to enhance the features we need in a targeted manner.

[0186] Experimental results: Figure 2 In

[15] , we compared frequency-domain (Fast Fourier Transform high frequency, Discrete Cosine Transform low frequency) and spatial-domain attacks on the CelebA dataset. The "Original" column shows the unaltered image, while subsequent columns show the attack results, with the difference magnified fivefold. Frequency-domain attacks (FFT, DCT) demonstrate a clear advantage over spatial attacks.

[0187] Because the perturbations are confined to the high-frequency region, the FFT attack is almost imperceptible, achieving the best stealth (PSNR=35.7, SSIM=0.94). The DCT attack is more noticeable due to the structured low-frequency perturbations, but achieves a higher attack success rate (0.94) at the expense of lower stealth (PSNR=29.4, SSIM=0.88). In contrast, the spatial attack introduces highly localized perturbations that are less resistant to diffusion denoising, making the attack less robust.

[0188] The diffusion model naturally mitigates local pixel variations, which explains why spatial attacks fail more often. Figure 3 In

[15] , frequency domain attacks maintain a higher attack success rate compared to spatial attacks, confirming their better persistence. Spectrum visualization further illustrates this difference.

[0189] After the FFT high-frequency attack, the increase in energy in the high-frequency region indicates the presence of hidden perturbations, which are still effective during denoising. In contrast, the DCT-based low-frequency perturbations change the global structure, resulting in higher robustness but also easier to detect (ASR = 0.95, PSNR = 29.4).

[0190] Figure 4 The performance comparison on CIFAR-10 and CelebA datasets under different attack scenarios is shown. DIFF represents the best method in different frequency domains.

[0191] Ablation studies:

[0192] The effect of DIFF:

[0193] Experimental results show that the frequency-domain backdoor attack method DIFF achieves superior attack success rate (ASR) and stealth (peak signal-to-noise ratio, PSNR) compared to TroiDiff. In all attack scenarios, DIFF consistently achieves a higher attack success rate, reaching 98.87% on the CIFAR-10 dataset and 97.92% on the CelebA dataset. Notably, DIFF* exhibits the strongest attack effect in the Out-D2D task, confirming its robustness under different data distributions. In addition, DIFF achieves a significantly higher peak signal-to-noise ratio than TroiDiff (e.g., 35.7 vs. 22.8 on the CIFAR-10 dataset), indicating that it is better resistant to visual degradation, more stealthy, and more difficult to detect.

[0194] Figure 5 and Figure 6This further highlights the effectiveness of frequency-domain attacks. As the attack strength increases, the attack success rates of all methods significantly improve, confirming that stronger perturbations are more reliably triggering the backdoor. Low-frequency attacks (Discrete Cosine Transform Low Frequency, DCTLow, Discrete Wavelet Transform Low Frequency, DWTLow) achieve the highest attack success rates at all attack strengths, outperforming high-frequency attacks (Fast Fourier Transform High Frequency, FFTHigh, Discrete Cosine Transform High Frequency, DCTHigh) at lower strengths. For example, at w = 0.3, Discrete Cosine Transform Low Frequency (0.50) and Discrete Wavelet Transform High Frequency (0.48) surpass Fast Fourier Transform High Frequency (0.45) and Fast Fourier Transform Low Frequency (0.38), demonstrating that modifying low-frequency components can produce more effective backdoor triggering. At w = 0.9, the attack success rates of most attacks approach perfect (close to 1.0), demonstrating the high effectiveness of frequency-domain perturbations at the optimal attack strength. The impact of frequency support size (Ω) on attack performance is also significant. Larger Ω allows perturbations to propagate across a wider range of frequency components, improving attack effectiveness but potentially reducing stealth. Conversely, smaller Ω limits perturbations to a specific frequency band, preserving stealth but requiring a higher w to maintain effectiveness. This is consistent with observations that adaptive frequency attacks (Fast Fourier Transform Adaptive, FFTAdapt, Discrete Cosine Transform Adaptive, DCTAdapt, Discrete Wavelet Transform Adaptive, DWTAdapt) dynamically adjust Ω and w to optimize attack success rate and imperceptibility.

[0195] Effects of adaptive noise scheduling:

[0196] Our FID-based optimization strategy dynamically adjusts frequency support (Ω) and perturbation strength (w) to optimize the attack success rate-peak signal-to-noise ratio (PSNR) trade-off. Compared to static attacks, adaptive methods improve stealth without compromising attack performance. For example, FFT high-frequency adaptation (FFTHigh-Adapt) improves the PSNR from 32.3 to 35.7 while maintaining the attack success rate at 0.93, and DCT high-frequency adaptation (DCTHigh-Adapt) improves the PSNR from 31.8 to 34.2, confirming that adaptive perturbations reduce perceptibility while maintaining the attack success rate.

[0197] Adaptive low-frequency attacks (FFT low-frequency adaptation, DCT low-frequency adaptation, and DWT low-frequency adaptation) achieve a high attack success rate (>0.92) while significantly improving the peak signal-to-noise ratio and structural similarity index (SSIM). DWT high-frequency adaptation outperforms its static counterpart, demonstrating that our approach balances stealth and attack success rate. These results confirm that adaptive scheduling is a key advancement over static frequency-based backdoor attacks.

[0198] Figure 6 The attack success rate-peak signal-to-noise ratio trade-off for different attack methods is shown. A clear inverse relationship can be observed: as the attack success rate increases, the peak signal-to-noise ratio decreases, which means that stronger backdoor triggers introduce more perceptible perturbations. The FFT high-frequency attack achieves the highest peak signal-to-noise ratio (35.7dB) while maintaining an attack success rate of 0.93, making it both highly covert and effective. In contrast, the FFT low-frequency attack and the DCT low-frequency attack achieve a slightly higher attack success rate (0.95), but at the cost of a lower peak signal-to-noise ratio (28.5dB, 29.4dB), indicating that the perturbation is more noticeable. This is consistent with our hypothesis that low-frequency perturbations are easier to detect due to their impact on the global image structure.

[0199] Figure 7 The trade-offs of attack success rate, peak signal-to-noise ratio, structural similarity index and Fréchet inception distance (FID) for different frequency attack methods are shown, and band-pass filter (BPM).

[0200] Based on the previous discussion and experimental evidence, the impact of different bandpass frequency domains varies across tasks. As shown in the figure, high-frequency attacks maximize the attack success rate but significantly degrade image quality (peak signal-to-noise ratio and structural similarity index). In contrast, low-frequency attacks achieve higher peak signal-to-noise ratio and structural similarity index, ensuring better concealment at the expense of slightly lower attack success rate. The bandpass approach provides a balanced trade-off between attack success rate and concealment.

[0201] A hidden backdoor attack system based on a denoising model, comprising:

[0202] A construction module, configured to construct a diffusion model, wherein the diffusion model includes a denoised diffusion probability model or a denoised diffusion implicit model;

[0203] A first acquisition module is used to acquire a training noise image;

[0204] a conversion module, configured to perform frequency conversion on the training noise image to obtain image frequency features;

[0205] Setting module, used to set the perturbation function;

[0206] A disturbance module, configured to obtain a disturbance feature according to the image frequency feature and the disturbance function;

[0207] a definition module, configured to obtain a spatial domain signal according to the disturbance characteristics, and define a trigger according to the spatial domain signal and an image frequency characteristic;

[0208] An implantation module, configured to implant a trigger into the diffusion model to obtain a modified diffusion model;

[0209] a modification module, configured to obtain, based on the modified diffusion model, a forward process for processing a training noisy image into a clean image and a posterior process for processing a clean image into a training noisy image;

[0210] A calculation module, configured to obtain a clean loss and a poisoned loss according to the forward process, the posterior process, and the loss function, and obtain a total loss according to the clean loss and the poisoned loss;

[0211] an adjustment module, adjusting parameters of a disturbance function according to the total loss to obtain an adjusted diffusion model;

[0212] The output module is used to input the actual noise image into the adjusted diffusion model to obtain an enhanced image.

[0213] An embodiment of the present application also discloses a terminal device, including a memory and a processor. The memory stores a computer program that can be run on the processor. When the processor loads and executes the computer program, it adopts an image enhancement method for hidden backdoor attacks based on a denoising model.

[0214] Among them, the terminal device can be a computer device such as a desktop computer, a laptop computer or a cloud server, and the terminal device includes but is not limited to a processor and a memory. For example, the terminal device can also include input and output devices, network access devices and buses, etc.

[0215] Among them, the processor can adopt a central processing unit (CPU). Of course, according to actual usage, other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. can also be adopted. The general-purpose processor can adopt a microprocessor or any conventional processor, etc., and this application does not impose any restrictions on this.

[0216] Among them, the memory can be an internal storage unit of the terminal device, such as the hard disk or memory of the terminal device, or it can be an external storage device of the terminal device, such as a plug-in hard disk, smart memory card (SMC), secure digital card (SD) or flash memory card (FC) equipped on the terminal device, etc., and the memory can also be a combination of the internal storage unit and the external storage device of the terminal device. The memory is used to store computer programs and other programs and data required by the terminal device. The memory can also be used to temporarily store data that has been output or is to be output. This application does not impose any restrictions on this.

[0217] Among them, through this terminal device, an image enhancement method for hidden backdoor attack based on a denoising model in the above embodiment is stored in the memory of the terminal device, and is loaded and executed on the processor of the terminal device for easy use.

[0218] An embodiment of the present application further discloses a computer-readable storage medium, and the computer-readable storage medium stores a computer program, wherein when the computer program is executed by a processor, an image enhancement method for a hidden backdoor attack based on a denoising model in the above embodiment is adopted.

[0219] Among them, the computer program can be stored in a computer-readable medium, the computer program includes computer program code, the computer program code can be in the form of source code, object code, executable file or certain middleware, etc. The computer-readable medium includes any entity or device that can carry computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal and software distribution medium, etc. It should be noted that computer-readable medium includes but is not limited to the above-mentioned components.

[0220] Among them, through this computer-readable storage medium, an image enhancement method for hidden backdoor attack based on a denoising model in the above embodiment is stored in a computer-readable storage medium, and is loaded and executed on a processor to facilitate the storage and application of the above method.

[0221] Those skilled in the art will appreciate that the discussion of any of the above embodiments is merely illustrative and is not intended to imply that the scope of protection of this application is limited to these examples. Within the context of this application, the technical features of the above embodiments or different embodiments may be combined, the steps may be implemented in any order, and there are many other variations of different aspects of one or more of the above embodiments of this application, which are not provided in detail for the sake of clarity.

[0222] The one or more embodiments of this application are intended to encompass all such substitutions, modifications, and variations that fall within the broad scope of this application. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of one or more embodiments of this application should be included in the scope of protection of this application.

Claims

1. An image enhancement method for hidden backdoor attacks based on a denoising model, characterized by: include: Constructing a diffusion model, wherein the diffusion model includes a denoised diffusion probability model or a denoised diffusion implicit model; Acquire a training noise image, where the training noise image is a lesion image; Performing frequency transformation on the training noise image to obtain image frequency features; Set the perturbation function; Obtaining a disturbance feature according to the image frequency feature and the disturbance function; Obtaining a spatial domain signal according to the disturbance characteristics, and defining a trigger according to the spatial domain signal and an image frequency characteristic; implanting a trigger into the diffusion model to obtain a modified diffusion model; According to the modified diffusion model, a forward process of processing a training noisy image into a clean image and a posterior process of processing a clean image into a training noisy image are obtained; According to the forward process, the posterior process and the loss function, a clean loss and a poisoning loss are obtained respectively, and a total loss is obtained according to the clean loss and the poisoning loss; Adjusting the parameters of the disturbance function according to the total loss to obtain an adjusted diffusion model; The actual noise image is input into the adjusted diffusion model to obtain an enhanced image.

2. The image enhancement method for hidden backdoor attacks based on a denoising model as claimed in claim 1, characterized in that: The performing frequency transformation on the training noise image to obtain the image frequency feature comprises: Performing Fourier transform on the training noise image to obtain image frequency features; The Fourier transform process is: Where M and N are the image sizes, x and y are the pixel coordinates of the image in the spatial domain, u and v are the coordinates in the frequency domain, corresponding to the frequency components in the horizontal x direction and vertical y direction, respectively, and f represents the pixel value of the corresponding coordinate of the image; or performing a two-dimensional discrete cosine transform on the training noise image to obtain image frequency features; The two-dimensional discrete cosine transform process is:

3. The image enhancement method for hidden backdoor attacks based on a denoising model as claimed in claim 1, characterized in that: The perturbation function is expressed as: in, is the indicator function of frequency support, Ω={(u,v):u1≤u≤u2,v1≤v≤v2}, w is the coverage strength, M(u,v) is the normalization operation, and φ(u,v) is the perturbation function.

4. The image enhancement method for hidden backdoor attacks based on a denoising model as claimed in claim 1, characterized in that: Obtaining the disturbance feature according to the image frequency feature and the disturbance function includes: Multiplying the image frequency feature and the disturbance function to obtain a disturbance feature; The disturbance characteristic is expressed as: in, represents the frequency characteristics of the training noise image, and φ(u,v) is the perturbation function.

5. The image enhancement method for hidden backdoor attacks based on a denoising model as claimed in claim 1, characterized in that: Obtaining a spatial domain signal according to the disturbance feature, and defining a trigger according to the spatial domain signal and an image frequency feature includes: Obtain spatial domain signal transformation formula; Inputting the disturbance feature into the spatial domain signal conversion formula to obtain a spatial domain signal; The spatial domain signal conversion formula is: in, represents a two-dimensional frequency transform, represents its inverse transform, is the disturbance characteristic; The trigger is represented as: in, is the spatial domain signal of the disturbance feature, x T is the spatial domain signal of the original feature.

6. The image enhancement method for hidden backdoor attacks based on a denoising model as claimed in claim 1, characterized in that: The modified diffusion model is used to obtain a forward process of processing a training noisy image into a clean image and a posterior process of processing a training clean image into a noisy image, including: The forward process is expressed as: in, α t =1-β t , β t is variance scheduling, is the spatial domain signal; The posterior process is: Among them, x0 is the clean image, Noisy image, is the noise image at the previous moment, t is the time, I is the unit matrix, The mean vector of the posterior process.

7. The image enhancement method for hidden backdoor attacks based on a denoising model as claimed in claim 1, characterized in that: The process of obtaining a clean loss and a poisoning loss according to the forward process, the posterior process, and the loss function, and obtaining a total loss according to the clean loss and the poisoning loss includes: According to the posterior process, obtaining a posterior mean; The posterior mean is expressed as: in, g ω is a trigger, x0 is a clean image, is the noise image, α t =1-β t , β t is variance scheduling, The poisoning loss is expressed as: in, is the expected value, ∈ is the Gaussian noise actually added, ∈ θ is the noise predicted by the model, Losses due to poisoning; The clean loss is expressed as: in, For clean loss; The total loss is expressed as: Among them, D p For poisoned samples, D c For clean samples, is the indicator function.

8. A hidden backdoor attack system based on a denoising model, characterized by: include: A construction module, configured to construct a diffusion model, wherein the diffusion model includes a denoised diffusion probability model or a denoised diffusion implicit model; A first acquisition module is used to acquire a training noise image, where the training noise image is a lesion image; a conversion module, configured to perform frequency conversion on the training noise image to obtain image frequency features; Setting module, used to set the perturbation function; A disturbance module, configured to obtain a disturbance feature according to the image frequency feature and the disturbance function; a definition module, configured to obtain a spatial domain signal according to the disturbance characteristics, and define a trigger according to the spatial domain signal and an image frequency characteristic; An implantation module, configured to implant a trigger into the diffusion model to obtain a modified diffusion model; a modification module, configured to obtain, based on the modified diffusion model, a forward process for processing a training noisy image into a clean image and a posterior process for processing a clean image into a training noisy image; A calculation module, configured to obtain a clean loss and a poisoned loss according to the forward process, the posterior process, and the loss function, and obtain a total loss according to the clean loss and the poisoned loss; an adjustment module, adjusting parameters of a disturbance function according to the total loss to obtain an adjusted diffusion model; The output module is used to input the actual noise image into the adjusted diffusion model to obtain an enhanced image.

9. A terminal device comprising a memory and a processor, characterized in that: The memory stores a computer program that can be run on the processor. When the processor loads and executes the computer program, the method according to any one of claims 1 to 7 is adopted.

10. A computer-readable storage medium storing a computer program, wherein: When the computer program is loaded and executed by a processor, the method according to any one of claims 1 to 7 is adopted.

Citation Information

Cited By

  • Distribution line abnormity monitoring and early warning method and system

    CN121459523A