Neural network backdoor attack method based on phase offset

By generating a backdoor trigger by performing frequency domain phase shift in the YCrCb color space, the problem of insufficient visibility and robustness of the backdoor attack method in the prior art is solved, and a backdoor attack with high concealment and high success rate is achieved.

CN120495720APending Publication Date: 2025-08-15JIANGSU OCEAN UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510422950.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

The existing backdoor attack methods are easily visually recognized, and are not robust enough in image preprocessing and security detection, making it difficult to take into account both concealment and attack success rate.

Method used

By converting the image from the RGB channel to the YCrCb channel, and performing two-dimensional discrete Fourier transform on the chroma channel with low visual sensitivity, selecting the high and low frequency ranges for phase offset, generating a trigger and using the LPIPS metric for perceptual filtering, forming a global trigger.

Benefits of technology

It significantly improves the concealment and attack success rate of backdoor attacks, enhances the adaptability to image transformation and detection strategies, and maintains the normal performance of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120495720A_ABST
    Figure CN120495720A_ABST
Patent Text Reader

Abstract

The invention discloses a neural network backdoor attack method based on phase deviation, and the method achieves the injection of an invisible backdoor trigger into an image through the fine adjustment of the phase information of a specific channel in a frequency domain. The method comprises the following steps: converting an image from RGB into a YCrCb color space, performing two-dimensional discrete Fourier transform on a selected channel, selecting phases of high-frequency and low-frequency components for disturbance, and then ensuring that a poisoning image is highly approximate to an original image in visual appearance through inverse transform and LPIPS perception filtering. According to the method, the implanted triggers are scattered and hidden, conventional detection means can be effectively avoided, normal classification performance of clean samples is kept, and high attack success rate and high adversarial robustness are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of image security, and in particular relates to a neural network backdoor attack method based on phase shift. Background Art

[0002] In recent years, deep neural networks (DNNs) have been widely used in fields such as computer vision, natural language processing, and biomedicine due to their superior feature extraction and pattern learning capabilities, significantly driving advances in key technologies such as medical diagnosis, autonomous driving, and financial forecasting. However, training high-performance DNN models requires extensive computing resources, massive amounts of labeled data, and sophisticated parameter optimization, resulting in high computational costs and stringent infrastructure requirements. Consequently, developers are increasingly relying on third-party computing platforms (such as cloud computing services) for training and deployment to reduce costs and technical barriers. While these platforms offer good scalability and convenience, users lack full control over the training process and data security, making them vulnerable to attackers who can tamper with data or model parameters to implant backdoor attacks during the training phase, posing potential risks.

[0003] Backdoor attacks are a subtle and emerging cybersecurity threat. Attackers silently insert malicious behavior into model training, causing the tampered model to behave almost identically to an unattacked model when processing normal input, making it difficult to detect through conventional verification methods. However, when the input contains a specific trigger designed by the attacker, the hidden backdoor is activated, and the model predicts the target output according to the attacker's pre-set target. This attack can be carried out without real-time input modification and solely based on pre-set triggers, making it extremely secretive and dangerous.

[0004] LPIPS (Learned Perceptual Image Patch Similarity) is a metric for evaluating the perceptual similarity of images. Its core idea is to use a pre-trained deep neural network to extract feature representations of images at different levels and perform distance calculations on these features to measure the similarity or difference between images. Specifically, LPIPS first inputs two images into the same or a pre-trained network with the same structure (such as VGG or AlexNet), extracting multi-level feature maps at the output of several layers of the network. It then performs L2 distance or other distance metrics on the feature maps of the corresponding layers point by point, and performs a weighted aggregation of these distances to obtain the final perceptual difference score. Compared with traditional mean square error (MSE) or structural similarity (SSIM), LPIPS is closer to the judgment results of the human visual system and has higher sensitivity and accuracy in detecting subtle perceptual differences.

[0005] In the research and practice of backdoor attacks, various image-based technical solutions have emerged. For example, Chinese Patent Publication No. CN118379534A describes a contour-based image classification backdoor attack method. This method extracts image edges and encodes contours from a target sample dataset to generate triggers, which are then added to the target image. The model is then trained with a benign sample dataset. When an image with the trigger is input into the backdoor model, the backdoor is activated and outputs an incorrect label, while maintaining normal classification for normal inputs. However, while these triggers based on explicit contours or feature stickers can enable backdoor attacks, they still have some visibility in the image's appearance and may be difficult to detect using manual or simple image detection methods. Chinese Patent Publication No. CN119378006A describes a backdoor attack method based on image steganography. This method utilizes an image steganography network and an image transformation network to generate malicious images that adapt to various spatial transformations. Backpropagation is used to continuously optimize the steganographic loss function, improving the attack success rate in deformed scenarios. This method enhances the concealment of the trigger and ensures that the trigger remains effective even after image transformations. However, since such steganographic methods often involve more complex network structure training and additional loss function design, higher requirements are placed on the implementation environment and computing costs. Chinese patent publication number CN117951690A is a random frequency domain perturbation backdoor attack method that combines high and low frequencies. The method first sets a trigger in the low-frequency component and then randomly injects perturbations into the high-frequency component, thereby achieving higher flexibility and a stable attack success rate between different network structures and data sets. This method improves the backdoor's concealment and robustness to adversarial forces in the form of frequency domain triggers. However, this solution mainly relies on random perturbations and introduces feature changes in high and low frequency areas. In certain scenarios, there may still be detectable feature patterns or statistical differences.

[0006] In addition to the methods disclosed in the aforementioned patent documents, the academic community has also proposed a variety of backdoor attack techniques using explicit decals, watermarks, and noise embedding. However, most existing solutions have the following limitations: Visible features are too obvious: Some attack triggers based on decals, contours, or special marks still have significant visual features, which can easily arouse suspicion; Insufficient adaptability to complex image transformations: After image scaling, rotation, lighting changes, or other pre-processing operations, the effectiveness of the trigger may decrease; It is difficult to balance stealth and attack success rate: Some methods sacrifice attack success rate while improving stealth, or increase trigger visibility while strengthening the attack effect; Insufficient ability to counter detection tools: With the improvement of model security detection methods (such as statistical screening, frequency domain filtering, deep feature comparison, etc.), some triggers are easily screened or defended. Summary of the Invention

[0007] Based on the above technical background and the deficiencies of the existing technology, the present invention proposes a neural network backdoor attack method based on phase shift. In order to solve the problems in the existing technology that backdoor triggers often leave recognizable traces in the visible space domain or local feature domain, lack sufficient concealment, and lack robustness in the face of image preprocessing and security detection, the frequency domain phase information of the specified channel is fine-tuned after the color space conversion, and the trigger is dispersed into the texture of the entire image by using high and low frequency fusion, making it more difficult to be detected visually. At the same time, it is also easier to cope with image transformation and various detection strategies, thereby improving the balance between concealment and attack success rate. This method helps to overcome the deficiencies of existing backdoor attack schemes in terms of visible features, robustness and concealment, and provides new challenges and ideas for the research of neural network model security analysis and defense technology. The phase-shifted neural network backdoor attack method described in the present invention specifically includes the following steps:

[0008] S1: Obtain and import raw image data: collect and preprocess raw images in RGB format to obtain a dataset for backdoor attack, and perform conventional normalization on the dataset;

[0009] S2: Perform color space conversion: Convert the RGB image obtained in step S1 pixel by pixel into the YCrCb color space, where Y represents the luminance component, and Cr and Cb represent the chrominance components; the RGB to YCrCb conversion follows the following formula or its equivalent form:

[0010] Y=0.299R+0.587G+0.114B

[0011] C b =128-0.168736R-0.331264G+0.5B,

[0012] C r =128+0.5R-0.418688G-0.081312B

[0013] Among them, R represents the red channel value, G represents the green channel value, B represents the blue channel value; Y represents the brightness component, C b and C r Represents the chrominance component;

[0014] S3: Perform a two-dimensional discrete Fourier transform (DFT) on the specified channel and extract the phase information: Select the Cr channel and the Cb channel to perform DFT; suppose the image space domain under this channel is represented as Y(x,y), and the formula for its two-dimensional discrete Fourier transform is:

[0015]

[0016] in, is the frequency domain, u and v represent the horizontal and vertical coordinates of the image in the frequency domain, x and y represent the horizontal and vertical coordinates of the image in the spatial domain, M and N represent the width and height of the image, and j is the imaginary unit;

[0017] S4: Select high and low frequency ranges and perform phase shift: Select several frequency points in the high and low frequency components of the frequency domain, wherein the high frequency component focuses on the edges and details of the image, and the low frequency component focuses on the overall structure of the image;

[0018] S5: Inverse Fourier transform and generate poisoned image: Perform inverse discrete Fourier transform on the phase-modified frequency domain coefficients to return to the spatial domain, merge the spatial domain pixel values of the selected channel after phase shift with the unmodified channel to form a complete image in YCrCb format, and perform inverse conversion from YCrCb to RGB to obtain a poisoned image that is visually similar to the original image;

[0019] S6: Perceptual filtering using LPIPS metric: The infected image and the original image are fed into the LPIPS metric algorithm to calculate the difference in their deep feature space. If the difference exceeds a preset threshold, the image is considered to have a visible anomaly at the perceptual level and is filtered. Otherwise, it is retained for subsequent training.

[0020] As a preferred technical solution of the present invention, the image selected channel in step S3 includes the Y channel, and the phase information obtained by discrete Fourier transform of the Y channel is calculated. Step S3 uses the following two-parameter method to avoid ambiguity of the phase angle in the second and third quadrants, and the calculation formula is:

[0021]

[0022] Where Re and Im are the real and imaginary parts of the calculated complex number, respectively, and θ(u,v) is the phase information.

[0023] As a technical preferred solution of the present invention, step S4 further divides the frequency domain information of the image, the high-frequency information is used to capture the edge and texture of the image, and the low-frequency information is used to represent the large-scale structure of the image, wherein the phase value of the high-frequency signal changes rapidly and exhibits a high degree of randomness. When its phase value is modified to the target value θ H (u,v), the strong concealment of the trigger is achieved. Modify the phase by θ L (u,v) is extended to a low-frequency signal with a more gradual change, thus forming a compound trigger that meets the following conditions:

[0024]

[0025] Among them, ω H and ω LRepresents the thresholds for the selected high-frequency and low-frequency components, respectively.

[0026] As a preferred technical solution of the present invention, the subsequent training in step S6 includes the following steps:

[0027] The poisoned images retained after screening in step S6 are labeled as backdoor target labels and constitute a mixed training set together with normal images; the mixed training set is trained using a conventional deep neural network training process to obtain a target model with a frequency-domain phase backdoor implanted; when the input image contains the same or similar phase perturbation as that in S4, the target model outputs a pre-set attacker target label during the inference phase, thereby achieving a covert backdoor attack.

[0028] Compared with the related prior art, the beneficial effects of the present invention are:

[0029] Higher stealth: By converting images from RGB channels to YCrCb channels and performing Fourier transform on the chroma channels with lower visual sensitivity, the backdoor triggers generated by the present invention can significantly reduce visual perceptibility, surpassing the stealth of traditional backdoor attacks that rely on local pixel modification or explicit marking.

[0030] Stronger attack effect: This invention modifies the phase information of a specified frequency in the frequency domain, so that the triggers are dispersed throughout the image in the form of fine textures and globally arranged. Compared with local triggers limited to a certain area of the image, this has a higher success rate and a wider range of applicability.

[0031] Greater naturalness and dispersion: The globally distributed texture design makes the trigger more natural in appearance, making it difficult to detect through conventional detection methods or simple visual comparison methods. It can also flexibly adapt to different types of image content and subsequent pre-processing operations.

[0032] Good versatility and robustness: The present invention can adjust the phase of both high-frequency and low-frequency components simultaneously, which helps to enhance the survivability of the backdoor under common data transformations (such as compression, cropping, filtering, etc.), and further improve the sustained effectiveness and stability of the attack. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] Figure 1 This is a method flow chart of a neural network backdoor attack method based on phase shift of the present invention;

[0034] Figure 2 Schematic diagram of a phase shift backdoor attack according to an embodiment of the present invention;

[0035] Figure 3 This is a heat map comparison diagram of the FDPS of the embodiment provided by the present invention and the original image;

[0036] Figure 4 This is an entropy value diagram of different backdoor attack methods facing STRIP defense according to the embodiment provided by the present invention. DETAILED DESCRIPTION

[0037] The present invention is further described below with reference to the accompanying drawings and examples. However, the present invention can be implemented in many different ways and should not be construed as limited to the illustrated embodiments; rather, these embodiments provide those skilled in the art with implementation methods that meet applicable legal requirements.

[0038] Example 1: As shown in the attached Figure 1 As shown, taking the backdoor attack on a convolutional neural network (CNN) in an image classification task as an example, the specific implementation process of the neural network backdoor attack method based on frequency domain phase shift is described in detail. The following steps correspond to the core process in the aforementioned claims and are described in more detail for this embodiment:

[0039] S1: Obtain and import raw image data: collect and preprocess raw images in RGB format to obtain a dataset for backdoor attack, and perform conventional normalization on the dataset;

[0040] S2: Perform color space conversion: Convert the RGB image obtained in step S1 pixel by pixel into the YCrCb color space, where Y represents the luminance component, and Cr and Cb represent the chrominance components; the RGB to YCrCb conversion follows the following formula or its equivalent form:

[0041] Y=0.299R+0.587G+0.114B

[0042] C b =128-0.168736R-0.331264G+0.5B,

[0043] C r =128+0.5R-0.418688G-0.081312B

[0044] Among them, R represents the red channel value, G represents the green channel value, B represents the blue channel value; Y represents the brightness component, C b and C r Represents the chrominance component;

[0045] S3: Perform a two-dimensional discrete Fourier transform (DFT) on the specified channel and extract the phase information: Select the Cr channel and the Cb channel to perform DFT; suppose the image space domain under this channel is represented as Y(x,y), and the formula for its two-dimensional discrete Fourier transform is:

[0046]

[0047] in, is the frequency domain, u and v represent the horizontal and vertical coordinates of the image in the frequency domain, x and y represent the horizontal and vertical coordinates of the image in the spatial domain, M and N represent the width and height of the image, and j is the imaginary unit;

[0048] The image selected channels include the Y channel, and the phase information is obtained by performing discrete Fourier transform on the Y channel. The following two-parameter method is used to calculate to avoid ambiguity in the phase angle in the second and third quadrants. The calculation formula is:

[0049]

[0050] Where Re and Im are the real and imaginary parts of the calculated complex number, respectively, and θ(u,v) is the phase information.

[0051] S4: Select high and low frequency ranges and perform phase shift: Select several frequency points in the high and low frequency components of the frequency domain, the high frequency component focuses on the edges and details of the image, and the low frequency component focuses on the overall structure of the image; divide the frequency domain information of the image, the high frequency information is used to capture the edges and textures of the image, and the low frequency information is used to represent the large-scale structure of the image, wherein the phase value of the high frequency signal changes rapidly and shows a high degree of randomness. When its phase value is modified to the target value θ H (u,v), the strong concealment of the trigger is achieved. Modify the phase by θ L (u,v) is extended to a low-frequency signal with a more gradual change, thus forming a compound trigger that meets the following conditions:

[0052]

[0053] Among them, ω H and ω L Represents the thresholds for the selected high-frequency and low-frequency components, respectively.

[0054] S5: Inverse Fourier transform and generate poisoned image: Perform inverse discrete Fourier transform on the phase-modified frequency domain coefficients back to the spatial domain, merge the spatial domain pixel values of the selected channel after phase shift with the unmodified channel to form a complete image in YCrCb format, and obtain a poisoned image that is visually similar to the original image through inverse conversion from YCrCb to RGB; the specific calculation formula is:

[0055]

[0056] Among them, f′(u,v) represents the pixel value of the channel in the spatial domain after phase shift; F′(u,v) is the frequency domain data corresponding to the phase shift;

[0057] S6: Perceptual filtering using LPIPS metric: The infected image and the original image are fed into the LPIPS metric algorithm to calculate the difference in their deep feature space. If the difference exceeds a preset threshold, the image is considered to have a visible anomaly at the perceptual level and is filtered. Otherwise, it is retained for subsequent training.

[0058] Example 2: As shown in the attached Figure 2 As shown in Figure 2, this embodiment evaluates the performance of the proposed phase-shifted neural network backdoor attack method (FDPS) on three datasets: Cifar10, Gtsrb, and ImageNet. ImageNet is a large dataset, and we randomly selected 10 categories for the experiment. In terms of models, this embodiment selected the mainstream image classification models ResNet18 and VGG19_bn for testing. Figure 3 As shown in Figure 3, the differences between different images under GradCam visualization defense methods are not significantly different from the heat map with added triggers and the original image.

[0059] Evaluation metrics include: Attack Success Rate (ASR): During the inference phase, if a backdoor trigger is embedded in the input image, the model outputs the category specified by the target attacker. ASR is the proportion of times the trigger successfully activates the backdoor and produces the attacker's expected output (i.e., the correct "misclassification" rate for trigger-activated inputs). Clean Accuracy (BA): The normal classification accuracy that the model can maintain on unmodified clean inputs. When the model's accuracy on clean data drops significantly, users often become suspicious of the model and abandon its use. Therefore, BA is also an important metric for measuring the quality of backdoor attacks. Similarity Metrics: To measure the stealthiness of the attack, we use Structural Similarity (SSIM) and LPIPS, two commonly used metrics for measuring image quality and perceptual distance, to quantify the degree of difference in visual features between images before and after trigger insertion. The closer the values, the smaller the image difference and the harder it is to detect the backdoor trigger.

[0060] Table 1 shows the comparative results of our method against various baseline methods on three datasets. As can be seen, FDPS achieves superior attack success rates (ASR) compared to baseline methods across multiple datasets and models, typically reaching around 99%. Furthermore, FDPS only minimally reduces the classification accuracy (BA) of clean images, effectively balancing attack effectiveness with performance on standard tasks.

[0061] Table 1: Benign sample rate (BA) and attack success rate (ASR) of different attack methods

[0062]

[0063] As shown in Table 1, FDPS can achieve an ASR of almost 99% on multiple datasets, with minimal impact on the accuracy of clean images. This demonstrates that our proposed method can maintain high accuracy for normal model reasoning while ensuring a high success rate for backdoor attacks.

[0064] As attached Figure 4 To further verify the visual concealment of the trigger, this example uses two common metrics, SSIM and LPIPS, to measure the similarity of the images before and after the trigger is embedded. Table 2 shows the comparison results of FDPS with other methods on Cifar10, Gtsrb, and ImageNet in terms of SSIM and LPIPS.

[0065] Table 2: Comparison of SSIM and LPIPS

[0066]

[0067] As shown in Table 2, FDPS achieves the best performance in the LPIPS metric, demonstrating minimal change in perceived distance and thus possessing strong stealth capabilities. In terms of SSIM, FDPS is very close to BadNets, with a slight difference but still maintaining a high level. Overall, FDPS maintains a very high visual similarity with the original image after embedding the backdoor trigger.

[0068] FDPS's resistance to three mainstream backdoor defense methods—Grad-CAM, Neural Cleanse, and STRIP—demonstrates the ability of our frequency-domain phase shift trigger to conceal itself from existing mainstream detection techniques. Grad-CAM uses backpropagation to weight the gradients of the convolutional layer, generating a heatmap that highlights the image regions the model focuses on during discrimination. If the model's focus on regions differs significantly between clean and triggered images, it is likely to indicate the presence of a backdoor. Neural Cleanse gradually optimizes a noise mask during model inference, searching for the minimum change that causes the model to misclassify the target class, thereby measuring the anomaly index. If the minimum trigger mask for a particular class is significantly smaller than that for other classes, or the anomaly index exceeds a threshold of 2, it is considered a backdoor attack. STRIP superimposes random clean images on suspicious inputs and observes the consistency of the model output. If the output entropy decreases significantly, it indicates the presence of a backdoor trigger.

[0069] Table 3: Comparison of outliers in different attack methods

[0070] Methods Clean BadNets Blend WaNet Ftrojan FDPS Index 0.76 4.28 3.26 2.32 1.85 1.73

[0071] In Table 3, the anomaly index of FDPS is significantly lower than the threshold of 2; while BadNets, Blend and WaNet can all be detected by NeuralCleanse.

[0072] This example demonstrates the high efficiency, stealth, and effective evasion of mainstream defenses of a phase-shift-based backdoor attack method, demonstrated through systematic experiments on three different datasets. Compared to traditional backdoor methods based on pixel or amplitude domains, this method leverages the imperceptible nature of phase information in the image's frequency domain, significantly improving the backdoor's stealth while maintaining the model's performance, making network security threats more subtle.

[0073] By applying a slight phase shift to specific frequency regions, the present invention can "integrate" the trigger into the underlying structure of the image without disrupting the overall brightness or local visual features. To further ensure high concealment, this embodiment also filters the generated trigger samples, removing those that differ from the original image by a threshold. This results in a highly concealed and fully functional backdoor trigger pattern in the final poisoned training set.

[0074] The backdoor attack described in this invention demonstrates significant superiority over other common techniques across various datasets and model architectures, particularly in the human-visible domain and with minimal perceptible interference using mainstream detection methods. This demonstrates the significant advantages and practical value of this technology in the field of deep learning model backdoor research, providing new research directions for subsequent security protection and adversarial detection, while also raising the bar for model security audits in industry applications.

[0075] The above embodiments merely illustrate several implementations of the present invention, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention. It should be noted that a person skilled in the art would be able to make numerous variations and improvements without departing from the spirit of the present invention, and all such variations and improvements fall within the scope of protection of the present invention.

Claims

1. A neural network backdoor attack method based on phase shift, characterized by: The method comprises the following steps: S1: Obtain and import raw image data: collect and preprocess raw images in RGB format to obtain a dataset for backdoor attack, and perform conventional normalization on the dataset; S2: Perform color space conversion: Convert the RGB image obtained in step S1 pixel by pixel into the YCrCb color space, where Y represents the luminance component, and Cr and Cb represent the chrominance components; the RGB to YCrCb conversion follows the following formula or its equivalent form: Y=0.299R+0.587G+0.114B C b =128-0.168736R-0.331264G+0.5B, C r =128+0.5R-0.418688G-0.081312B Among them, R represents the red channel value, G represents the green channel value, B represents the blue channel value; Y represents the brightness component, C b and C r Represents the chrominance component; S3: Perform a two-dimensional discrete Fourier transform (DFT) on the specified channel and extract the phase information: Select the Cr channel and the Cb channel to perform DFT; suppose the image space domain under this channel is represented as Y(x,y), and the formula for its two-dimensional discrete Fourier transform is: in, is the frequency domain, u and v represent the horizontal and vertical coordinates of the image in the frequency domain, x and y represent the horizontal and vertical coordinates of the image in the spatial domain, M and N represent the width and height of the image, and j is the imaginary unit; S4: Select high and low frequency ranges and perform phase shift: Select several frequency points in the high and low frequency components of the frequency domain, wherein the high frequency component focuses on the edges and details of the image, and the low frequency component focuses on the overall structure of the image; S5: Inverse Fourier transform and generate poisoned image: Perform inverse discrete Fourier transform on the phase-modified frequency domain coefficients to return to the spatial domain, merge the spatial domain pixel values of the selected channel after phase shift with the unmodified channel to form a complete image in YCrCb format, and perform inverse conversion from YCrCb to RGB to obtain a poisoned image that is visually similar to the original image; S6: Perceptual filtering using LPIPS metric: The infected image and the original image are fed into the LPIPS metric algorithm to calculate the difference in their deep feature space. If the difference exceeds a preset threshold, the image is considered to have a visible anomaly at the perceptual level and is filtered. Otherwise, it is retained for subsequent training.

2. The neural network backdoor attack method based on phase shift according to claim 1, characterized in that: The image selected channel in step S3 includes a Y channel, and the phase information is obtained by performing discrete Fourier transform on the Y channel.

3. The neural network backdoor attack method based on phase shift according to claim 1, characterized in that: Step S3 uses the following two-parameter method to avoid ambiguity of the phase angle in the second and third quadrants. The calculation formula is: θ(u,v)=atan2(Im(F(u,v)),Re(F(u,v))), Where Re and Im are the real and imaginary parts of the calculated complex number, respectively, and θ(u,v) is the phase information.

4. The neural network backdoor attack method based on phase shift according to claim 1, characterized in that: Step S4 further divides the frequency domain information of the image. The high-frequency information is used to capture the edge and texture of the image, and the low-frequency information is used to represent the large-scale structure of the image. The phase value of the high-frequency signal changes rapidly and exhibits high randomness. When its phase value is modified to the target value θ H (u,v), strong concealment of the trigger is achieved.

5. The neural network backdoor attack method based on phase shift according to claim 4, characterized in that: Step S4 further modifies the phase by θ L (u,v) is extended to a low-frequency signal with a more gradual change, thus forming a compound trigger that meets the following conditions: Among them, ω H and ω L Represents the thresholds for the selected high-frequency and low-frequency components, respectively.

6. The neural network backdoor attack method based on phase shift according to claim 1, characterized in that: The subsequent training in step S6 includes the following steps: The poisoned images retained after the screening in step S6 are labeled as backdoor target labels and together with the normal images form a mixed training set; The mixed training set is trained using a conventional deep neural network training process to obtain a target model implanted with a frequency domain phase backdoor; When the input image contains the same or similar phase offset as that in step S4, the target model outputs a preset attacker target label during the inference phase, thereby achieving a covert backdoor attack.

Citation Information

Patent Citations

  • Random frequency domain disturbance backdoor attack method for high-frequency injection of low-frequency trigger

    CN117951690A

  • Image classification backdoor attack method, device and equipment based on image contour

    CN118379534A

  • Backdoor attack method and system based on image steganography

    CN119378006A