Suspicious model backdoor category positioning method based on model assimilation

By locating suspicious categories in the dataset and cleaning and balancing dataset generation, the high cost and false detection rate problems of existing backdoor attack detection methods are solved, and efficient backdoor attack detection and model robustness are achieved.

CN120259784AActive Publication Date: 2025-07-04HANGZHOU JUNTONG FUTURE TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510730424.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-03
Publication Date
2025-07-04
Estimated Expiration
2045-06-03

AI Technical Summary

Technical Problem

The existing backdoor attack detection methods have high calculation cost, strong dependence on attack methods, high error detection rate, and long training time, making it difficult to achieve effective detection in multiple backdoor attack scenarios.

Method used

By measuring the attention pattern differences in different categories in the dataset, covariance discriminant analysis is used to locate suspicious categories, perform data cleaning and balance dataset generation, and retrain the model to improve detection accuracy and robustness.

Benefits of technology

It significantly improves the accuracy and robustness of backdoor attack detection, reduces the computational complexity and training time, and ensures the continuous security of the model in practical applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259784A_ABST
    Figure CN120259784A_ABST
Patent Text Reader

Abstract

The invention relates to the field of machine learning, in particular to a model assimilation-based suspicious model backdoor category positioning method, which comprises the following steps of: performing assimilation degree calculation on a model by utilizing an image data set, and judging whether the model is attacked by a backdoor or not; when the model is subjected to backdoor attack, covariance discriminant analysis is utilized to calculate a covariance index of each category so as to locate a suspicious category; performing data cleaning according to the suspicious category to obtain a clean data set; generating a balanced data set according to the clean data set; and retraining the model according to the balanced data set. According to the method, the suspicious poisoning type in the data set can be positioned by measuring the attention mode difference of different types in the data set, the backdoor model and the poisoning type are effectively detected in various backdoor attack scenes without depending on known backdoor information, and the backdoor attack detection accuracy and robustness are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of machine learning, and particularly to a method for locating suspicious model backdoor categories based on model assimilation. Background Art

[0002] With the wide application of deep neural network DNN in the field of computer vision, the security issue of the model has attracted increasing attention. Among them, backdoor attacks have become the main threat to the model. Backdoor attacks use specific triggers, causing the model to produce misclassifications when the trigger appears in the prediction stage. Due to the diversity of backdoor triggers, the model owner cannot obtain knowledge related to backdoor attacks, resulting in high concealment of backdoor attacks.

[0003] Currently, the defense against backdoor attacks mainly focuses on methods such as model purification and model detection. However, the above methods have problems such as high computational cost, strong dependence on the attack method, and high false detection rate. In addition, for other backdoor defense methods such as model fine-tuning and feature separation, although the robustness of the model is improved to a certain extent, its computational complexity is high, resulting in a long training time, and a certain number of clean samples and poisoned samples are also required for experiments, which also has certain limitations. Therefore, how to effectively detect backdoor models and poisoned types in multiple backdoor attack scenarios without relying on known backdoor information is of great significance for improving the accuracy and robustness of backdoor attack detection and reducing the computational complexity and training time of backdoor attack detection. Summary of the Invention

[0004] Aiming at the problems of high computational cost, strong dependence on the attack method, and high false detection rate existing in the existing backdoor attack detection methods, and aiming to utilize the unique manifestation of the assimilation phenomenon in the backdoor model to locate suspicious poisoned categories by measuring the difference in attention patterns of different categories in the dataset, the present invention provides a method for locating suspicious model backdoor categories based on model assimilation, and the method includes the following steps: Step S1: Input an image dataset containing several categories, calculate the assimilation degree of the model using the image dataset, and determine whether the model is under a backdoor attack; Step S2: When the model is under a backdoor attack, use covariance discriminant analysis to calculate the covariance index of each category, thereby locating the suspicious category; Step S3: According to the suspicious category, perform data cleaning to obtain a clean dataset; Step S4: Generate a balanced dataset according to the clean dataset; retrain the model according to the balanced dataset.

[0005] Preferably, in step S1, calculating the assimilation degree of the model using the image dataset and determining whether the model is under a backdoor attack specifically includes: Obtain several intermediate layers from the model to form a set of feature extraction layers; For each category of samples in the image dataset, generate a feature map corresponding to the sample according to each intermediate layer in the set of feature extraction layers, and calculate the attention map corresponding to each feature map; According to the attention map, perform cross-category attention map comparison on the image dataset; According to the result of the cross-category attention map comparison, perform assimilation degree calculation to obtain the assimilation score of each category in the image dataset, so as to determine whether the model is under a backdoor attack.

[0006] Preferably, in step S1, according to the attention map, performing cross-category attention map comparison on the image dataset is specifically: For any two categories in the image dataset, calculate the Frobenius norm difference between the attention maps corresponding to each intermediate layer of the two categories.

[0007] Preferably, in step S1, according to the result of the cross-category attention map comparison, performing assimilation degree calculation to obtain the assimilation score of each category in the image dataset, so as to determine whether the model is under a backdoor attack is specifically: Calculate the standard deviation of the Frobenius norm differences between the attention maps corresponding to all intermediate layers of the two categories to obtain the assimilation score of each category in the image dataset; Judge whether the model is under a backdoor attack according to the assimilation score of each category.

[0008] Preferably, in step S2, when the model is under a backdoor attack, use covariance discriminant analysis to calculate the covariance index of each category, so as to locate the suspicious category, specifically: When the model is under a backdoor attack, calculate the covariance matrix of the attention maps of each sample in the category; According to the covariance matrices of the attention maps of all samples in the category, calculate the average covariance matrix of the category; For each sample in the category, calculate the covariance discriminant score of the sample according to the covariance matrix of the sample and the average covariance matrix of the category; and analyze the top ten highest covariance discriminant scores to locate the suspicious category.

[0009] Preferably, in step S2, calculating the covariance matrix of the attention maps of each sample in the category is specifically: Extract the attention maps corresponding to all intermediate layers of each sample in the category, and flatten the attention maps into vectors; Calculate the covariance matrix of the attention map for each sample in the category according to the vector.

[0010] Preferably, in step S2, for each sample in the category, calculate the covariance discrimination score of the sample according to the covariance matrix of the sample and the average covariance matrix of the category; and analyze the top ten highest covariance discrimination scores to locate the suspicious category, specifically: For each sample in the category, calculate the Frobenius norm difference between the covariance matrix of the sample and the average covariance matrix of the category to obtain the covariance discrimination score of the sample; Take the top ten highest covariance discrimination scores. If the samples corresponding to the top ten highest covariance discrimination scores are concentrated in a certain category, determine that category as the suspicious category.

[0011] Preferably, in step S3, perform data cleaning according to the suspicious category to obtain a clean data set, specifically: Clean the samples of the suspicious category to obtain a clean data set; Alternatively, after separately isolating the suspicious category from the image data set, perform anomaly detection on the samples in the suspicious category to determine whether the samples are embedded with backdoor triggers; if so, remove the samples from the image data set; if not, lift the separate isolation of the samples to obtain a clean data set.

[0012] Preferably, in step S4, generate a balanced data set according to the clean data set, specifically: According to the samples removed from the clean data set, generate non-poisoned samples that match the characteristics of the removed samples through data augmentation and generative adversarial networks; Supplement the non-poisoned samples to the clean data set to generate a balanced data set.

[0013] Preferably, in step S4, retrain the model according to the balanced data set, specifically: Retrain the model using the balanced data set and introduce a regularization-based robustness constraint during the training process.

[0014] Compared with the prior art, the present invention has the following beneficial effects: First, through multi-layer feature mapping extraction and attention map generation, the present invention calculates the assimilation degree of each category using the Frobenius norm difference, and can accurately capture the attention pattern assimilation phenomenon caused by backdoor attacks in the model; by analyzing the assimilation scores of different categories in the model, it can effectively identify whether the model is under backdoor attack and locate potential poisoning types; Second, the present invention adopts the covariance discriminant analysis (CDA) method to calculate the covariance matrix of each sample and compare it with the average covariance matrix of its category, which can capture the subtle structural differences between samples; the covariance discriminant analysis method is used to sort the samples, accurately locate the poisoned samples, and significantly improve the accuracy of backdoor attack detection; Third, the present invention combines sample elimination and isolation mechanisms to ensure the purity of the model training data set by screening and eliminating samples with high covariance discriminant analysis scores; it also generates high-quality samples similar to real data samples through data enhancement and generative adversarial networks (GANs), balances the distribution of data of different categories, and prevents the model from overfitting due to data imbalance; Fourth, the present invention further enhances the defense capability of the model by retraining the model and introducing robustness constraints, ensuring that the model no longer relies on backdoor trigger features in actual application scenarios; combined with real-time backdoor detection and dynamic update mechanisms, the continuous security and robustness of the model during operation are ensured. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work. Among them: Figure 1 It is a flow chart of the method for locating suspicious model backdoor categories based on model assimilation provided by the present invention. DETAILED DESCRIPTION

[0016] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the specific implementation methods of the present invention are described in detail below in conjunction with the accompanying drawings. It is understood that the specific embodiments described herein are only used to explain the present invention, rather than to limit the present invention. It should also be noted that, for ease of description, only the parts related to the present invention rather than all structures are shown in the accompanying drawings. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.

[0017] The terms "include" and "have" and any variations thereof in the present invention are intended to cover non-exclusive inclusions. For example, a process, method, product or device comprising a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products or devices.

[0018] References herein to "embodiments" mean that particular features, structures, or characteristics described in connection with the embodiments can be included in at least one embodiment of the invention. The phrase occurring in various places in the specification is not necessarily referring to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive of other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.

[0019] Please refer to Figure 1 as shown, the present invention provides a method for locating suspicious model backdoor categories based on model assimilation, and the method includes the following steps: Step S1, input an image data set including several categories, calculate the assimilation degree of the model using the image data set, and determine whether the model is under a backdoor attack.

[0020] Further, in step S1, calculating the assimilation degree of the model using the image data set and determining whether the model is under a backdoor attack specifically includes: Obtain several intermediate layers from the model to form a feature extraction layer set; For each category of samples in the image data set, generate a feature map corresponding to the sample according to each intermediate layer in the feature extraction layer set, and calculate the attention map corresponding to each feature map; Perform cross-category attention map comparison on the image data set according to the attention map; According to the result of the cross-category attention map comparison, perform assimilation degree calculation to obtain the assimilation score of each category in the image data set, thereby determining whether the model is under a backdoor attack.

[0021] Further, in step S1, performing cross-category attention map comparison on the image data set according to the attention map specifically includes: For any two categories in the image data set, calculate the Frobenius norm difference between the attention maps corresponding to the two categories for each intermediate layer.

[0022] Further, in step S1, performing assimilation degree calculation according to the result of the cross-category attention map comparison to obtain the assimilation score of each category in the image data set, thereby determining whether the model is under a backdoor attack, specifically includes: Calculate the standard deviation of the Frobenius norm differences between the attention maps corresponding to all intermediate layers of the two categories to obtain the assimilation score of each category in the image data set; Determine whether the model is under a backdoor attack according to the assimilation score of each category.

[0023] In practical applications, first input an image data set including several categories into the model, and the image data set includes data that may contain backdoor samples.

[0024] For a certain data sample X in the image dataset, first perform forward propagation through the convolutional layer of the deep neural network in the model to generate the feature maps of each intermediate layer of the model. Let l1, l2,..., l i be several intermediate layers obtained from the image classification model, then let be the selected set of feature extraction layers. For each intermediate layer l i , the feature map i of the intermediate layer l can be obtained, where represents the feature map of the intermediate layer l i , with a size of C×H×W, where C represents the number of channels, and H and W represent the height and width of the feature map.

[0025] Preferably, the feature maps of each data sample can be extracted from multiple intermediate layers of the model (such as the 5th layer, 15th layer, 30th layer, and 40th layer), where the selection of the above intermediate layers is based on their ability to better represent the hierarchical information of the attention pattern.

[0026] For each feature map , calculate the corresponding attention map using the following formula, In the above formula, represents the element power of the feature map , C represents the number of channels, ‖‖ F represents the Frobenius norm, represents a constant to prevent division by zero; through the above process, it can be ensured that attention maps related to the samples are generated on multiple intermediate layers.

[0027] For each category i and category j in the image dataset, calculate the difference of the Frobenius norms between the attention maps of each intermediate layer l, and its formula is as follows: In the above formula, and represent the feature maps of category i and category j, ‖‖ F represents the Frobenius norm. By calculating the difference of the Frobenius norms between the attention maps of multiple intermediate layers, the degree of attention map assimilation between each category i and category j can be obtained.

[0028] Based on the difference of the Frobenius norms across categories (i.e., two different categories in the image dataset), calculate the assimilation score of category i using the following formula , In the above formula, represents the standard deviation across categories, and C represents the number of channels. When the assimilation score of category i is larger, it indicates that the difference between the attention maps of category i and other categories in the image dataset is smaller, that is, category i is more likely to be affected by the backdoor attack.

[0029] By analyzing the assimilation scores of each category in the image dataset, evaluate whether the samples in each category show overall assimilation. For example, the assimilation score of each category can be compared with a preset assimilation score threshold. When the assimilation score is greater than the preset assimilation score threshold, it is determined that the model is under a backdoor attack related to this category; otherwise, it is determined that the model is not under a backdoor attack.

[0030] Step S2, when the model is under a backdoor attack, use covariance discriminant analysis to calculate the covariance index of each category, so as to locate the suspicious category.

[0031] Furthermore, in step S2, when the model is under a backdoor attack, use covariance discriminant analysis to calculate the covariance index of each category, so as to locate the suspicious category, specifically: When the model is under a backdoor attack, calculate the covariance matrix of the attention maps of each sample in the category; According to the covariance matrix of the attention maps of all samples in the category, calculate the average covariance matrix of the category; For each sample in the category, calculate the covariance discriminant score of the sample according to the covariance matrix of the sample and the average covariance matrix of the category; and analyze the top ten highest covariance discriminant scores to locate the suspicious category.

[0032] Furthermore, in step S2, calculate the covariance matrix of the attention maps of each sample in the category, specifically: Extract the attention maps of each sample in the category corresponding to all intermediate layers, and flatten the attention maps into vectors; According to the vectors, calculate the covariance matrix of the attention maps of each sample in the category.

[0033] Furthermore, in step S2, for each sample in the category, calculate the covariance discriminant score of the sample according to the covariance matrix of the sample and the average covariance matrix of the category; and analyze the top ten highest covariance discriminant scores to locate the suspicious category, specifically: For each sample in the category, calculate the Frobenius norm difference between the covariance matrix of the sample and the average covariance matrix of the category to obtain the covariance discriminant score of the sample; Take the top ten highest covariance discriminant scores. If the samples corresponding to the top ten highest covariance discriminant scores are concentrated in a certain category, the category is determined to be a suspicious category.

[0034] In practical applications, for each sample, extract its attention map in each intermediate layer , and flatten it into a vector , used for further covariance analysis.

[0035] Specifically, the covariance matrix of the attention map of each sample is calculated using the following formula: , In the above formula, Represents the covariance matrix of the vectorized attention map, which describes the structural relationship between different regions of the attention map.

[0036] In calculating the average covariance matrix for each category , the mean covariance matrix It is obtained by averaging the covariance matrices of all samples in each category.

[0037] For sample X, through its covariance matrix The average covariance matrix with its category The Frobenius norm difference between the two is used to calculate the covariance discriminant score of sample X. , the calculation formula is as follows, In the above formula, ‖‖ F represents the Frobenius norm.

[0038] When the covariance discriminant score of sample X The larger the value is, the greater the deviation of sample C from the normal structural pattern of its category is, and the more likely it is to be attacked by a backdoor. For example, the covariance discriminant score of each sample X can be compared with the preset discriminant score threshold. When the covariance discriminant score is greater than the preset discriminant score threshold, sample X is judged to be a suspicious sample; otherwise, sample X is judged not to be a suspicious sample. For another example, for all samples under the same category, all samples are arranged in order from high to low according to their respective covariance discriminant scores. In the above arrangement, the samples with higher covariance discriminant scores (i.e., the samples with the higher arrangement positions) are more likely to be affected by backdoor attacks. In addition, by analyzing the samples with the highest covariance discriminant scores, the categories that are most likely to be attacked by backdoors are identified; generally speaking, samples with high covariance discriminant scores are concentrated in a specific category, indicating that the category may contain a backdoor; preferably, the top ten highest covariance discriminant scores are taken. If the samples corresponding to the top ten highest covariance discriminant scores are concentrated in a certain category, the category is determined to be a suspicious category.

[0039] Step S3: Clean the data according to the suspicious categories to obtain a clean data set.

[0040] Furthermore, in step S3, data cleaning is performed according to the suspicious categories to obtain a clean data set, specifically: Clean samples of suspicious categories to obtain a clean data set; Alternatively, after isolating the suspicious category from the image dataset, anomaly detection is performed on the samples in the suspicious category to determine whether the sample is embedded with a backdoor trigger; if so, the sample is removed from the image dataset; if not, the separate isolation of the sample is lifted to obtain a clean dataset.

[0041] In practical applications, samples with a covariance discriminant score greater than a preset discriminant score threshold or samples with the highest covariance discriminant score can be extracted. These samples are considered to be the most likely to be affected by backdoor attacks, that is, these samples are marked as suspicious samples. If there are a large number of suspicious samples in the image data set (i.e., the training data set), these suspicious samples can be directly removed from the image data set to obtain a clean data set. Alternatively, these suspicious samples can be processed by isolation and re-detection, that is, after isolating the suspicious samples from the image data set, anomaly detection is performed on the suspicious samples (such as feature-based adversarial detection methods) to determine whether the suspicious samples are embedded with backdoor triggers; if so, the suspicious samples are removed from the image data set; if not, the separate isolation of the suspicious samples is lifted to obtain a clean data set.

[0042] Step S4, generating a balanced data set based on the clean data set; and retraining the model based on the balanced data set.

[0043] Further, in step S4, a balanced dataset is generated based on the clean dataset, specifically as follows: According to the samples excluded from the clean dataset, non-poisoned samples that match the characteristics of the excluded samples are generated through data augmentation and generative adversarial networks; The non-poisoned samples are supplemented to the clean dataset to generate a balanced dataset.

[0044] Further, in step S4, the model is retrained based on the balanced dataset, specifically as follows: The model is retrained using the balanced dataset, and a regularization-based robustness constraint is introduced during the training process.

[0045] In practical applications, for the obtained clean dataset, new samples can be generated through data augmentation and generative adversarial networks (GANs) to supplement the excluded samples. Specifically, high-quality images similar to the originally excluded samples are generated using GANs. The generator and discriminator of GANs are trained adversarially, and a large number of non-poisoned samples (i.e., non-poisoned images) that match the characteristics of the excluded samples can be generated to ensure that the distribution of the dataset will not be imbalanced due to sample exclusion. For the categories with relatively high covariance discriminant scores (i.e., backdoor poisoning categories), by generating new non-poisoned samples, the number of samples in different categories is balanced to ensure that the model will not overfit to certain categories due to data imbalance.

[0046] When the model is retrained using the samples of the balanced dataset, it can be ensured that the model no longer depends on the backdoor trigger features, thereby reducing the possibility of the model being attacked again. Additionally, during the model retraining process, a regularization-based robustness constraint is introduced to enhance the model's resistance to abnormal data, where the above regularization can adopt L2 regularization or other adversarial regularization methods. Continuously monitor the performance of the model on the balanced dataset during the model retraining process to ensure that the model has a robust defense ability in real application scenarios.

[0047] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of adding a necessary general hardware platform, and of course, it can also be implemented by a combination of hardware and software. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a computer product. The present invention can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) that contain computer-usable program codes.

[0048] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than limiting it. Other embodiments can also be adopted. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features. And these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for locating suspicious model backdoor categories based on model assimilation, characterized in that, The method includes the following steps: Step S1: Input an image dataset containing several categories, calculate the assimilation degree of the model using the image dataset, and determine whether the model is under a backdoor attack; Step S2: When the model is under a backdoor attack, use covariance discriminant analysis to calculate the covariance index of each category, thereby locating the suspicious category; Step S3: According to the suspicious category, perform data cleaning to obtain a clean dataset; Step S4: Generate a balanced dataset based on the clean dataset; retrain the model based on the balanced dataset.

2. The method according to claim 1, wherein In step S1, calculating the assimilation degree of the model using the image dataset and determining whether the model is under a backdoor attack is specifically as follows: Obtain several intermediate layers from the model to form a feature extraction layer set; For each category of samples in the image dataset, generate a feature map corresponding to the sample according to each intermediate layer in the feature extraction layer set, and calculate the attention map corresponding to each feature map; According to the attention map, perform cross-category attention map comparison on the image dataset; According to the result of the cross-category attention map comparison, perform assimilation degree calculation to obtain the assimilation score of each category in the image dataset, thereby determining whether the model is under a backdoor attack.

3. The method according to claim 2, wherein In step S1, performing cross-category attention map comparison on the image dataset according to the attention map is specifically as follows: For any two categories in the image dataset, calculate the Frobenius norm difference between the attention maps corresponding to each intermediate layer of the two categories.

4. The method according to claim 3, wherein In step S1, performing assimilation degree calculation according to the result of the cross-category attention map comparison to obtain the assimilation score of each category in the image dataset, thereby determining whether the model is under a backdoor attack is specifically as follows: Calculate the standard deviation of the Frobenius norm differences between the attention maps corresponding to all intermediate layers of the two categories to obtain the assimilation score of each category in the image dataset; Determine whether the model is under a backdoor attack according to the assimilation score of each category.

5. The method according to claim 4, wherein In step S2, when the model is under a backdoor attack, using covariance discriminant analysis to calculate the covariance index of each category, thereby locating the suspicious category is specifically as follows: When the model is under a backdoor attack, calculate the covariance matrix of the attention maps of each sample in the category; According to the covariance matrices of the attention maps of all samples in the category, calculate the average covariance matrix of the category; For each sample in the category, calculate the covariance discriminant score of the sample according to the covariance matrix of the sample and the average covariance matrix of the category; And take the analysis of the top ten highest covariance discriminant scores to locate the suspicious category.

6. The method according to claim 5, wherein In step S2, calculate the covariance matrix of the attention maps of each sample in the category, specifically: Extract the attention maps of each sample in the category corresponding to all intermediate layers, and flatten the attention maps into vectors; Calculate the covariance matrix of the attention maps of each sample in the category according to the vectors.

7. The method according to claim 6, wherein In step S2, for each sample in the category, calculate the covariance discrimination score of the sample according to the covariance matrix of the sample and the average covariance matrix of the category; and analyze the top ten highest covariance discrimination scores to locate the suspicious category, specifically: For each sample in the category, calculate the Frobenius norm difference between the covariance matrix of the sample and the average covariance matrix of the category to obtain the covariance discrimination score of the sample; Take the top ten highest covariance discrimination scores. If the samples corresponding to the top ten highest covariance discrimination scores are concentrated in a certain category, determine that the category is a suspicious category.

8. The method according to claim 7, wherein In step S3, perform data cleaning according to the suspicious category to obtain a clean data set, specifically: Clean the samples in the suspicious category to obtain a clean data set; Alternatively, after separately isolating the suspicious category from the image data set, perform anomaly detection on the samples in the suspicious category to determine whether the samples are embedded with a backdoor trigger; if so, remove the samples from the image data set; if not, lift the separate isolation of the samples to obtain a clean data set.

9. The method according to claim 8, wherein In step S4, generate a balanced data set according to the clean data set, specifically: According to the samples removed from the clean data set, generate non-poisoned samples that match the characteristics of the removed samples through data augmentation and generative adversarial networks; Supplement the non-poisoned samples to the clean data set to generate a balanced data set.

10. The method according to claim 9, wherein In step S4, retrain the model according to the balanced data set, specifically: Retrain the model using the balanced data set and introduce a regularization-based robustness constraint during the training process.

Citation Information

Patent Citations

  • A Method and System for Resisting Neural Network Backdoor Attacks Based on Image Feature Analysis

    CN113205115B

  • Voiceprint recognition backdoor attack defense method based on feature clustering analysis and feature dimension reduction

    CN115331661A

  • Method for generating backdoor attack defense model based on target detection

    CN115632843A

  • Backdoor attack design and evaluation system and method of SAR image DNN classifier

    CN116524291A

  • Backdoor attack detection method based on data complexity measurement

    CN118172624A

Cited By

  • Backdoor attack detection method based on artificial intelligence

    CN121151096A