Sample filtering backdoor defense method, device and equipment for data poisoning
Through the combination of t-SNE and OPTICS algorithms, backdoor attacks in the prior art that cannot effectively identify and clear non-insert triggers are solved, and poisoned sample filtering is realized in the training stage, improving the security and accuracy of the model.
Patent Information
- Application Number
- CN202510333261.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2025-07-04
AI Technical Summary
The prior art is in response to backdoor attacks, especially non-insert triggers, with limited defense effects and inability to effectively remove poisoned texts during the training phase, resulting in the impact of model security and accuracy.
The t-SNE dimensionality reduction technology is used to retain the structural features of the text embedded vector, and combined with the OPTICS clustering algorithm, potential poisoning samples are identified and filtered, high-dimensional features are extracted through the RoBERTa model, and t-SNE algorithm is used to reduce the dimensionality, and OPTICS algorithm is used to cluster to identify and filter poisoning samples in low-density areas.
Effectively identify and filter poisoned samples, improve the safety and accuracy of the model during the training stage, prevent backdoor injection, and maintain the performance of the model on clean data.
Smart Images

Figure CN120256984A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of natural language processing, and specifically to a sample filtering backdoor defense method, device, and equipment for data poisoning. Background Art
[0002] With the rapid development of artificial intelligence (AI) and natural language processing (NLP) technologies, deep neural networks (DNNs) have been widely used in many fields such as text classification, question answering, and neural machine translation. To save computing resources and improve efficiency, more and more developers choose to use datasets and pre-trained models provided by third-party platforms. Although this approach can significantly shorten the training time, it also brings potential security risks.
[0003] Language models are vulnerable to malicious attacks, especially backdoor attacks, which have become a major security threat in NLP tasks in recent years and have received wide attention. In a backdoor attack, an attacker embeds a backdoor into an NLP model by injecting a small number of poisoned samples into the training data. During the inference stage, as long as the input data contains a specific trigger pattern, the model will output the target category preset by the attacker, while maintaining good overall performance on clean samples, making backdoor attacks extremely difficult to detect.
[0004] However, existing technologies still have deficiencies in dealing with backdoor attacks. Poisoned sample detection methods with triggers are the most common backdoor defense methods, mainly divided into two categories: dataset cleaning and poisoned text detection. Both of these methods can achieve good backdoor defense effects, but there are some drawbacks: the former has limited defense effects against non-inserted triggers (such as triggers based on text style or syntactic structure); the latter can only identify poisoned text during the inference stage and cannot effectively remove poisoned text during the training stage, failing to prevent the injection of backdoors from the source. Summary of the Invention
[0005] To overcome the deficiencies of the above-mentioned existing technologies, the present invention provides a sample filtering backdoor defense method, device, and equipment for data poisoning, which retains the structural features in the text embedding vectors through t-SNE dimensionality reduction technology and combines the OPTICS clustering algorithm to identify potential anomalies in the samples, and can more effectively separate poisoned samples from normal samples. Compared with traditional defense methods, t-OPTICS shows higher robustness when dealing with non-inserted triggers and can identify and filter poisoned samples during the training stage to prevent the injection of backdoors from the source.
[0006] In a first aspect, the present invention provides the following technical solution: A sample filtering backdoor defense method for data poisoning, including:
[0007] S1. Collect the original text dataset, and use the RoBERTa model to extract the semantic features of the text samples in the text dataset and convert them into high-dimensional embedding vectors;
[0008] S2. Use the t-SNE algorithm to reduce the high-dimensional embedding vectors to low-dimensional embedding vectors and retain the similarity between samples;
[0009] The process of retaining the similarity between samples is as follows: Calculate the Gaussian similarity between each pair of the high-dimensional embedding vectors, reduce the high-dimensional embedding vectors to low-dimensional embedding vectors, calculate the t-similarity between each pair of the low-dimensional vectors, and minimize the KL divergence between the Gaussian similarity and the t-similarity to retain the sample similarity;
[0010] S3. Perform density clustering on the low-dimensional embedding vectors through the OPTICS algorithm;
[0011] S4. After clustering is completed, filter the poisoned samples according to the clustering results, filter the poisoned samples in the low-density area, and retain the normal data.
[0012] Further, the specific steps of step S1 are as follows:
[0013] S101. Select and collect the original text dataset for model training, read the text data and remove the leading and trailing spaces;
[0014] S102. Input the text samples into the pre-trained RoBERTa model, select the output of the last hidden layer of the RoBERTa model as the embedding representation of the text samples, and store it as a high-dimensional embedding vector matrix, where each row represents the vector of a text sample.
[0015] Further, the specific steps of step S2 are as follows:
[0016] S201. The t-SNE algorithm uses the Gaussian distribution to calculate the Gaussian similarity between each pair of high-dimensional embedding vectors and reduces the high-dimensional embedding vectors to low-dimensional embedding vectors;
[0017] S202. The t-SNE algorithm uses the t-distribution to calculate the t-similarity between each pair of low-dimensional embedding vectors;
[0018] S203. t-SNE minimizes the Kullback-Leibler divergence (KL divergence) between the high-dimensional embedding vectors and the low-dimensional embedding vectors; the minimization process is iteratively optimized by the gradient descent method, and the objective function of the minimization is defined as follows:
[0019]
[0020] Among them, C represents the KL divergence, which measures the distribution difference between the high-dimensional embedding vectors and the low-dimensional embedding vectors. P and Q respectively represent the similarity distributions of the corresponding high-dimensional embedding vectors and low-dimensional embedding vectors between each pair of text samples. KL(P||Q) represents the difference between P and Q; p ij is the Gaussian similarity of the high-dimensional embedding vectors of text sample i and text sample j, and q ij is the t-similarity of the low-dimensional embedding vectors of text sample i and text sample j;
[0021] Furthermore, the specific steps of step S3 are as follows:
[0022] S301. Set the search radius to determine the neighborhood of the text samples; set the minimum number of neighborhood points MinPts, which represents the minimum number of neighborhood points required for a point to become a core point;
[0023] S302. Create an ordered queue A and a queue B for storing the final clustering results;
[0024] S303. If all text samples have been processed, end the clustering process; otherwise, select an unprocessed core text sample to start clustering. Starting from this sample, find all directly density-reachable text samples. If the directly density-reachable text sample is not in queue B, put it into queue A and sort it according to the reachable distance.
[0025] S304. If A is an empty queue, return to S303 to select the next unprocessed core text sample; otherwise, take out the text sample with the smallest reachable distance from queue A, store it in queue B, and perform the following processing:
[0026] ① Judge whether the embedding vector is a core text sample; if not, return to S303; if so, find all directly density-reachable samples of the embedding vector;
[0027] ② For the density-reachable samples, check whether they exist in queue B. If they already exist, skip this sample;
[0028] ③ If the text sample already exists in queue A and the current reachable distance is smaller than the original reachable distance, update the distance value of the text sample and re-sort queue A;
[0029] ④ If the text sample does not exist in queue A, insert it into queue A and re-sort.
[0030] S305. Repeat S303 and S304 until all samples have been clustered;
[0031] S306. Set the clustering threshold according to the reachable distance of the clustering results and the samples in queue B, and finally obtain the clustering results.
[0032] Further, the specific steps of S4 are as follows:
[0033] S401. According to the clustering results, first identify the maximum predicted cluster corresponding to each true label. That is, select the sample cluster with the highest density in each class as the representative cluster of normal samples. In this way, the regions with higher density in the clustering are regarded as the representatives of normal samples.
[0034] S402. For the regions with lower density in the clustering, discard all samples including potential poisoned text samples, and only retain the core text samples in the clustering, removing the marginal text samples and noise points;
[0035] S403. Finally, obtain a clean data set containing only credible text samples for subsequent analysis or model training, reducing the interference of poisoned text samples.
[0036] In a second aspect, based on the above-mentioned sample filtering backdoor defense method for data poisoning, an embodiment of the present invention further provides a sample filtering backdoor defense device for data poisoning, including:
[0037] A text feature extraction module based on the RoBERTa model, which is used to extract and store the high-dimensional features of data samples;
[0038] A dimensionality reduction module based on the t-SNE algorithm, which uses the t-SNE algorithm to reduce the high-dimensional features to low-dimensional features, retaining the similarity relationship between samples;
[0039] A clustering module based on the OPTICS algorithm, which uses the OPTICS algorithm to cluster the dimensionality-reduced features and identify the density clusters of samples.
[0040] A poisoned sample data filtering module, which screens samples according to the clustering results, discards potential poisoned samples in low-density regions, and retains credible data.
[0041] In a third aspect, an embodiment of the present invention provides a sample filtering backdoor defense device for data poisoning, including: a processor, a memory, a bus, and a computer program stored on the memory and executable on the processor;
[0042] Wherein, the processor and the memory communicate with each other through the bus;
[0043] When the processor executes the computer program, it implements the sample filtering backdoor defense method for data poisoning as described above.
[0044] Compared with the prior art, the beneficial effects of the present invention are:
[0045] (1) By combining the t-SNE dimensionality reduction technique and the OPTICS clustering algorithm, the present invention can effectively identify and filter out poisoned samples in the training phase. Compared with traditional backdoor defense methods, this method can not only cope with inserted triggers but also effectively handle non-inserted triggers (such as triggers based on text style or syntactic structure), thus preventing backdoor injection at the source and enhancing the security of the model.
[0046] (2) Through clustering analysis, the present invention can distinguish normal samples in high-density regions from poisoned samples in low-density regions, ensuring the integrity of normal samples while filtering out poisoned samples. This enables the model to be undisturbed by poisoned samples during training and maintain high practicality and accuracy. Especially in natural language processing tasks, it can maintain the performance of the model on clean data. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0048] Figure 1 It is a flowchart of the method of the present invention.
[0049] Figure 2 It is a structural diagram of the device of the present invention.
[0050] Figure 3 It is a schematic diagram of the system of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0051] To achieve the above objectives, the present invention is realized through the following technical solutions. The present invention provides a sample filtering backdoor defense method and device against data poisoning, which is applied to the field of natural language processing. The method includes: First, perform representation learning on the data set based on the RoBERTa model, extract and store the high-dimensional features of the data samples. Second, perform dimensionality reduction based on the t-SNE algorithm to reduce the high-dimensional features to low-dimensional features. Then, use the OPTICS algorithm to cluster the dimensionality-reduced features. Finally, screen the samples through the clustering results, filter the poisoned samples in the low-density area, and retain the credible normal data. Through the technical solution of the present invention, it is possible to filter poisoned samples while retaining normal samples during the training phase, ensuring that the model is not interfered by poisoned samples during the training process and maintaining high practicability and accuracy. The present invention provides a sample filtering backdoor defense method and device against data poisoning. The device includes: a text feature extraction module based on the RoBERTa model, a dimensionality reduction module based on the t-SNE algorithm, a clustering module based on the OPTICS algorithm, and a poisoned sample data filtering module.
[0052] Embodiment 1:
[0053] Figure 1 Fig. 7 shows a sample filtering backdoor defense method against data poisoning provided by an embodiment of the present invention, including the following steps:
[0054] S1. Use the RoBERTa model to extract semantic features and convert them into high-dimensional vector representations;
[0055] Select and collect the original text data set for model training, read the text data and remove the leading and trailing spaces.
[0056] Input each text sample into the pre-trained RoBERTa model. The RoBERTa model converts the text into a high-dimensional embedding vector of a fixed dimension through its embedding layer. Each text sample processed by the model will generate a 512-dimensional vector, which is used to represent the semantic features of the text.
[0057] Select the output of the last hidden layer of the RoBERTa model as the embedding representation of each text sample, and store it as a high-dimensional vector matrix, where each row represents the high-dimensional embedding vector of a text sample.
[0058] S2. Use the t-SNE algorithm for dimensionality reduction to retain the similarity between samples;
[0059] First, the t-SNE algorithm uses a Gaussian distribution in the high-dimensional space to calculate the conditional probability between each pair of text samples to reflect their similarity. These conditional probabilities represent the relative relationship between data points. Similar points have a higher probability, while dissimilar points have a lower probability. The points at a distance from point x i of all points xj The Gaussian similarity can be calculated as follows:
[0060]
[0061] where p j|i represents the similarity between text samples x i and x j , ∥x i -x j ∥ 2 represents the square of the Euclidean distance between text samples x i and x j , and is the variance of the Gaussian distribution. It should be noted that the similarity between points x i and x j is not symmetric, so the average of p j|i and p i|j is taken to obtain the final similarity:
[0062]
[0063] ij where n represents the total number of samples, and p i represents the similarity between text samples i and j. Next, t-SNE uses the t-distribution in the low-dimensional space to calculate the similarity between text samples, ensuring that the similarity relationship in the high-dimensional space is preserved in the low-dimensional space. The t-similarity of all points in the low-dimensional space is given by the formula:
[0064] where ∥y j -y 2 ∥ is the square of the Euclidean distance between text samples i and j in the low-dimensional space. t-SNE minimizes the Kullback-Leibler divergence (KL divergence) between the high-dimensional space and the low-dimensional space, so that similar text samples are also as close as possible in the low-dimensional space, while different text samples are as far apart as possible. The objective function is defined as follows:
[0066]
[0067] The smaller the KL divergence, the better the embedding result in the low-dimensional space can preserve the similarity relationship in the high-dimensional space. This process is iteratively optimized by the gradient descent method, with the goal of making the layout of data points in the low-dimensional space as accurately as possible reflect the similarity relationship in the high-dimensional space.
[0068] S3. Perform density clustering through the OPTICS algorithm;
[0069] On the low-dimensional embedded vectors of the dimensionality-reduced text, the OPTICS algorithm is used for clustering. First, set the search radius to determine the neighborhood of the samples; set the minimum number of points in the neighborhood MinPts, which represents the minimum number of neighboring points required for a point to become a core point.
[0070] S31) Create two queues: one is the ordered queue A, and the other is the queue B that stores the final clustering results.
[0071] 32) If all samples have been processed, end the clustering process. Otherwise, select an unprocessed core sample (i.e., an unprocessed low-dimensional embedded vector) to start clustering. Starting from this sample, find all its directly density-reachable embedded vectors. If these embedded vectors are not in B, put them into A and sort them by the reachable distance.
[0072] 33) If A is an empty sequence, return to step 2 to select the next unprocessed core sample. Otherwise, take out the sample with the minimum reachable distance from A, store it in the B queue, and perform the following processing:
[0073] ① Judge whether this embedded vector is a core sample. If not, return to step 3; if this sample is a core sample, find all the directly density-reachable samples of this embedded vector.
[0074] ② For the density-reachable samples, check whether they have been added to B. If they already exist, skip this sample; if not, continue to the next step.
[0075] ③ If this sample is already in the A queue and the current distance is smaller than the original distance, update the distance value of this sample and re-sort the A queue.
[0076] ④ If this sample is not in the A queue, insert it and re-sort.
[0077] 34) Repeat steps 32) and 33) until all samples have been processed.
[0078] 35) According to the reachable distance of the clustering results and the samples in the B queue, set the clustering threshold to finally obtain the clustering results.
[0079] S4. Filter the poisoned samples in the low-density area and retain the normal data.
[0080] After clustering is completed, filter the poisoned samples according to the clustering results. Assume that the number of poisoned samples is usually small. Therefore, through clustering analysis, poisoned samples can be effectively distinguished from normal samples.
[0081] 41) According to the clustering results, first identify the maximum predicted cluster corresponding to each true label. That is, select the sample cluster with the highest density in each class as the representative cluster of normal samples. In this way, the regions with higher density in the clustering are regarded as the representatives of normal samples.
[0082] 42) For the regions with lower density in the clustering, discard their samples (including potential poisoned samples). Through this filtering step, ensure that only the core samples in the clustering are retained, and the marginal samples and noise points are removed.
[0083] Finally, obtain a clean processed dataset. By training the model using the clean dataset and the original dataset respectively, and evaluating the model using the clean test dataset and the poisoned test dataset, to compare the performance of the model on poisoned data and clean data, so as to evaluate whether the poisoned samples have been effectively removed. The evaluation methods include the following:
[0084] ΔASR: Reduction in the attack success rate (ASR, classification accuracy of test samples filled with triggers), reflecting the model's defense ability against attacks;
[0085] ΔCACC: Reduction in clean accuracy (CACC, accuracy of the model on normal test samples), measuring the performance of the model on normal samples.
[0086] Higher ΔASR and lower ΔCACC indicate that this method effectively removes poisoned samples and maintains the performance of normal samples.
[0087] Example 2:
[0088] A sample filtering backdoor defense device against data poisoning provided by an embodiment of the present invention, in combination with Figure 2 , includes: a text feature extraction module based on the RoBERTa model; a dimensionality reduction module based on the t-SNE algorithm; a clustering module based on the OPTICS algorithm; a poisoned sample data filtering module, where:
[0089] The text feature extraction module based on the RoBERTa model is used to extract and store high-dimensional features of data samples;
[0090] The dimensionality reduction module based on the T-sne algorithm is used to reduce the high-dimensional features to low-dimensional features;
[0091] The clustering module based on the OPTICS algorithm uses the OPTICS algorithm to cluster the low-dimensional features.
[0092] The poisoned sample data filtering module screens samples according to the clustering results, discards potential poisoned samples in the low-density regions, and retains trustworthy data.
[0093] In this way, it is possible to effectively distinguish poisoned samples from clean samples, filter out potential poisoned samples, while retaining normal data, improving the security of the model, and ensuring that while defending against backdoor attacks, high accuracy and practicality can still be maintained in actual tasks.
[0094] Embodiment 3:
[0095] An embodiment of the present invention provides a sample filtering backdoor defense device against data poisoning, in combination with Figure 3 , including: a processor, a memory, a bus, and a computer program stored on the memory and executable on the processor;
[0096] Among them, the processor and the memory communicate with each other through the bus;
[0097] When the processor executes the computer program, the method of Embodiment 1 is implemented, for example, including: performing representation learning on the data set using the RoBERTa model; reducing high-dimensional features to low-dimensional features based on the t-SNE algorithm; clustering data samples based on the OPTICS algorithm; screening samples according to the clustering results, discarding potential poisoned samples in low-density regions, and retaining trustworthy data.
[0098] The present invention provides a sample filtering backdoor defense method, device and equipment against data poisoning, which are applied to the field of natural language processing. By combining the RoBERTa model, the t-SNE dimensionality reduction algorithm and the OPTICS clustering algorithm, poisoned samples can be effectively filtered out and normal data can be retained, thereby improving the security and practicality of the model.
[0099] As described above, only the specific embodiments of the present application are provided, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claimed rights.
Claims
1. A sample filtering backdoor defense method against data poisoning, comprising: S1. Collect the original text dataset, and use the RoBERTa model to extract the semantic features of the text samples in the text dataset and convert them into high-dimensional embedding vectors; S2. Use the t-SNE algorithm to reduce the dimensionality of the high-dimensional embedding vectors into low-dimensional embedding vectors and retain the similarity between samples; The process of retaining the similarity between samples is as follows: Calculate the Gaussian similarity between each pair of the high-dimensional embedding vectors, reduce the dimensionality of the high-dimensional embedding vectors into low-dimensional embedding vectors, calculate the t-similarity between each pair of the low-dimensional vectors, and minimize the KL divergence between the Gaussian similarity and the t-similarity to retain the sample similarity; S3. Perform density clustering on the low-dimensional embedding vectors through the OPTICS algorithm; S4. After clustering is completed, filter the poisoned samples according to the clustering results, filter the poisoned samples in the low-density area, and retain the normal data.
2. The sample filtering backdoor defense method against data poisoning according to claim 1, the specific steps of S1 are as follows: S101. Select and collect the original text dataset for model training, read the text data and remove the leading and trailing spaces; S102. Input the text sample into the pre-trained RoBERTa model, select the output of the last hidden layer of the RoBERTa model as the embedding representation of the text sample, and store it as a high-dimensional embedding vector matrix, where each row represents the vector of a text sample.
3. The sample filtering backdoor defense method against data poisoning according to claim 1, the specific steps of S2 are: S201. The t-SNE algorithm uses the Gaussian distribution to calculate the Gaussian similarity between each pair of high-dimensional embedding vectors and reduces the dimensionality of the high-dimensional embedding vectors into low-dimensional embedding vectors; S202. The t-SNE algorithm uses the t-distribution to calculate the t-similarity between each pair of low-dimensional embedding vectors; S203. t-SNE minimizes the KL divergence between the high-dimensional embedding vectors and the low-dimensional embedding vectors; the minimization process is iteratively optimized by the gradient descent method, and the objective function of the minimization is defined as follows: Among them, Let \(C\) denote the Kullback-Leibler divergence, which measures the distribution difference between the high-dimensional embedding vectors and the low-dimensional embedding vectors. \(P\) and \(Q\) respectively represent the similarity distributions of the corresponding high-dimensional embedding vectors and low-dimensional embedding vectors between each pair of text samples, and \(KL(P||Q)\) represents the difference between \(P\) and \(Q\); \(p\) ij is the Gaussian similarity of the high-dimensional embedding vectors of text sample \(i\) and text sample \(j\), and \(q\) ij is the t-similarity of the low-dimensional embedding vectors of text sample \(i\) and text sample \(j\).
4. The sample filtering backdoor defense method against data poisoning according to claim 1, the specific steps of S3 are: S301. Set the search radius to determine the neighborhood of the text sample; set the minimum number of neighborhood points MinPts, which represents the minimum number of neighborhood points required for a point to become a core point; S302. Create an ordered queue A and a queue B for storing the final clustering results; S303. If all text samples have been processed, end the clustering process; otherwise, select an unprocessed core text sample to start clustering. Starting from this sample, find all directly density-reachable text samples. If the directly density-reachable text samples are not in queue B, put them into queue A and sort them according to the reachable distance; S304. If A is an empty queue, return to S303 to select the next unprocessed core text sample; otherwise, take out the text sample with the minimum reachable distance from queue A, store it in queue B, and perform the following processing: ① Determine whether the embedding vector is a core text sample; if not, return to S303; if so, find all directly density-reachable samples of the embedding vector; ② For the density-reachable samples, check whether they exist in queue B; if they already exist, skip this sample; ③ If the text sample already exists in queue A and the current reachable distance is smaller than the original reachable distance, update the distance value of the text sample and reorder queue A; ④ If the text sample does not exist in queue A, insert it into queue A and reorder; S305. Repeat S303 and S304 until all samples have been clustered; S306. Set the clustering threshold according to the reachable distance of the clustering result and the samples in queue B, and finally obtain the clustering result.
5. According to the sample filtering backdoor defense method against data poisoning described in claim 1, the specific steps of S4 are as follows: S401. According to the clustering result, first identify the maximum predicted cluster corresponding to each true label, that is, select the sample cluster with the highest density in each class as the representative cluster of normal samples; the area with higher density in the clustering is regarded as the representative of normal samples; S402. For the area with lower density in the clustering, discard all samples including potential poisoned text samples, and only retain the core text samples in the clustering, removing marginal text samples and noise points; S403. Finally, obtain a clean data set containing only trustworthy text samples for subsequent analysis or model training, reducing the interference of poisoned text samples.
6. A sample filtering backdoor defense device against data poisoning, characterized in that, Including: A text feature extraction module based on the RoBERTa model, used to extract and store high-dimensional features of data samples; A dimensionality reduction module based on the t-SNE algorithm, using the t-SNE algorithm to reduce the high-dimensional features to low-dimensional features, retaining the similarity relationship between samples; A clustering module based on the OPTICS algorithm, using the OPTICS algorithm to cluster the dimensionality-reduced features and identify the density clusters of samples; A poisoned sample data filtering module, which filters samples according to the clustering result, discards potential poisoned samples in the low-density area, and retains trustworthy data.
7. A sample filtering backdoor defense device against data poisoning, characterized in that, Including: A processor, a memory, a bus, and a computer program stored on the memory and executable on the processor; Wherein, the processor and the memory communicate with each other through the bus; When the processor executes the computer program, it implements the sample filtering backdoor defense method against data poisoning according to any one of claims 1-5.