Data processing method and device, electronic equipment and storage medium
By preprocessing, clustering analysis and abnormal detection of text data sets, identifying and removing poisoned samples, the problem that the existing technology cannot fully respond to text data backdoor attacks is solved, and efficient and accurate defense effects are achieved.
Patent Information
- Application Number
- CN202510167923.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-14
- Publication Date
- 2025-06-06
AI Technical Summary
The existing technology is difficult to fully respond to various trigger types and attack methods for text data backdoor attacks, resulting in poor defense capabilities.
By obtaining the sample data set, preprocessing and clustering analysis, screening and verification using the trained anomaly detection model, combining the cluster analysis results to determine the poisoned sample and remove it from the data set.
It improves the ability to identify different types of poisoned samples, is suitable for different types of text data sets, and can identify different types and levels of poisoned samples. It has low false alarm rate, fast processing speed, and is effectively defended against text backdoor attacks.
Smart Images

Figure CN120105095A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of data processing technology, and in particular to a data processing method, device, electronic device and storage medium. Background Art
[0002] When training large-scale deep learning models, the quality of the dataset is directly related to the performance and reliability of the model. However, malicious actors may take advantage of this feature and inject specially designed samples (i.e., poisoned samples) into the training data to induce the model to produce incorrect responses to certain inputs or leak sensitive information in the future. This type of attack is called a "backdoor attack" or "data poisoning attack."
[0003] Among the data processing related technologies, the current defense technologies against text data backdoor attacks mostly focus on certain specific types of triggers, but it is difficult to comprehensively cover all potential attack methods. The defense capabilities are relatively poor when facing complex grammatical styles and semantic-level attacks. Summary of the invention
[0004] The present disclosure provides a data processing method and device, an electronic device and a storage medium, the main purpose of which is to solve the problem that the current defense method cannot cope with a full range of trigger types and attack methods, resulting in poor defense capabilities.
[0005] According to a first aspect of the present disclosure, a data processing method is provided, comprising:
[0006] Get a sample dataset;
[0007] Preprocessing the sample data set to obtain a first data set;
[0008] Performing cluster analysis on the first data set to generate a cluster analysis result including a cluster identifier to which each outlier belongs;
[0009] Using the trained anomaly detection model, screening and verifying the first data set to obtain anomaly detection results;
[0010] Determine the poisoned samples in the sample data set by comparing the cluster analysis result and the anomaly detection result;
[0011] The poisoned sample is removed from the sample data set to obtain a target data set.
[0012] Optionally, performing cluster analysis on the first data set to generate a cluster analysis result including a cluster identifier to which each outlier belongs includes:
[0013] Using a density-based clustering algorithm, determining the optimal number of clusters for the first data set and processing noise points, wherein the noise points are points far away from any cluster;
[0014] Determining, according to the clustering results obtained by the density-based clustering algorithm, a main cluster belonging to normal samples in the first data set;
[0015] Agglomerative hierarchical clustering is performed separately for each selected cluster in the main clusters to obtain the cluster analysis result.
[0016] Optionally, performing agglomerative hierarchical clustering separately for each selected cluster in the main clusters to obtain the cluster analysis result includes:
[0017] Calculating an average distance and / or similarity between internal samples of each selected cluster, and determining the internal consistency of each selected cluster according to the average distance and / or the similarity;
[0018] The cluster analysis result is obtained by comparing the internal consistency of each selected cluster with a preset threshold.
[0019] Optionally, obtaining the cluster analysis result by comparing the internal consistency of each selected cluster with a preset threshold comprises:
[0020] If the internal consistency of the cluster to be marked is lower than the preset threshold, it is determined that there are abnormal points in the cluster to be marked, and the cluster to be marked is marked as an abnormal sample.
[0021] Optionally, the anomaly detection model includes an isolation forest model and a support vector machine model;
[0022] The training process of the anomaly detection model includes:
[0023] According to the cluster analysis result, filtering the first data set to obtain a second data set;
[0024] Training the isolation forest model based on the second data set;
[0025] Using a pre-trained isolation forest model to predict the first data set, and determining the data set marked as normal by the pre-trained isolation forest model as a third data set;
[0026] The support vector machine model is trained based on the third data set.
[0027] Optionally, the using of the trained anomaly detection model to screen and verify the first data set to obtain an anomaly detection result includes:
[0028] Using the pre-trained isolation forest model to predict the first data set, and marking sample data in the first data set according to the prediction result to obtain a first marking result;
[0029] Using the trained support vector machine model to predict the first data set, and marking sample data in the first data set according to the prediction result to obtain a second marking result;
[0030] The abnormality detection result is determined according to the first marking result and the second marking result.
[0031] Optionally, preprocessing the sample data set to obtain the first data set includes:
[0032] Stepwise splitting the sample data set into word levels;
[0033] Use the pre-trained word embedding model to map each word into a dense vector of fixed dimension;
[0034] Using nonlinear dimensionality reduction techniques, vector embedding with dimensions higher than a preset dimension is mapped into a space with dimensions lower than a preset dimension.
[0035] According to a second aspect of the present disclosure, there is provided a data processing device, comprising:
[0036] An acquisition module is used to obtain a sample data set;
[0037] A processing module, used for preprocessing the sample data set to obtain a first data set;
[0038] An analysis module, configured to perform cluster analysis on the first data set to generate a cluster analysis result including a cluster identifier to which each outlier belongs;
[0039] A detection module, used to screen and verify the first data set using the trained anomaly detection model to obtain an anomaly detection result;
[0040] A determination module, used to determine the poisoned samples in the sample data set by comparing the cluster analysis result and the anomaly detection result;
[0041] The removal module is used to remove the poisoned sample from the sample data set to obtain a target data set.
[0042] Optionally, the analysis module is also used to use a density-based clustering algorithm to determine the optimal number of clusters of the first data set and process noise points, where the noise points are points far away from any cluster; based on the clustering results obtained by the density-based clustering algorithm, determine the main clusters belonging to normal samples in the first data set; and perform agglomerative hierarchical clustering separately for each selected cluster in the main clusters to obtain the cluster analysis results.
[0043] Optionally, the analysis module is also used to calculate the average distance and / or similarity between internal samples of each selected cluster, and determine the internal consistency of each selected cluster based on the average distance and / or the similarity; and obtain the cluster analysis result by comparing the internal consistency of each selected cluster with a preset threshold.
[0044] Optionally, the detection module is further configured to determine that an abnormal point exists in the cluster to be marked if the internal consistency of the cluster to be marked is lower than the preset threshold, and mark the cluster to be marked as an abnormal sample.
[0045] Optionally, the anomaly detection model includes an isolation forest model and a support vector machine model; accordingly, the detection module is also used to filter the first data set according to the cluster analysis results to obtain a second data set; train the isolation forest model based on the second data set; use the pre-trained isolation forest model to predict the first data set, and determine the data set marked as normal by the pre-trained isolation forest model as a third data set; and train the support vector machine model based on the third data set.
[0046] Optionally, the detection module is also used to use the pre-trained isolation forest model to predict the first data set, and mark the sample data in the first data set according to the prediction result to obtain a first marking result; use the trained support vector machine model to predict the first data set, and mark the sample data in the first data set according to the prediction result to obtain a second marking result; determine the anomaly detection result based on the first marking result and the second marking result.
[0047] Optionally, the processing module is also used to gradually split the sample data set into word levels; use a pre-trained word embedding model to map each word into a dense vector of a fixed dimension; use nonlinear dimensionality reduction technology to embed and map vectors higher than a preset dimension into a space lower than a preset dimension.
[0048] According to a third aspect of the present disclosure, there is provided an electronic device, including:
[0049] at least one processor; and
[0050] a memory communicatively connected to the at least one processor; wherein,
[0051] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method described in the first aspect.
[0052] According to a fourth aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to execute the method described in the first aspect.
[0053] According to a fifth aspect of the present disclosure, a computer program product is provided, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements the method as described in the first aspect above.
[0054] The data processing method and device, electronic device and storage medium provided by the present disclosure first obtain a sample data set; pre-process the sample data set to obtain a first data set; then perform cluster analysis on the first data set to generate a cluster analysis result including a cluster identifier to which each abnormal point belongs; then use the trained anomaly detection model to screen and verify the first data set to obtain an anomaly detection result; determine the poisoned samples in the sample data set by comparing the cluster analysis results and the anomaly detection results; remove the poisoned samples from the sample data set to obtain a target data set. Compared with the related art, the present application cleans the sample data set into a target data set by comprehensively applying a clustering algorithm and anomaly detection model, thereby improving the ability to recognize different types of poisoned samples, being applicable to different types of text data sets, and being able to recognize poisoned samples of different types and levels, including triggers at the word level, sentence level, and more complex grammatical styles and semantic levels, with a low false alarm rate and fast processing speed, and can effectively defend against text backdoor attacks.
[0055] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present application, nor is it intended to limit the scope of the present application. Other features of the present application will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] The accompanying drawings are used to better understand the present solution and do not constitute a limitation of the present disclosure.
[0057] Figure 1 A flowchart of a data processing method provided by an embodiment of the present disclosure;
[0058] Figure 2 A schematic diagram of an example process provided by an embodiment of the present disclosure;
[0059] Figure 3A flowchart of a data processing method provided by the present disclosure;
[0060] Figure 4 A schematic diagram of an example process provided by the present disclosure;
[0061] Figure 5 A schematic diagram of the structure of another data processing device provided by an embodiment of the present disclosure;
[0062] Figure 6 A schematic block diagram of an exemplary electronic device provided for an embodiment of the present disclosure. DETAILED DESCRIPTION
[0063] The following is a description of exemplary embodiments of the present disclosure in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be recognized by those of ordinary skill in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0064] The data processing method and apparatus, electronic device, and storage medium according to the embodiments of the present disclosure are described below with reference to the accompanying drawings.
[0065] Figure 1 A flowchart of a data processing method provided by an embodiment of the present disclosure.
[0066] like Figure 1 As shown, the method comprises the following steps:
[0067] Step 101: Obtain a sample data set.
[0068] A sample dataset is a collection of text records, usually used for natural language processing (NLP) tasks such as text classification, sentiment analysis, topic modeling, machine translation, etc. It can contain various types of text content, such as news articles, product reviews, social media posts, emails, book chapters, etc.
[0069] Step 102: preprocess the sample data set to obtain a first data set.
[0070] The original sample data set usually contains noise, redundant information, and inconsistent formats. By removing labels, special characters, punctuation marks, etc., the noise in the text can be reduced to make the data cleaner. You can also handle missing values in the data set. If there are fewer missing values, you can choose to directly delete records with missing values. You can also use the mean, median, or mode to fill missing values, or interpolate according to the context to avoid affecting the accuracy of the data set. Classifying text content can help identify and organize information. You can also convert text data into numerical vectors. Considering the problems of data sparsity and computational complexity in high-dimensional space, high-dimensional word embeddings are mapped to lower-dimensional spaces for easy input into machine learning models.
[0071] By preprocessing the sample data, the data quality can be improved. The obtained high-quality first data set can support more complex analysis methods, making subsequent model training and analysis more accurate and efficient.
[0072] Step 103: Perform cluster analysis on the first data set to generate a cluster analysis result including a cluster identifier to which each abnormal point belongs.
[0073] Cluster analysis is an unsupervised learning method used to divide objects in a data set into multiple groups or clusters, so that objects in the same cluster have high similarities, while objects in different clusters have great differences. Then, the patterns and structures within each cluster are identified. Normal data usually form tight and consistent clusters, while poisoned samples may form independent small clusters or be distributed on the edges of existing clusters. At this stage, cluster analysis can help identify potential structures and patterns in the data set, as well as possible poisoned data clusters.
[0074] Step 104 , using the trained anomaly detection model, screen and verify the first data set to obtain anomaly detection results.
[0075] Anomaly detection models are used to identify data points in a dataset that are significantly different from the majority of the data. These data points may be poisoned data or outliers. Poisoned samples usually have different characteristics from normal samples, so they can be identified by anomaly detection methods.
[0076] Step 105, by comparing the cluster analysis results and the anomaly detection results, the poisoned samples in the sample data set are determined.
[0077] By comparing the cluster analysis results and anomaly detection results, we can directly compare the cluster labels and anomaly labels, and the anomaly samples with the merged labels are poisoned samples.
[0078] For example, the cluster label to which each sample belongs can be obtained from cluster analysis, and the label of whether each sample is an outlier can be obtained from anomaly detection. Then, the cluster label and anomaly label of each sample can be compared: if a sample is labeled as an outlier and the cluster it belongs to has a low consistency, it is considered to be a poisoned sample; if a sample is labeled as an outlier but the cluster it belongs to has a high internal consistency, further analysis is required, as it may be a normal outlier rather than a poisoned sample.
[0079] Step 106, remove the poisoned samples from the sample data set to obtain the target data set.
[0080] After removing the samples confirmed to be poisoned from the sample data set, the remaining data set is a clean data set without poisoned samples, i.e., the target data set. After removing the poisoned samples, the remaining data set can also be verified to ensure its quality and consistency.
[0081] The data processing method in this embodiment is one of the effective methods to defend against backdoor attacks on text data. The purpose of cleaning the sample data set into the target data set is to clean up abnormal samples in the sample data and ensure the purity of the data set, so as to protect the NLP model from the influence of malicious samples and ensure its security and reliability in practical applications.
[0082] In order to better understand the data processing method in this embodiment, Figure 2 As shown, Figure 2 An example process diagram provided for the present disclosure includes preprocessing a poisoning data set, wherein the poisoning data set is a sample data set mixed with some poisoning data. In the data preprocessing stage, the data will be cleaned, sorted and normalized, which may include deleting duplicate data, filling missing values, converting data formats, text to vector conversion, nonlinear dimensionality reduction and other operations, and ensuring that the data set is suitable for subsequent analysis. Then, the preprocessed data will enter the cluster analysis stage, and after density-based cluster analysis, hierarchical cluster analysis and other analysis operations, the cluster analysis results are finally obtained. Then, in the anomaly detection stage, a variety of anomaly detection algorithms can be used to identify possible poisoning data. Finally, according to the results of cluster analysis and anomaly detection, the identified poisoning data is removed from the data set or corrected. In this way, the data set finally obtained after completing the entire data cleaning process is a clean data set, which can not only improve the reliability of the data set, but also ensure the effect of subsequent analysis and model training using the data set.
[0083] The data processing method and device, electronic device and storage medium provided by the present disclosure first obtain a sample data set; pre-process the sample data set to obtain a first data set; then perform cluster analysis on the first data set to generate a cluster analysis result including a cluster identifier to which each abnormal point belongs; then use the trained anomaly detection model to screen and verify the first data set to obtain an anomaly detection result; determine the poisoned samples in the sample data set by comparing the cluster analysis results and the anomaly detection results; remove the poisoned samples from the sample data set to obtain a target data set. Compared with the related art, the embodiment of the present disclosure cleans the sample data set into a target data set by comprehensively applying a clustering algorithm and anomaly detection model, thereby improving the ability to recognize different types of poisoned samples, being applicable to different types of text data sets, and being able to recognize poisoned samples of different types and levels, including triggers at the word level, sentence level, and more complex grammatical styles and semantic levels, with a low false alarm rate and fast processing speed, and can effectively defend against text backdoor attacks.
[0084] As a refinement of step 102, when preprocessing the sample data set to obtain the first data set, it can be implemented in the following ways but not limited to: gradually splitting the sample data set into word levels; using a pre-trained word embedding model to map each word into a dense vector of a fixed dimension; using nonlinear dimensionality reduction technology to embed and map vectors higher than a preset dimension into a space lower than a preset dimension.
[0085] For the input sample dataset, first split it into sentence level and then further subdivide it into word level. This step can use standard word segmentation tools (such as Natural Language Toolkit (NLTK), spaCy, etc.) to adapt to different language structures, and remove stop words, punctuation and other non-alphabetic characters to reduce noise and improve model performance. Use a pre-trained word embedding model (such as Word2Vec) to map each word to a dense vector of fixed dimension. Generate sentence-level or document-level vector representations through simple averaging, weighted averaging (TF-IDF weights) or more complex pooling operations (such as maximum pooling or attention mechanisms).
[0086] Considering the data sparsity and computational complexity in high-dimensional space, nonlinear dimensionality reduction techniques (such as Uniform Manifold Approximation and Projection (UMAP)) can be used to map high-dimensional word embeddings into a lower-dimensional space. According to the actual application scenario, the key parameters of UMAP, such as the number of neighbors (n_neighbors) and the minimum distance (min_dist), are set to control the size and separation of clusters.
[0087] This embodiment uses advanced dimensionality reduction technology and semantic-aware vectorization methods to retain more original data features and ensure that the system can still operate stably in the face of complex and changing attacks.
[0088] As a refinement of step 103, when performing cluster analysis on the first data set to generate a cluster analysis result including a cluster identifier to which each abnormal point belongs, the following methods may be used but are not limited to: Figure 3 As shown, Figure 3 A flowchart of a data processing process provided by an embodiment of the present disclosure includes:
[0089] Step 201 : using a density-based clustering algorithm, determine the optimal number of clusters for the first data set and process noise points, where the noise points are points far away from any cluster.
[0090] Apply the density-based clustering algorithm (Hierarchical Density-Based Spatial Clustering of Applications with Noise, HDBSCAN) on the reduced-dimensional data to automatically determine the optimal number of clusters and process noise points. HDBSCAN builds clusters by defining "core points" (i.e., points with a sufficient number of neighboring points in their neighborhood). Set appropriate minimum number (min_cluster_size) and minimum number of neighbors (min_samples) parameters according to the actual application scenario to balance the number and quality of clusters. Points far away from any cluster are regarded as noise, i.e., potential poisoning samples, and are labeled 1; normal samples are labeled 0.
[0091] Step 202: determining the main clusters belonging to normal samples in the first data set according to the clustering results obtained by the density-based clustering algorithm.
[0092] From the HDBSCAN clustering results, we select the main clusters that are obviously normal samples. These clusters are usually larger in size, higher in density, and have higher similarity between samples within them. By analyzing these clusters, we can determine which clusters are mainly composed of normal samples.
[0093] Step 203 , agglomerative hierarchical clustering is performed separately for each selected cluster in the main clusters to obtain cluster analysis results.
[0094] For each selected cluster, Ward's agglomerative hierarchical clustering is performed separately to further refine the internal structure. Ward's is a hierarchical clustering algorithm that measures the distance between clusters by calculating the sum of squares between samples within a cluster, thereby merging the closest clusters until the stopping condition is met.
[0095] As a refinement of step 203, when performing agglomerative hierarchical clustering separately for each selected cluster in the main cluster to obtain the cluster analysis result, it can be achieved by but not limited to the following methods: calculating the average distance and / or similarity between internal samples of each selected cluster, and determining the internal consistency of each selected cluster based on the average distance and / or similarity; and obtaining the cluster analysis result by comparing the internal consistency of each selected cluster with a preset threshold.
[0096] The average distance between the internal samples of each selected cluster can be determined by distance calculation methods such as Euclidean distance, Manhattan distance, and cosine similarity. The smaller the distance, the higher the internal consistency. Correspondingly, the similarity between the internal samples of each selected cluster can be determined by similarity calculation methods such as Pearson correlation coefficient and Spearman rank correlation coefficient. The higher the similarity, the stronger the internal consistency. In order to link internal consistency with clustering quality, one or more preset thresholds can be set, which should be set according to data characteristics and clustering purposes. If the average distance is less than a certain threshold, the cluster is considered to be internally consistent; if the average similarity is greater than a certain threshold, the cluster is considered to be internally consistent. Based on the evaluation results, the clustering quality is evaluated and each selected cluster is marked.
[0097] By comprehensively applying multiple clustering algorithms, the ability to identify different types of poisoning samples can be improved, especially those non-intrusive triggers that are highly concealed and difficult to detect.
[0098] As a refinement of the above embodiment, when obtaining the cluster analysis result by comparing the internal consistency of each selected cluster with the preset threshold, it can be implemented in but not limited to the following way: if the internal consistency of the cluster to be marked is lower than the preset threshold, it is determined that there are abnormal points in the cluster to be marked, and the cluster to be marked is marked as an abnormal sample.
[0099] For each cluster, the average distance or similarity between its internal samples is calculated. A reasonable preset threshold is set for the internal consistency of the cluster based on actual conditions or historical experience. When the internal consistency of a cluster is lower than the threshold, it is considered that there are abnormal points in the cluster and marked (normal sample label is 0, abnormal sample label is 1).
[0100] As a refinement of step 104, when executing the training of the anomaly detection model, it can be implemented in but not limited to the following ways: according to the clustering analysis results, the first data set is filtered to obtain the second data set; the isolation forest model is trained based on the second data set; the first data set is predicted using the pre-trained isolation forest model, and the data set marked as normal by the pre-trained isolation forest model is determined as the third data set; and the support vector machine model is trained based on the third data set.
[0101] Isolation Forest is an anomaly detection algorithm based on random forests, which is used to quickly mark potential anomalies in a dataset. Through the screening of the Isolation Forest model, the scope of potential poisoning data in the dataset can be further narrowed. The Isolation Forest model is trained on a clean dataset (the second dataset) after clustering analysis and filtering of the first dataset, and the trained model is used to predict all the original data (including data not marked as abnormal by HDBSCAN) to identify potential anomalies. Data points that are considered abnormal by the Isolation Forest are marked (labeled as 1), and these data points are considered to be potential poisoning samples. Samples marked as normal (labeled as 0) are retained to form a smaller and cleaner dataset (the third dataset) for further verification in the next step. According to the actual application scenario, adjust the key parameters of the Isolation Forest, such as the number of trees (n_estimators), sample ratio (contamination), etc., to optimize the detection performance.
[0102] One-Class SVM is a support vector machine algorithm used to train a model to identify anomalies that are different from normal samples. Training the One-Class SVM model on the normal sample subset (the third dataset) screened by the isolation forest ensures that the model only learns the characteristics of normal samples.
[0103] As a refinement of the above embodiment, when using a trained anomaly detection model to screen and verify the first data set to obtain anomaly detection results, it can be implemented in but not limited to the following ways: using a pre-trained isolation forest model to predict the first data set, and marking the sample data in the first data set according to the prediction results to obtain a first marking result; using a trained support vector machine model to predict the first data set, and marking the sample data in the first data set according to the prediction results to obtain a second marking result; determining the anomaly detection result based on the first marking result and the second marking result.
[0104] By using the trained One-Class SVM model to predict the first data set (including those samples that are not marked as abnormal by the isolation forest), and marking the prediction results (normal sample label is 0, abnormal sample label is 1), combining the results of the isolation forest and One-Class SVM, when a sample is marked as abnormal by both models, the sample marked as abnormal is finally obtained. Anomaly detection results.
[0105] The disclosed embodiment proposes a data processing method that combines clustering algorithms such as HDBSCAN and Ward's hierarchical clustering, and anomaly detection models such as isolation forest and One-Class SVM, for identifying and filtering backdoor attack samples in text. The method of this embodiment improves the ability to identify different types of poisoned samples through multi-level and multi-angle analysis. At the same time, it optimizes the detection strategy, reduces the misjudgment of normal samples, and maintains the high availability and practicality of the system.
[0106] In order to better understand the data processing method in this embodiment, Figure 4 As shown, Figure 4 A schematic flow chart of an example provided by the present disclosure, first, in order to process poisoning data, words or text data can be mapped into vector form through word embedding mapping, so as to facilitate subsequent processing and analysis. Then, the pre-trained model is used to further map the words into vectors, which may have higher dimensions. In order to reduce the computational complexity and improve the clustering effect, the next step can be dimensionality reduction, that is, the high-dimensional vector is reduced to a low-dimensional space through the UMAP algorithm, so as to better visualize and analyze the data. After the dimensionality reduction process, the main clusters and noise points in the data are identified by the HDBSCAN algorithm. Through HDBSCAN clustering, similar samples in the data set can be clustered together to form different clusters, and noise points that are not similar to most data are identified at the same time. Then, Ward's method is used to refine each main cluster, further subdivide the main clusters, and improve the accuracy and stability of clustering. After the refinement process, an isolated forest is used for preliminary screening to further narrow the scope of potential poisoning data. Then, the relatively clean and reliable normal sample subset screened out in the previous step is further verified using One-Class SVM. Through this step of verification, we can further confirm which samples are poisoned and which are normal. Finally, based on the verification results of One-Class SVM, we remove the samples confirmed to be poisoned and save the clean data set. In this way, after a series of processing and analysis, we finally get a relatively clean and reliable target data set, ensuring the high availability and practicality of the system.
[0107] In some specific embodiments, the above data processing method can be applied to a social media comment dataset. Social media platforms are one of the common targets of text backdoor attacks because they contain a large amount of user-generated content. Attackers may embed specific patterns or triggers in comments, causing machine learning models to exhibit abnormal behavior when encountering these patterns. This embodiment aims to detect and filter out poisoned social media comments. First, a dataset of 10,000 social media comments was collected, of which about 5% of the data were artificially injected with backdoor samples; each word was mapped to a 300-dimensional vector using a pre-trained Word2Vec model; for each sentence, a weighted average strategy (based on TF-IDF weights) was used to generate a sentence-level vector representation, and then UMAP was used to reduce the high-dimensional word embedding to a 2-dimensional space, with parameters set to n_neighbors=15 and min_dist=0.1. In the cluster analysis of the data, min_cluster_size=10 and min_samples=5 were set, and the main clusters and noise points were identified by HDBSCAN; Ward's hierarchical clustering was performed separately for each main cluster selected by HDBSCAN to further refine the internal structure and obtain the cluster analysis result dateset_cluster. n_estimators=100 and contamination=0.05 were configured in the isolation forest to quickly mark potential anomalies, and then One-Class SVM was used for in-depth verification, including selecting a normal sample subset for training, selecting the RBF kernel function, and setting the parameters nu=0.05 and gamma='scale'. Combining the results of the isolation forest and One-Class SVM, when a sample is marked as abnormal by both models at the same time, the sample marked as abnormal is finally obtained. The anomaly detection results are compared with the cluster analysis results and the anomaly detection results. The abnormal samples with label 1, i.e., poisoning samples, are merged, and the comments confirmed as poisoning are removed. The remaining is a clean data set. The data processing results of this embodiment are as follows: true positive rate (TPR): 98%, that is, 98% of the backdoor samples are correctly identified; false positive rate (FPR): 2%, that is, only 2% of the normal samples are misjudged as abnormal; processing speed: the entire process takes about 10 minutes and is completed within a reasonable time frame.
[0108] In some specific embodiments, the above data processing method can also be applied to an email communication dataset. Email communication involves sensitive information and is an important commercial and private communication channel. Therefore, it is crucial to protect email data from backdoor attacks. This embodiment aims to improve the security of the email system. First, a dataset containing 20,000 emails is collected. On average, each email contains about 200 words, of which about 2% of the data is artificially injected with backdoor samples; each word is mapped to a 300-dimensional vector using the FastText model (WikiNews corpus); for each email, a simple average strategy is used to generate an email-level vector representation, and then UMAP is used to reduce the high-dimensional word embedding to a 2-dimensional space, with parameters set to n_neighbors=10 and min_dist=0.1. In the clustering analysis process of the data, min_cluster_size=5 and min_samples=3 are set, and the main clusters and noise points are identified by HDBSCAN; Ward's hierarchical clustering is performed separately for each main cluster selected by HDBSCAN to further refine the internal structure and obtain the clustering analysis result dateset_cluster. In the isolation forest, n_estimators=150 and contamination=0.02 are configured to quickly mark potential anomalies, and then One-Class SVM is used for in-depth verification, including selecting a subset of normal samples for training, selecting the RBF kernel function, and setting the parameters to nu=0.02 and gamma='scale'. Combining the results of the isolation forest and One-Class SVM, when a sample is marked as abnormal by both models at the same time, the sample marked as abnormal is the total anomaly detection result. Comparing the cluster analysis results and the anomaly detection results, the abnormal samples with label 1, i.e., the poisoned samples, are merged, and the emails confirmed as poisoned are removed, and the remaining is a clean data set. The data processing results of this embodiment are as follows: True Positive Rate (TPR): 96%, i.e., 96% of the backdoor samples are correctly identified; False Positive Rate (FPR): 1%, i.e., only 1% of the normal samples are misjudged as abnormal; Processing speed: The entire process takes about 15 minutes and is completed within a reasonable time frame.
[0109] Through the above specific embodiments, it can be seen that the method of this embodiment has shown good performance on different types of text data sets: in each specific embodiment, the true positive rate exceeds 96%, indicating that the method of this embodiment can effectively identify most backdoor samples; the false positive rate is controlled below 2%, reducing the interference with normal data; although the data set is large, the entire process can be completed within a reasonable time and is suitable for actual application scenarios; it can be seen that the method of this embodiment is not only suitable for shorter texts such as social media comments, but also for longer texts such as email communications, and has wide applicability.
[0110] In summary, the embodiments of the present disclosure can achieve the following effects:
[0111] The disclosed embodiment utilizes nonlinear dimensionality reduction techniques such as UMAP to reduce the data dimension while retaining the features of the original data to the maximum extent, ensuring that the system can still operate stably in the face of complex and changeable attacks. It is not only applicable to short texts (such as social media comments), but also to long texts (such as news articles, email communications), demonstrating wide applicability and flexibility. By integrating clustering algorithms such as HDBSCAN (density-based clustering), Ward's hierarchical clustering, and anomaly detection models such as isolation forest and One-Class SVM, it is possible to identify poisoning samples of different types and levels, including word-level, sentence-level, and more complex grammatical styles and semantic-level triggers. By setting reasonable parameters and thresholds, and combining consistency evaluations of multiple models, the misjudgment of normal samples is effectively reduced, maintaining the high availability and practicality of the system.
[0112] Corresponding to the above data processing method, the present invention also provides a data processing device. Since the device embodiment of the present invention corresponds to the above method embodiment, details not disclosed in the device embodiment can be referred to the above method embodiment, and will not be repeated in the present invention.
[0113] Figure 5 A structural diagram of a data processing device provided by an embodiment of the present disclosure is shown in FIG. Figure 5 As shown, including:
[0114] An acquisition module 31 is used to acquire a sample data set;
[0115] A processing module 32, configured to preprocess the sample data set to obtain a first data set;
[0116] An analysis module 33, configured to perform cluster analysis on the first data set to generate a cluster analysis result including a cluster identifier to which each outlier belongs;
[0117] A detection module 34, configured to screen and verify the first data set using the trained anomaly detection model to obtain an anomaly detection result;
[0118] A determination module 35, configured to determine the poisoned samples in the sample data set by comparing the cluster analysis result and the anomaly detection result;
[0119] The removal module 36 is used to remove the poisoned sample from the sample data set to obtain a target data set.
[0120] The data processing device provided by the present disclosure first obtains a sample data set; pre-processes the sample data set to obtain a first data set; then performs cluster analysis on the first data set to generate a cluster analysis result including a cluster identifier to which each abnormal point belongs; then uses a trained anomaly detection model to screen and verify the first data set to obtain an anomaly detection result; determines the poisoned samples in the sample data set by comparing the cluster analysis result and the anomaly detection result; removes the poisoned samples from the sample data set to obtain a target data set. Compared with the related art, the embodiment of the present disclosure cleans the sample data set into a target data set by comprehensively applying a clustering algorithm and anomaly detection model, thereby improving the ability to identify different types of poisoned samples, being applicable to different types of text data sets, and being able to identify poisoned samples of different types and levels, including triggers at the word level, sentence level, and more complex grammatical styles and semantic levels, with a low false alarm rate and fast processing speed, and can effectively defend against text backdoor attacks.
[0121] Furthermore, in a possible implementation of the embodiment of the present disclosure, as Figure 5 As shown, the analysis module 33 is also used to use a density-based clustering algorithm to determine the optimal number of clusters of the first data set and process noise points, where the noise points are points far away from any cluster; based on the clustering results obtained by the density-based clustering algorithm, the main clusters belonging to normal samples in the first data set are determined; and agglomerative hierarchical clustering is performed separately for each selected cluster in the main clusters to obtain the cluster analysis results.
[0122] Furthermore, in a possible implementation of the embodiment of the present disclosure, as Figure 5 As shown, the analysis module 33 is also used to calculate the average distance and / or similarity between the internal samples of each selected cluster, and determine the internal consistency of each selected cluster based on the average distance and / or the similarity; and obtain the cluster analysis result by comparing the internal consistency of each selected cluster with a preset threshold.
[0123] Furthermore, in a possible implementation of the embodiment of the present disclosure, as Figure 5 As shown, the detection module 34 is further configured to determine that an abnormal point exists in the cluster to be marked if the internal consistency of the cluster to be marked is lower than the preset threshold, and mark the cluster to be marked as an abnormal sample.
[0124] Furthermore, in a possible implementation of the embodiment of the present disclosure, as Figure 5As shown, the anomaly detection model includes an isolation forest model and a support vector machine model; accordingly, the detection module 34 is also used to filter the first data set according to the cluster analysis result to obtain a second data set; train the isolation forest model based on the second data set; use the pre-trained isolation forest model to predict the first data set, and determine the data set marked as normal by the pre-trained isolation forest model as a third data set; and train the support vector machine model based on the third data set.
[0125] Furthermore, in a possible implementation of the embodiment of the present disclosure, as Figure 5 As shown, the detection module 34 is also used to use the pre-trained isolation forest model to predict the first data set, and mark the sample data in the first data set according to the prediction result to obtain a first marking result; use the trained support vector machine model to predict the first data set, and mark the sample data in the first data set according to the prediction result to obtain a second marking result; determine the anomaly detection result based on the first marking result and the second marking result.
[0126] Furthermore, in a possible implementation of the embodiment of the present disclosure, as Figure 5 As shown, the processing module 32 is also used to gradually split the sample data set into word levels; use a pre-trained word embedding model to map each word into a dense vector of a fixed dimension; use nonlinear dimensionality reduction technology to embed and map vectors higher than a preset dimension into a space lower than a preset dimension.
[0127] It should be noted that the above explanation of the method embodiment is also applicable to the device of the embodiment of the present disclosure, and the principle is the same, which is no longer limited in the embodiment of the present disclosure.
[0128] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device and a readable storage medium.
[0129] Figure 6 A schematic block diagram of an example electronic device 400 that can be used to implement an embodiment of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or required herein.
[0130] like Figure 6 As shown, the device 400 includes a computing unit 401, which can perform various appropriate actions and processes according to a computer program stored in a ROM (Read-Only Memory) 402 or a computer program loaded from a storage unit 408 to a RAM (Random Access Memory) 403. In the RAM 403, various programs and data required for the operation of the device 400 can also be stored. The computing unit 401, the ROM 402, and the RAM 403 are connected to each other via a bus 404. An I / O (Input / Output) interface 405 is also connected to the bus 404.
[0131] A number of components in the device 400 are connected to the I / O interface 405, including: an input unit 406, such as a keyboard, a mouse, etc.; an output unit 407, such as various types of displays, speakers, etc.; a storage unit 408, such as a disk, an optical disk, etc.; and a communication unit 409, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 409 allows the device 400 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0132] The computing unit 401 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 401 include, but are not limited to, a CPU (Central Processing Unit), a GPU (Graphic Processing Units), various dedicated AI (Artificial Intelligence) computing chips, various computing units running machine learning model algorithms, a DSP (Digital Signal Processor), and any appropriate processor, controller, microcontroller, etc. The computing unit 401 performs the various methods and processes described above, such as a data processing method. For example, in some embodiments, the data processing method may be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit 408. In some embodiments, part or all of the computer program may be loaded and / or installed on the device 400 via the ROM 402 and / or the communication unit 409. When the computer program is loaded into the RAM 403 and executed by the computing unit 401, one or more steps of the method described above may be performed. Alternatively, in other embodiments, the computing unit 401 may be configured to execute the aforementioned data processing method in any other appropriate manner (for example, by means of firmware).
[0133] Various embodiments of the systems and techniques described above herein may be implemented in digital electronic circuit systems, integrated circuit systems, FPGAs (Field Programmable Gate Arrays), ASICs (Application-Specific Integrated Circuits), ASSPs (Application Specific Standard Products), SOCs (System On Chips), CPLDs (Complex Programmable Logic Devices), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include: being implemented in one or more computer programs that may be executed and / or interpreted on a programmable system including at least one programmable processor that may be a dedicated or general-purpose programmable processor that may receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0134] The program code for implementing the method of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that the program code, when executed by the processor or controller, enables the functions / operations specified in the flow chart and / or block diagram to be implemented. The program code may be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.
[0135] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or device, or any suitable combination of the foregoing. More specific examples of machine-readable storage media may include electrical connections based on one or more lines, portable computer disks, hard disks, RAM, ROM, EPROM (Electrically Programmable Read-Only-Memory) or flash memory, optical fiber, CD-ROM (Compact Disc Read-Only Memory), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0136] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (Cathode-Ray Tube) or LCD (Liquid Crystal Display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0137] The systems and techniques described herein may be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: LAN (Local Area Network), WAN (Wide Area Network), the Internet, and blockchain networks.
[0138] A computer system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The relationship between the client and the server is generated by computer programs running on the corresponding computers and having a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system to solve the defects of difficult management and weak business scalability in traditional physical hosts and VPS services ("Virtual Private Server", or "VPS" for short). The server may also be a server of a distributed system, or a server combined with a blockchain.
[0139] It should be noted that artificial intelligence is a discipline that studies how computers can simulate certain human thought processes and intelligent behaviors (such as learning, reasoning, thinking, planning, etc.), and includes both hardware-level and software-level technologies. Artificial intelligence hardware technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, and big data processing; artificial intelligence software technologies mainly include computer vision technology, speech recognition technology, natural language processing technology, as well as machine learning / deep learning, big data processing technology, knowledge graph technology, and other major directions.
[0140] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps recorded in this disclosure can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this document does not limit this.
[0141] The above specific implementations do not constitute a limitation on the protection scope of the present disclosure. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent substitution and improvement made within the spirit and principle of the present disclosure shall be included in the protection scope of the present disclosure.
Claims
1. A data processing method, characterized in that: include: Get a sample dataset; Preprocessing the sample data set to obtain a first data set; Performing cluster analysis on the first data set to generate a cluster analysis result including a cluster identifier to which each outlier belongs; Using the trained anomaly detection model, screening and verifying the first data set to obtain anomaly detection results; Determine the poisoned samples in the sample data set by comparing the cluster analysis result and the anomaly detection result; The poisoned sample is removed from the sample data set to obtain a target data set.
2. The method according to claim 1, characterized in that The step of performing cluster analysis on the first data set to generate a cluster analysis result including a cluster identifier to which each outlier belongs includes: Using a density-based clustering algorithm, determining the optimal number of clusters for the first data set and processing noise points, wherein the noise points are points far away from any cluster; Determining, according to the clustering results obtained by the density-based clustering algorithm, a main cluster belonging to normal samples in the first data set; Agglomerative hierarchical clustering is performed separately for each selected cluster in the main clusters to obtain the cluster analysis result.
3. The method according to claim 2, characterized in that The step of performing agglomerative hierarchical clustering separately for each selected cluster in the main clusters to obtain the cluster analysis result includes: Calculating an average distance and / or similarity between internal samples of each selected cluster, and determining the internal consistency of each selected cluster according to the average distance and / or the similarity; The cluster analysis result is obtained by comparing the internal consistency of each selected cluster with a preset threshold.
4. The method according to claim 3, characterized in that The obtaining of the cluster analysis result by comparing the internal consistency of each selected cluster with a preset threshold comprises: If the internal consistency of the cluster to be marked is lower than the preset threshold, it is determined that there are abnormal points in the cluster to be marked, and the cluster to be marked is marked as an abnormal sample.
5. The method according to claim 1, characterized in that The anomaly detection model includes an isolation forest model and a support vector machine model; The training process of the anomaly detection model includes: According to the cluster analysis result, filtering the first data set to obtain a second data set; Training the isolation forest model based on the second data set; Using a pre-trained isolation forest model to predict the first data set, and determining the data set marked as normal by the pre-trained isolation forest model as a third data set; The support vector machine model is trained based on the third data set.
6. The method according to claim 5, characterized in that The first data set is screened and verified by using the trained anomaly detection model to obtain an anomaly detection result including: Using the pre-trained isolation forest model to predict the first data set, and marking sample data in the first data set according to the prediction result to obtain a first marking result; Using the trained support vector machine model to predict the first data set, and marking sample data in the first data set according to the prediction result to obtain a second marking result; The abnormality detection result is determined according to the first marking result and the second marking result.
7. The method according to claim 1, characterized in that The preprocessing of the sample data set to obtain a first data set comprises: Stepwise splitting the sample data set into word levels; Use the pre-trained word embedding model to map each word into a dense vector of fixed dimension; Using nonlinear dimensionality reduction techniques, vector embedding with dimensions higher than a preset dimension is mapped into a space with dimensions lower than a preset dimension.
8. A data processing device, characterized in that: include: An acquisition module is used to obtain a sample data set; A processing module, used for preprocessing the sample data set to obtain a first data set; An analysis module, configured to perform cluster analysis on the first data set to generate a cluster analysis result including a cluster identifier to which each outlier belongs; A detection module, used to screen and verify the first data set using the trained anomaly detection model to obtain an anomaly detection result; A determination module, used to determine the poisoned samples in the sample data set by comparing the cluster analysis result and the anomaly detection result; The removal module is used to remove the poisoned sample from the sample data set to obtain a target data set.
9. An electronic device, characterized in that: include: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 7.
10. A non-transitory computer-readable storage medium storing computer instructions, characterized in that: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-7.