Data cleaning method and device and storage medium
By using a generative adversarial network model based on dynamic detection thresholds in data cleaning, the problems of low efficiency and poor adaptability of toxic text data in the prior art are solved, and the cleaning and identification effects of high-quality data are achieved.
Patent Information
- Application Number
- CN202311706361.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-12
- Publication Date
- 2025-06-13
AI Technical Summary
When processing toxic text data, the prior art has problems such as low recognition efficiency, lack of information and inability to adapt to dynamic changes in the network.
A generative adversarial network model based on dynamic detection threshold is adopted to generate cleaned data through mutual game and learning between generators and discriminators, improve data quality and enhance recognition capabilities.
It realizes efficient repair of dirty data, improves the quality of original data, significantly improves the effect of identifying dirty data, and adapts to dynamic changes in the network.
Smart Images

Figure CN120144568A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the fields of artificial intelligence and data processing, and particularly to a data cleaning method, apparatus, and storage medium. Background Art
[0002] Data quality problems generally exist in all walks of life. On the one hand, with the popularization and development of the Internet, there is more and more information on the Internet, including a large number of toxic texts, which have a great negative impact on the network environment and the mental health of users. On the other hand, with the rapid development of the artificial intelligence (AI) industry, the scale of the dataset used for training also shows an exponential upward trend. However, the collected data may have problems such as missing, noisy, and duplicate data, and cannot be directly used for large models. Instead, it needs to be cleaned, labeled, etc. before a dataset for large models can be generated.
[0003] In related technologies, the methods for cleaning toxic text data can be roughly divided into: rule-based cleaning, statistics-based cleaning, and deep learning-based cleaning methods, but they have various problems. For example, rule-based and statistics-based cleaning methods ignore the semantic information of the text and are not suitable for complex situations. Deep learning-based cleaning methods generally require a large amount of labeled data and model training time. However, it is difficult to obtain toxic text data. In the absence of corresponding data, the model cleaning efficiency is low. In addition, it is also difficult to process unknown types of toxic text.
[0004] In addition, the general processing method for existing text cleaning methods after identifying toxic data is filtering operations such as deletion, but this will cause information loss and incompleteness of the text to a certain extent. Even deleting toxic data when the data volume is insufficient will make the data volume even more lacking. Summary of the Invention
[0005] To overcome the problems in related technologies, the present disclosure provides a data cleaning method, apparatus, and storage medium.
[0006] According to the first aspect of the embodiments of the present disclosure, a data cleaning method is provided, and the method includes:
[0007] Obtain the dataset to be cleaned; call a preset generative adversarial network model to process the dataset to be cleaned to obtain the cleaned data. Among them, the generator in the generative adversarial network model is used to generate the cleaned data, and the discriminator in the generative adversarial network model is used to determine the probability that the cleaned data generated by the generator is the target type data, and the target type data is dirty data or clean data; in the training stage of the generative adversarial network model, the generator is used to generate the second type of data based on the first type of data, the first type of data is the sample data to be cleaned, and the second type of data is the sample pseudo data that makes the discriminator recognize as the target type data; in the training stage of the generative adversarial network model, the discriminator is used to distinguish the second type of data and the first type of data to determine that the second type of data is not the target type data.
[0008] In one implementation, the discriminator determines that the second type of data is not the target type data based on a dynamic detection threshold, so as to determine that the second type of data is not the target type data.
[0009] In one implementation, the discriminator uses the following method to distinguish the second type of data and the first type of data based on a dynamic detection threshold, so as to determine that the second type of data is not the target type data: input the second type of data and the first type of data into the discriminator; determine the recognition accuracy of the discriminator, and the recognition accuracy is the accuracy of the discriminator to distinguish the second type of data and the first type of data based on the first detection threshold; adjust the first detection threshold based on the recognition accuracy to obtain the second detection threshold; distinguish the second type of data and the first type of data based on the second detection threshold; iteratively execute the above process until the discriminator determines that the second type of data is not the target type data.
[0010] In one implementation, determining the recognition accuracy of the discriminator includes: determining the first quantity of the first type of data determined by the discriminator as the target type data based on the first detection threshold, and determining the second quantity of the second type of data determined by the discriminator as the target type data based on the first detection threshold; determining the sum value of the first quantity and the second quantity, and determining the ratio of the sum value to the third quantity as the accuracy of the discriminator to distinguish the second type of data and the first type of data based on the first detection threshold; the third quantity is the total quantity of the first type of data and the second type of data.
[0011] In one implementation, adjusting the first detection threshold based on the recognition accuracy rate to obtain a second detection threshold includes: if the recognition accuracy rate is greater than or equal to the accuracy rate threshold, using the first detection threshold as the second detection threshold; if the recognition accuracy rate is less than the accuracy rate threshold, determining the second detection threshold based on the first detection threshold and the recognition accuracy rate. In one implementation, the following formula is used to determine the second detection threshold based on the first detection threshold and the recognition accuracy rate: T t = T t-1 + α * (Acc - T t-1 ); where t represents the current iteration number, T is a preset initial threshold, α is the learning rate, and Acc is the current model training accuracy rate.
[0012] In one implementation, an original data set obtained by collection is acquired. The original data set includes multiple pieces of original data, and the format of each piece of original data is text; the original data set is preprocessed to obtain a data set to be cleaned. The data set to be cleaned includes multiple pieces of data to be cleaned, and the format of each piece of data to be cleaned is a vector.
[0013] In one implementation, preprocessing the original data set to obtain a data set to be cleaned includes: performing first preprocessing on the original data set to obtain a first data set; the first preprocessing includes any one or more of removing stop words, removing special characters, and removing garbled characters; performing word segmentation processing on the first data set to obtain a second data set; and vectorizing the text data in the second data set to obtain the data set to be cleaned.
[0014] According to the second aspect of the embodiments of the present disclosure, a data cleaning device is provided. The device includes:
[0015] An acquisition unit for acquiring a data set to be cleaned; an execution unit for calling a preset generative adversarial network model to process the data set to be cleaned to obtain cleaned data. In the generative adversarial network model, the generator is used to generate the cleaned data, and the discriminator in the generative adversarial network model is used to determine the probability that the cleaned data generated by the generator is target type data. The target type data is dirty data or clean data; in the training stage of the generative adversarial network model, the generator is used to generate a second type of data based on a first type of data. The first type of data is sample data to be cleaned, and the second type of data is sample pseudo data that makes the discriminator recognize as the target type data; in the training stage of the generative adversarial network model, the discriminator is used to distinguish the second type of data from the first type of data to determine the second type of data as non-target type data.
[0016] In one embodiment, the discriminator determines the second type of data as non-target type data based on a dynamic detection threshold, so as to determine the second type of data as non-target type data. In one embodiment, the discriminator, through an execution unit, adopts the following method to distinguish the second type of data and the first type of data based on a dynamic detection threshold, so as to determine the second type of data as non-target type data: input the second type of data and the first type of data into the discriminator of the generative adversarial network prediction model; determine the recognition accuracy rate of the discriminator, where the recognition accuracy rate is the accuracy rate of the discriminator to distinguish the second type of data and the first type of data based on a first detection threshold; dynamically adjust the first detection threshold based on the recognition accuracy rate to obtain a second detection threshold; distinguish the second type of data and the first type of data based on the second detection threshold; iteratively execute the above process until the discriminator determines the second type of data as non-target type data.
[0017] In one embodiment, the execution unit adopts the following method to determine the recognition accuracy rate of the discriminator: determine a first quantity that the discriminator determines the first type of data as target type data based on the first detection threshold, and determine a second quantity that the discriminator determines the second type of data as target type data based on the first detection threshold; determine the sum value of the first quantity and the second quantity, and determine the ratio of the sum value to a third quantity as the accuracy rate of the discriminator to distinguish the second type of data and the first type of data based on the first detection threshold; the third quantity is the total quantity of the first type of data and the second type of data.
[0018] In one embodiment, the execution unit adopts the following method to dynamically adjust the first detection threshold based on the accuracy rate to obtain a second detection threshold: if the recognition accuracy rate is greater than or equal to an accuracy rate threshold, then use the first detection threshold as the second detection threshold; if the recognition accuracy rate is less than the accuracy rate threshold, then determine the second detection threshold based on the first detection threshold and the recognition accuracy rate.
[0019] In one embodiment, the execution unit adopts the following formula to determine the second detection threshold based on the first detection threshold and the recognition accuracy rate: T t =T t-1 +α*(Acc - T t-1 ); where t represents the current iteration number, T is a preset initial threshold, α is a learning rate, and Acc is the current model training accuracy rate.
[0020] In one embodiment, the obtaining unit obtains the dataset to be cleaned in the following manner, including: obtaining the original dataset collected, where the original dataset includes multiple pieces of original data, and the format of each piece of original data is text; preprocessing the original dataset to obtain the dataset to be cleaned, where the dataset to be cleaned includes multiple pieces of data to be cleaned, and the format of each piece of data to be cleaned is a vector.
[0021] In one embodiment, the obtaining unit preprocesses the original dataset in the following manner to obtain the dataset to be cleaned, including: performing a first preprocessing on the original dataset to obtain a first dataset; the first preprocessing includes: removing any one or more of stop words, special characters, and garbled characters; performing word segmentation on the first dataset to obtain a second dataset; vectorizing the text data in the second dataset to obtain the dataset to be cleaned.
[0022] According to a third aspect of the embodiments of the present disclosure, there is provided a data cleaning device, including: a processor; a memory for storing instructions executable by the processor; wherein, the processor is configured to: execute the data cleaning method described in the first aspect or any one of the embodiments in the first aspect.
[0023] According to a fourth aspect of the embodiments of the present disclosure, there is provided a storage medium, in which instructions are stored, and when the instructions in the storage medium are executed by a processor of a terminal, the terminal can execute the data cleaning method described in the first aspect or any one of the embodiments in the first aspect.
[0024] The technical solutions provided by the embodiments of the present disclosure may include the following beneficial effects: Through the present disclosure, a generative adversarial network model is trained based on a dynamic detection threshold, and the data to be cleaned is passed through the trained generative adversarial network model to generate the cleaned data, that is, the dirty data can be repaired to obtain clean data, improving the quality of the original data and significantly enhancing the effect in identifying dirty data.
[0025] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] The accompanying drawings herein are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure.
[0027] Figure 1 is a flowchart of a data cleaning method shown according to an exemplary embodiment.
[0028] Figure 2It is a framework diagram of a GAN model shown according to an exemplary embodiment.
[0029] Figure 3 It is a generative adversarial network model built with Keras shown according to an exemplary embodiment.
[0030] Figure 4 It is a flowchart of training a generative adversarial network model based on a dynamic detection threshold shown according to an exemplary embodiment.
[0031] Figure 5 It is a flowchart of determining the accuracy of a discriminator in distinguishing between a second type of data and a first type of data based on a first detection threshold shown according to an exemplary embodiment.
[0032] Figure 6 It is a flowchart of dynamically adjusting the first detection threshold based on the accuracy to obtain a second detection threshold shown according to an exemplary embodiment.
[0033] Figure 7 It is a flowchart of GAN-enhanced toxic data cleaning based on adaptive threshold adjustment shown according to an exemplary embodiment.
[0034] Figure 8 It is a block diagram of a data cleaning device shown according to an exemplary embodiment.
[0035] Figure 9 It is a block diagram of a device for data cleaning shown according to an exemplary embodiment.
[0036] Figure 10 It is a block diagram of a device for data cleaning shown according to an exemplary embodiment. Detailed implementation manners
[0037] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present disclosure.
[0038] Currently, data quality problems are widespread in all walks of life. On the one hand, with the popularization and development of the Internet, there is more and more information on the Internet, including a large amount of toxic text, such as cyber violence, personal attacks, discrimination, or unhealthy language, etc. These toxic texts have a great negative impact on the network environment and the mental health of users. Therefore, researching how to effectively clean toxic texts and improve the quality of the network environment has become an urgent problem to be solved.
[0039] On the other hand, with the rapid development of the AI industry, the scale of the datasets used for training has also shown an exponential upward trend. Large AI models are regarded as the new industrial revolution. New revolutionary technologies mean new innovations, and innovations mean new opportunities. As the fuel, data, in the form of high-quality, large-scale, and rich datasets, has become an essential part of the competition among large models. The most crucial part of the dataset is the data with high relevance, diversity, and quality to the model tasks. Considering that the collected data may have problems such as missing, noisy, and duplicate data, massive amounts of data cannot be directly used for large models. Instead, it needs to go through processes such as cleaning and annotation to generate a dataset that can be used by large models, and then combined with algorithms, computing power, etc., to be truly used in large models. Taking the third-generation Generative Pre-trained Transformer-3 (GPT-3) as an example, its original data volume was 45TB, while the high-quality data after cleaning was 570GB. Taking this as a reference, only about 1% of the original data after cleaning became the data in the corpus. Therefore, data will become the key to the differential competition of large models, and cleaning is essential.
[0040] Toxic text can be interpreted as text containing unreasonable and incorrect logic or spreading negative energy. Sometimes it is equivalent to dirty data, which not only reduces the quality of the data but also causes certain negative impacts. Existing toxic data cleaning methods can be roughly divided into: rule-based cleaning, statistics-based cleaning, deep learning-based cleaning, and other methods.
[0041] Rule-based cleaning methods mainly use predefined rules such as regular expressions and keyword matching to identify and process toxic text. For example, deleting sensitive words and blocking advertisements. Such methods are simple and easy to implement, but have high requirements for the writing and maintenance of rules and may not be able to adapt to the constantly changing toxic text. In addition, there may be cases of missed reports and false reports in the rules.
[0042] Statistics-based cleaning methods, for example, represent the text as a bag-of-words model, and then use statistical methods such as Term Frequency-Inverse Document Frequency (TF-IDF) and chi-square test to identify and process toxic text. Generally, statistics-based methods cannot capture the context information in the text, resulting in poor cleaning effects. In addition, this method cannot handle new words outside the vocabulary.
[0043] Model-based methods, such as deep learning models like Recurrent Neural Network (RNN) and Long Short Term Memory (LSTM), are used to identify and process toxic texts. Deep learning-based methods can automatically learn the semantic and contextual information of texts, have strong recognition capabilities for implicit aggression and malice, and have strong generalization capabilities, being able to adapt to new toxic texts. However, the training of deep neural networks requires a large amount of computing resources and time, and generally uses predefined fixed thresholds for identification, making it difficult to adapt to the ever-changing toxic data.
[0044] There are various forms and types of toxic texts, and the existing toxic text cleaning methods mainly have the following problems:
[0045] (1) Rule-based and statistical cleaning methods ignore the semantic information of texts and are not applicable to complex situations. Due to the various forms and types of toxic texts, rule-based methods have high maintenance costs and poor generalization, and cleaning methods such as those based on Latent Dirichlet Allocation (LDA) require feature extraction of a large amount of texts, with high costs and poor cleaning effects.
[0046] (2) Deep learning-based cleaning methods generally require a large amount of labeled data and model training time, while it is difficult to obtain toxic text data. In the absence of corresponding data, the model cleaning efficiency is low. In addition, it is also difficult to process unknown types of toxic texts.
[0047] (3) Existing toxic text cleaning methods generally have problems in dealing with the dynamic changes of the network. Since the forms and types of toxic texts on the network are constantly changing, existing cleaning methods cannot adapt to these changes in a timely manner, resulting in poor cleaning effects.
[0048] (4) The general processing method of existing text cleaning methods after identifying toxic data is filtering operations such as deletion, but this will cause information loss and incompleteness to a certain extent. Even deleting toxic data when the data volume is insufficient will make the data volume even more lacking.
[0049] Based on this, the embodiments of the present disclosure provide a data cleaning method, training a generative adversarial network model based on a dynamic detection threshold, and generating cleaned data from the data to be cleaned through the trained Generative Adversarial Networks (GAN) model, that is, dirty data can be repaired to obtain clean data, improving the quality of the original data and significantly enhancing the effect in identifying dirty data.
[0050] In the embodiments of the present disclosure, dirty data can be interpreted as data of types such as containing unreasonable error logic or spreading negative energy, which will reduce the quality of data. Clean data is the repaired data generated by the trained generative adversarial network model for dirty data, that is, clean data is data that does not contain types such as error logic or spreading negative energy.
[0051] Among them, in the embodiments of the present disclosure, dirty data is toxic data, and clean data is non-toxic data.
[0052] Figure 1 is a flowchart of a data cleaning method shown according to an exemplary embodiment, as Figure 1 shown, the method includes the following steps.
[0053] In step S11, obtain the dataset to be cleaned.
[0054] In the embodiments of the present disclosure, the collected data is preprocessed to obtain the data to be cleaned. Among them, the data to be cleaned may contain dirty data or may not contain dirty data.
[0055] In step S12, call the preset generative adversarial network model to process the dataset to be cleaned, and obtain the cleaned data.
[0056] In the embodiments of the present disclosure, during the usage stage of the trained model, the data to be cleaned is input into the generative adversarial network model, and the generator repairs the data to be cleaned to obtain the repaired data, that is, the cleaned data.
[0057] In the embodiments of the present disclosure, Figure 2 is a framework diagram of a GAN model shown according to an exemplary embodiment, as Figure 2 shown, the GAN framework includes a generator (Generator, G) and a discriminator (Discriminator, D). Both are deep learning networks and are inverse processes of each other. The generator G is responsible for generating a cleaned version of the toxic data, and the discriminator D is responsible for determining whether the generated data is a valid cleaned version, that is, determining whether it is toxic or non-toxic. Here, both the generator and the discriminator use a feedforward neural network model (Multilayer Perceptron, MLP). Among them, X is the data to be cleaned, Z is a random noise, and it is a random vector following a Gaussian distribution.
[0058] In the embodiments of the present disclosure, the generative adversarial network model mainly includes two modules: a generator and a discriminator. Among them, the generator is used to generate the cleaned data, and the discriminator is used to determine the probability that the cleaned data generated by the generator is the target type of data, where the target type of data is dirty data or clean data. In the training stage of the generative adversarial network model, the generator is used to generate the second type of data based on the first type of data, where the first type of data is the sample data to be cleaned, and the second type of data is the sample pseudo-data that enables the discriminator to identify as the target type of data; in the training stage of the generative adversarial network model, the discriminator is used to distinguish the second type of data from the first type of data to determine the second type of data as non-target type of data.
[0059] In the embodiments of the present disclosure, the confrontation and generation of the generative adversarial network model are reflected in that: the generator needs to try its best to make the discriminator recognize the samples generated by the generator as real, while the discriminator needs to try its best to recognize the samples generated by the generator as fake. The two are a process of mutual game and learning. The data generated by the generator is the second type of data. The discriminator inputs the first type of data and the second type of data, and by continuously adjusting the detection threshold of the discriminator, it determines whether the second type of data and the first type of data are toxic data or non-toxic data, and continuously adjusts the detection threshold to balance the data quality generated by the generator and the discrimination ability of the discriminator.
[0060] Figure 3 Shows a generative adversarial network model built with Keras according to an exemplary embodiment, as Figure 3 shown. As Figure 3 (Left) shows the structure of the generative network model: the input layer is random noise following a Gaussian distribution, with a vector dimension of 100. It passes through two Dense hidden layers with 256 and 512 neurons respectively, and at the same time uses the activation function (Rectified Linear Unit, ReLU) and the Dropout layer. Finally, a Dense layer with 1024 neurons is used as the output layer of the generator.
[0061] Figure 3 (Right) shows the structure of the discriminative network model: the input layer is the output pseudo-data of the generator, with a data dimension of 1024. It passes through two Dense hidden layers with 256 and 128 neurons respectively, uses the ReLU activation function and the Dropout layer. Finally, a Dense layer with one neuron is used, and the Sigmoid activation function is used to obtain the predicted probability value.
[0062] From the formula, the generator is responsible for generating non-toxic data, and the discriminator is responsible for distinguishing between toxic data and non-toxic data. In the initialization stage, both the generator and the discriminator are randomly parameterized neural networks. Specifically, the generator first uses random noise to generate virtual text through a series of neural network layers:
[0063] G(z) = g(z, W G )
[0064] where the random noise z is a random vector with a Gaussian distribution, and W G are the weight parameters of the generator G. At the same time, the discriminator D takes the original text x as input and judges whether the text is toxic:
[0065] D(x) = d(x, W D )
[0066] where W D are the weight parameters of the discriminator D.
[0067] The goal of the generator is to minimize the probability that the discriminator correctly classifies the generated data. In other words, the generator tries to generate non-toxic data that can "fool" the discriminator. The loss function of the generator can be expressed as:
[0068] L G = -logD(G(z))
[0069] The goal of the discriminator is to maximize the probability of correct classification. The discriminator not only has to ensure its basic classification ability, that is, accurately identifying real samples as real samples, but also has the ability to distinguish false samples. Its loss function is:
[0070] L D = -logD(x) - log(1 - D(G(z))
[0071] It can be found from the above loss functions that the calculations of the loss functions are all generated in D (the discriminator). Coupled with the fact that the output of D is generally a binary classification judgment, the binary cross-entropy function is generally adopted as a whole. In addition, the training of the generator and the discriminator is carried out alternately. In each round of training, first fix the generator and update the parameters of the discriminator to make it better identify real toxic texts and toxic texts generated by the generator. Then, fix the discriminator and update the parameters of the generator to make the toxic texts it generates more deceptive. The training process can be carried out through the gradient descent optimization algorithm. Through iterative improvement, the non-toxic data generated by the generator gets closer and closer to the real data distribution, and the classification ability of the discriminator is gradually improved.
[0072] In the embodiments of the present disclosure, the collected data is preprocessed to facilitate the processing of the deep learning network model. The GAN model can generate samples similar to real data through the generator, which helps to enhance the robustness of the model during the data cleaning process. Compared with other methods, GAN can better process different types of dirty data. Among them, in the actual use of the generative adversarial network model, the data generated by the generator is the data after cleaning the data to be cleaned, that is, most likely non-toxic data, and the real data is the data to be cleaned, which can be toxic data or non-toxic data.
[0073] In the embodiments of the present disclosure, the discriminator determines the second type of data as non-target type data based on the dynamic detection threshold, uses the dynamic detection threshold to determine the non-target type data, and dynamically adjusts the threshold according to the distribution characteristics of the data, avoiding the limitations brought by the fixed threshold, increasing the diversity of the data, and further improving the recognition ability of the generative adversarial network model.
[0074] Figure 4 It is a flowchart of training a generative adversarial network model based on a dynamic detection threshold shown according to an exemplary embodiment, as Figure 4 shown, and includes the following steps.
[0075] In step S21, the generator based on the generative adversarial network prediction model generates the second type of data.
[0076] In the embodiments of the present disclosure, during the training process, the generative adversarial network model is called the generative adversarial network prediction model, and the generator based on the generative adversarial network prediction model generates the second type of data.
[0077] In step S22, the second type of data and the first type of data are input into the discriminator of the generative adversarial network prediction model.
[0078] In the embodiments of the present disclosure, the sample pseudo data and the data to be cleaned sample are input into the discriminator of the generative adversarial network prediction model.
[0079] In step S23, determine the recognition accuracy of the discriminator to distinguish the second type of data and the first type of data based on the first detection threshold.
[0080] In the embodiments of the present disclosure, the discriminator obtains the corresponding probability values for the prediction of the input first type of data and the second type of data based on the initial detection threshold, converts the probabilities corresponding to the input first type of data and the second type of data into binary classification, and determines the recognition accuracy of the current model training based on the binary classification and the quantities of the first type of data and the second type of data, that is, the recognition accuracy of distinguishing the second type of data and the first type of data.
[0081] In step S24, dynamically adjust the first detection threshold based on the recognition accuracy to obtain the second detection threshold.
[0082] In the embodiments of the present disclosure, during the training process, the recognition accuracy (Accuracy, Acc) of the discriminator is calculated in real time, and then the detection threshold T is dynamically adjusted according to Acc to optimize the quality of the cleaned data generated by the generator. Additionally, a smaller T value can be used in the initial stage to accelerate the training speed. As the performance of the model continuously improves, the value of T is gradually increased to enhance the recognition effect of the discriminator.
[0083] In step S25, the generative adversarial network prediction model is iteratively trained based on the second detection threshold to obtain the generative adversarial network model.
[0084] In the embodiments of the present disclosure, the generative adversarial network prediction model is iteratively trained. In each iterative training, the Acc of the discriminator is calculated in real time, and then the threshold T is dynamically adjusted according to Acc to balance the text quality of the generator and the judgment ability of the discriminator, so as to obtain an optimized generative adversarial network model.
[0085] Figure 5 It is a flowchart showing the accuracy of determining that the discriminator distinguishes the second type of data and the first type of data based on the first detection threshold according to an exemplary embodiment, as Figure 5 shown, and includes the following steps.
[0086] In step S31, determine the first quantity of the discriminator determining the first type of data as the target type of data based on the first detection threshold, and determine the second quantity of the discriminator determining the second type of data as the target type of data based on the first detection threshold.
[0087] In the embodiments of the present disclosure, the target type of data is set as clean data, that is, non-toxic data.
[0088] In the embodiments of the present disclosure, determine the first quantity of the discriminator determining the first type of data as non-toxic data based on the initial detection threshold, and determine the second quantity of the discriminator determining the second type of data as non-toxic data based on the initial detection threshold.
[0089] In step S32, determine the sum value of the first quantity and the second quantity, and determine the ratio of the sum value to the third quantity as the accuracy of the discriminator distinguishing the second type of data and the first type of data based on the first detection threshold.
[0090] In the embodiments of the present disclosure, the third quantity is the total quantity of the first type of data and the second type of data.
[0091] In the embodiments of the present disclosure, let the real input data, that is, the first type of data, be x r , and the data generated by the generator, that is, the second type of data, be x g , and the total data volume, that is, the third quantity, be N = |x r | + |x g|, the discriminator predicts probability values D(x r ) and D(x g ), which respectively represent the probability that the input data is real data. According to the set threshold T, D(x r ) and D(x g ) are converted into binary classification, which is expressed by the mathematical formula:
[0092]
[0093]
[0094] Acc = (∑(y r == 1) + ∑(y g == 0)) / N
[0095] Among them, yr represents the classification result of the D prediction of the real data, yg represents the classification result of the D of the data generated by G, ∑ is the summation symbol, ∑(y r == 1) is the sum of the non-toxic data in the first type of data to obtain the first quantity, ∑(y g == 0)) is the sum of the non-toxic data in the second type of data to obtain the second quantity, N is the total of the real data and the generated data, that is, the third quantity, (∑(y r == 1) + ∑(y g == 0)) / N is the ratio of the sum of the first quantity and the second quantity to the third quantity, and then the training accuracy of the current model is obtained.
[0096] Figure 6 is a flowchart showing a method for dynamically adjusting a first detection threshold based on accuracy to obtain a second detection threshold according to an exemplary embodiment, as Figure 6 shown, including the following steps.
[0097] In step S41, the first detection threshold is dynamically adjusted based on the recognition accuracy to obtain the second detection threshold.
[0098] In step S42, if the recognition accuracy is greater than or equal to the accuracy threshold, the first detection threshold is used as the second detection threshold.
[0099] In step S43, if the recognition accuracy is less than the accuracy threshold, the second detection threshold is determined based on the first detection threshold and the recognition accuracy.
[0100] In the embodiment of the present disclosure, the following formula is used to determine the second detection threshold based on the first detection threshold and the recognition accuracy: T t = T t-1 + α * (Acc - T t-1), where \(t\) represents the current iteration number, \(T\) is a preset initial threshold, \(\alpha\) is the learning rate, and \(Acc\) is the current training accuracy of the model.
[0101] In the embodiments of the present disclosure, an identification accuracy threshold is set. If the accuracy of the current model training is greater than or equal to this accuracy threshold, the model continues to be trained with the first detection threshold.
[0102] In the embodiments of the present disclosure, if the accuracy of the current model training is less than this accuracy threshold, the second detection threshold is determined based on the first detection threshold and the identification accuracy.
[0103] In the embodiments of the present disclosure, the adaptive threshold adjustment strategy is a method of dynamically adjusting the toxic data detection threshold according to the performance metrics during the model training process. Here, the initial threshold \(T\) is set to 0.5. For the non-toxic probability output by the discriminator \(D\), if \(D(G(z))>T\), it is considered that the generated data is a valid cleaned version; otherwise, it is considered that there are still toxic components in the generated text. According to the accuracy during the training process, the threshold \(T\) is continuously adjusted to balance the text quality generated by the generator and the judgment ability of the discriminator. The specific adjustment method is as follows:
[0104] \(T\) t \(=\) t-1 \(T\) t-1 \(+\alpha\times(Acc - T)\)
[0105] where \(t\) represents the current iteration number, \(T\) is a preset initial threshold, \(\alpha\) is the learning rate, and \(Acc\) is the current training accuracy of the model, that is, the proportion of the discriminator correctly classifying real data and generated data.
[0106] In the embodiments of the present disclosure, the model is trained based on the adaptive threshold adjustment. According to the accuracy during the training process, the threshold \(T\) is continuously adjusted to balance the text quality generated by the generator and the judgment ability of the discriminator, so that the effect of this method in identifying toxic data is significantly improved.
[0107] In another embodiment of the present disclosure, during the process of training the GAN model, enhanced perturbations are introduced during the training process to improve the stability and performance of the model. Specifically, the following formula is used to calculate the perturbation:
[0108] \(x\) adv \(=x+\varepsilon f(x)\)
[0109]
[0110] Among them, G(x) is the denoised text generated by the generator, i.e., the pseudo-data. ε is the amplitude of the perturbation. Generally, the perturbation space is relatively small to avoid damaging the original samples. Here, ε is set to 0.3. f(x) is the perturbation function for the input text x. Here, the method adopted is to randomly replace each character of the input text with a perturbed character to generate a text x with perturbation. adv . After adding the perturbation, the corresponding loss functions of the generator and the discriminator also change. When training the generator G, its loss function becomes:
[0111] L G = -log(D(G(x))) - λlog(D(G(x adv )))
[0112] where D(G(x)) represents the output of the discriminator on the denoised text generated by the generator, and D(G(x adv )) represents the output of the discriminator on the text with adversarial perturbation generated by the generator. λ represents the regularization parameter, which is set to 0.1 as the penalty term.
[0113] Similarly, when training the discriminator D, the following loss function is optimized:
[0114] L D = -log(D(x)) - log(1 - D(G(x))) + λ(-log(D(x adv )) - log(1 - D(G(x adv ))))
[0115] where D(x) represents the output of the discriminator on the original input text x, D(G(x)) represents the output of the discriminator on the denoised text generated by the generator, D(x adv ) represents the output of the discriminator on the text with perturbation, and D(G(x adv )) represents the output of the discriminator on the text with perturbation generated by the generator.
[0116] In addition, in the test phase, the text in the test set is input into the generator G, and then the discriminator D is used to discriminate the pseudo-text generated by the generator. If it is a non-toxic text, it is directly output. If it is a toxic text, it can be retained as augmented data and used together with the original toxic text dataset to train the toxic text cleaning model.
[0117] In the embodiments of the present disclosure, during the GAN training process, the diversity of the generator is increased by adding noise perturbation, thereby improving the robustness and generalization ability of the model.
[0118] In the embodiments of the present disclosure, the dataset to be cleaned can be obtained in the following manner: Obtain the original dataset collected, where the original dataset can be a dataset including multiple pieces of original data, and the format of each piece of original data is text. Based on preprocessing the original dataset, obtain the dataset to be cleaned. Among them, the preprocessing is to convert the dataset to be processed in text format into the dataset to be cleaned in vector format.
[0119] In the embodiments of the present disclosure, the preprocessing of the original dataset includes the first preprocessing. Among them, the first preprocessing is to remove any one or more of stop words, special characters, and garbled characters, and obtain the first dataset based on the first preprocessing. Perform word segmentation on the first dataset to obtain the second dataset. Based on the Tokenizer class of the Keras open-source framework, vectorize each piece of text in the second dataset. Tokenizer is a class used to vectorize text or convert text into a sequence. Based on the vectorized corpus, the number of words, TF-IDF, etc., convert each text into an integer sequence. Specifically, assume the original text data is x, and the preprocessed text data is:
[0120] V = {v 1 , v 2 , …, v n}
[0121] Among them, v i is a vector with a dimension of 300, and n is the length of the processed text word sequence.
[0122] In the embodiments of the present disclosure, based on the vectorization of the text of the original dataset, the semantic generalization ability is enhanced, which is convenient for training the generative adversarial network model and further improving the recognition ability of the model.
[0123] Figure 7 is a flowchart of GAN enhanced toxic data cleaning based on adaptive threshold adjustment shown according to an exemplary embodiment. As Figure 7 shown, the present invention is divided into data preprocessing, model construction and training, and the adaptive threshold adjustment and GAN enhancement strategies adopted during training.
[0124] The present invention relates to the fields of artificial intelligence and data cleaning, particularly a method for cleaning toxic data in large models. The technical solution of the present invention includes the following aspects: First, perform data preprocessing on the original text, including word segmentation, stop word filtering, etc., and convert the data into a format for model training; Secondly, based on the GAN enhancement technology, use a generative adversarial network to transform and repair toxic data. Use an MLP as the generator, input the toxic text, and output the generated pseudo-data as non-toxic data to replace the original toxic data. The discriminator is also based on the MLP model, input the generated pseudo-non-toxic text, and output the non-toxic probability of this text. At the same time, during training, through an adaptive threshold adjustment strategy, according to the performance indicators in the model training process, dynamically adjust the threshold for toxic data detection and use the GAN enhancement technology, thereby improving the effect of data cleaning, and finally realizing a cleaned version of the toxic data and returning it.
[0125] In the fields of data mining and machine learning, data quality is one of the key factors affecting model performance. The present invention has wide application value in multiple application scenarios such as social media content review, financial risk control, and intelligent customer service. Currently, this solution has been applied to scenarios such as cleaning large model training datasets, intelligent question answering, and intelligent customer service.
[0126] The present invention solves the problems of low accuracy and poor effect in the existing technical methods when dealing with toxic data. On the one hand, it adopts a generative adversarial network GAN model, which mainly includes two modules: a generator and a discriminator. The confrontation and generation are reflected in: the generator needs to try its best to make the discriminator recognize the samples generated by the generator as real, while the discriminator needs to try its best to recognize the samples generated by the generator as fake. The two are a process of mutual game and learning. Among them, the training of the generator is unsupervised (generative) learning. Data cleaning can be carried out without a large amount of labeled data. At the same time, the data generated by GAN can further expand the subsequent training samples, making the present invention have advantages when the acquisition of toxic text data is limited; on the other hand, it adopts an adaptive threshold adjustment strategy, dynamically adjusts the threshold according to the distribution characteristics of the data, avoids the limitations brought by a fixed threshold, and at the same time, cleaning toxic data based on GAN can not only clean toxic data, but also increase the diversity of the data, further improving the model's recognition ability for different variants and implicit expressions of toxic texts, and having better robustness.
[0127] By scraping the highly discussed topics published on social media platforms, performing manual annotation and experiments with real comment data, and randomly extracting 1000 pieces of data as evaluation samples, it is statistically found that the ratio of toxic and non-toxic data is 1:5.
[0128] Table 1 Comparison of Toxic Data Recognition Effects
[0129] Recognition method Precision Recall F1 Traditional regular method 85% 80% 82% GAN 90% 86% 88% GAN enhancement 92% 88% 90% GAN enhancement + Adaptive threshold adjustment 94.55% 92.26% 93.4%
[0130] As can be seen from the above table, the toxic data cleaning method enhanced by GAN based on adaptive threshold adjustment has improved in terms of accuracy, recall rate, and F1 score compared to other methods, indicating that the method has significantly improved the effect in identifying toxic data. At the same time, in terms of toxic data cleaning, when the input is toxic data and the model identifies it as non-toxic data, the non-toxic data generated by the generator at this time is retained and manually evaluated, and the cleaning accuracy reaches 80%, that is, 80 out of 100 toxic data are cleaned and repaired into non-toxic data, which not only improves the quality of the original data, but also expands the amount of the original data and enriches the data diversity.
[0131] In summary, the toxic data cleaning method proposed by the present invention using GAN enhanced based on adaptive threshold adjustment is an efficient and high-quality cleaning method.
[0132] Based on the same concept, the embodiments of the present disclosure also provide a data cleaning device.
[0133] It can be understood that in order to implement the above functions, the data cleaning device provided by the embodiments of the present disclosure includes the corresponding hardware structure and / or software module for executing each function. Combining the units and algorithm steps of the examples disclosed in the embodiments of the present disclosure, the embodiments of the present disclosure can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in the way of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the technical solution of the embodiments of the present disclosure.
[0134] Figure 8 It is a block diagram of a data cleaning device 100 shown according to an exemplary embodiment. Referring to Figure 8 ,the device 100 includes an acquisition unit 101 and an execution unit 102.
[0135] The acquisition unit 101 is used to acquire the data set to be cleaned.
[0136] The execution unit 102 is used to call a preset generative adversarial network model to process the dataset to be cleaned, and obtain the cleaned data. The generator in the generative adversarial network model is used to generate the cleaned data, and the discriminator in the generative adversarial network model is used to determine the probability that the cleaned data generated by the generator is the target type data, where the target type data is dirty data or clean data; in the training stage of the generative adversarial network model, the generator is used to generate the second type of data based on the first type of data, the first type of data is the sample data to be cleaned, and the second type of data is the sample pseudo-data that makes the discriminator identify as the target type data; in the training stage of the generative adversarial network model, the discriminator is used to distinguish the second type of data from the first type of data to determine the second type of data as non-target type data.
[0137] In one implementation, the discriminator determines the second type of data as non-target type data based on a dynamic detection threshold, so as to determine the second type of data as non-target type data.
[0138] In one implementation, the discriminator in the execution unit 102 adopts the following method to distinguish the second type of data from the first type of data based on a dynamic detection threshold, so as to determine the second type of data as non-target type data: input the second type of data and the first type of data into the discriminator of the generative adversarial network prediction model; determine the recognition accuracy of the discriminator, where the recognition accuracy is the accuracy of the discriminator to distinguish the second type of data from the first type of data based on the first detection threshold; dynamically adjust the first detection threshold based on the recognition accuracy to obtain the second detection threshold; distinguish the second type of data from the first type of data based on the second detection threshold; iteratively execute the above process until the discriminator determines the second type of data as non-target type data.
[0139] In one implementation, the execution unit 102 adopts the following method to determine the recognition accuracy of the discriminator: determine the first quantity of the first type of data determined as the target type data by the discriminator based on the first detection threshold, and determine the second quantity of the second type of data determined as the target type data by the discriminator based on the first detection threshold; determine the sum value of the first quantity and the second quantity, and determine the ratio of the sum value to the third quantity as the accuracy of the discriminator to distinguish the second type of data from the first type of data based on the first detection threshold; the third quantity is the total quantity of the first type of data and the second type of data.
[0140] In one implementation, the execution unit 102 adopts the following method to dynamically adjust the first detection threshold based on the accuracy to obtain the second detection threshold: if the recognition accuracy is greater than or equal to the accuracy threshold, then use the first detection threshold as the second detection threshold; if the recognition accuracy is less than the accuracy threshold, then determine the second detection threshold based on the first detection threshold and the recognition accuracy.
[0141] In one embodiment, the execution unit 102 determines the second detection threshold based on the first detection threshold and the recognition accuracy rate by using the following formula: T t = T t-1 + α * (Acc - T t-1 ); where t represents the current iteration number, T is a preset initial threshold, α is the learning rate, and Acc is the current model training accuracy rate.
[0142] In one embodiment, the acquisition unit 101 acquires the dataset to be cleaned in the following manner, including: acquiring the original dataset obtained by collection, where the original dataset includes multiple pieces of original data, and the format of each piece of original data is text; preprocessing the original dataset to obtain the dataset to be cleaned, where the dataset to be cleaned includes multiple pieces of data to be cleaned, and the format of each piece of data to be cleaned is a vector.
[0143] In one embodiment, the acquisition unit 101 preprocesses the original dataset in the following manner to obtain the dataset to be cleaned, including: performing a first preprocessing on the original dataset to obtain a first dataset; the first preprocessing includes: removing any one or more of stop words, special characters, and garbled characters; performing word segmentation on the first dataset to obtain a second dataset; vectorizing the text data in the second dataset to obtain the dataset to be cleaned.
[0144] Regarding the device in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated herein.
[0145] Figure 9 is a block diagram of a device 200 for data cleaning shown according to an exemplary embodiment. The device 200 may be provided as a terminal. For example, the device 200 may be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.
[0146] Referring to Figure 9 , the device 200 may include one or more of the following components: a processing component 202, a memory 204, a power component 206, a multimedia component 208, an audio component 210, an input / output (I / O) interface 212, a sensor component 214, and a communication component 216.
[0147] The processing component 202 generally controls the overall operation of the device 200, such as operations associated with display, telephone calls, data communication, camera operations, and recording operations. The processing component 202 may include one or more processors 220 to execute instructions to complete all or part of the steps of the above-described methods. Additionally, the processing component 202 may include one or more modules to facilitate the interaction between the processing component 202 and other components. For example, the processing component 202 may include a multimedia module to facilitate the interaction between the multimedia component 208 and the processing component 202.
[0148] The memory 204 is configured to store various types of data to support the operation of the device 200. Examples of such data include instructions for any application or method operating on the device 200, contact data, phone book data, messages, pictures, videos, and the like. The memory 204 may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disk.
[0149] The power component 206 provides power to the various components of the device 200. The power component 206 may include a power management system, one or more power sources, and other components associated with generating, managing, and distributing power for the device 200.
[0150] The multimedia component 208 includes a screen that provides an output interface between the device 200 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors may sense not only the boundaries of the touch or swipe actions but also detect the duration and pressure associated with the touch or swipe operations. In some embodiments, the multimedia component 208 includes a front camera and / or a rear camera. When the device 200 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera may receive external multimedia data. Each of the front camera and the rear camera may be a fixed optical lens system or have a focal length and optical zoom capabilities.
[0151] The audio component 210 is configured to output and / or input audio signals. For example, the audio component 210 includes a microphone (MIC) that is configured to receive external audio signals when the device 200 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signal can be further stored in the memory 204 or transmitted via the communication component 216. In some embodiments, the audio component 210 further includes a speaker for outputting audio signals.
[0152] The I / O interface 212 provides an interface between the processing component 202 and peripheral interface modules, and the peripheral interface modules may be a keyboard, a click wheel, buttons, etc. These buttons may include but are not limited to: a home button, a volume button, a power button, and a lock button.
[0153] The sensor component 214 includes one or more sensors for providing an assessment of various aspects of the state of the device 200. For example, the sensor component 214 can detect the on / off state of the device 200, the relative positioning of components, such as the display and keypad of the device 200, the sensor component 214 can also detect a change in the position of the device 200 or a component of the device 200, the presence or absence of user contact with the device 200, the orientation or acceleration / deceleration of the device 200, and the temperature change of the device 200. The sensor component 214 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor component 214 may also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor component 214 may further include an acceleration sensor, a gyro sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.
[0154] The communication component 216 is configured to facilitate communication between the device 200 and other devices in a wired or wireless manner. The device 200 can access a wireless network based on a communication standard, such as WiFi, 2G, or 3G, or a combination thereof. In an exemplary embodiment, the communication component 216 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 216 further includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0155] In an exemplary embodiment, the apparatus 200 may be implemented by one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components for performing the above method.
[0156] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 204 including instructions, and the above instructions can be executed by a processor 220 of the apparatus 200 to complete the above method. For example, the non-transitory computer-readable storage medium may be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc.
[0157] Figure 10 FIG. is a block diagram of an apparatus 300 for data cleaning according to an exemplary embodiment. For example, the apparatus 300 may be provided as a server. Referring to Figure 10 , the apparatus 300 includes a processing component 322, which further includes one or more processors, and memory resources represented by a memory 332 for storing instructions executable by the processing component 322, such as application programs. The application programs stored in the memory 332 may include one or more modules each corresponding to a set of instructions. In addition, the processing component 322 is configured to execute instructions to perform the above method.
[0158] The apparatus 300 may also include a power component 326 configured to perform power management of the apparatus 300, a wired or wireless network interface 350 configured to connect the apparatus 300 to a network, and an input / output (I / O) interface 358. The apparatus 300 may operate based on an operating system stored in the memory 332, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, or the like.
[0159] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 332 including instructions, and the above instructions can be executed by a processing component 322 of the apparatus 300 to complete the above method. For example, the non-transitory computer-readable storage medium may be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc.
[0160] It can be understood that in this disclosure, "a plurality of" means two or more, and other quantifiers are similar thereto. "And / or" describes the relationship between associated objects and indicates that there can be three relationships. For example, A and / or B can mean: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally indicates that the associated objects before and after are in an "or" relationship. The singular forms of "a", "the", and "said" are also intended to include the plural forms unless the context clearly indicates otherwise.
[0161] It can be further understood that the terms "first", "second", etc. are used to describe various information, but such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other and do not indicate a specific order or degree of importance. In fact, the expressions such as "first" and "second" can be used interchangeably. For example, without departing from the scope of this disclosure, the first information can also be referred to as the second information, and similarly, the second information can also be referred to as the first information.
[0162] It can be further understood that the orientation or positional relationship indicated by terms such as "center", "longitudinal", "transverse", "front", "rear", "upper", "lower", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing this embodiment and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation.
[0163] It can be further understood that unless otherwise specified, "connection" includes direct connection without other components between the two, and also includes indirect connection with other elements between the two.
[0164] It can be further understood that although the operations are described in a specific order in the drawings in the embodiments of this disclosure, it should not be understood as requiring these operations to be performed in the specific order shown or in a serial order, or requiring all the operations shown to obtain the desired result. In a specific environment, multitasking and parallel processing may be beneficial.
[0165] Those skilled in the art will readily think of other embodiments of this disclosure after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common general knowledge or conventional technical means in this technical field that are not disclosed in this disclosure.
[0166] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.
Claims
1. A data cleaning method, characterized in that, comprising: obtaining a dataset to be cleaned; invoking a preset generative adversarial network model to process the dataset to be cleaned, obtaining cleaned data, wherein, in the generative adversarial network model, the generator is used to generate the cleaned data, and the discriminator in the generative adversarial network model is used to determine the probability that the cleaned data generated by the generator is target type data, and the target type data is dirty data or clean data; In the training stage of the generative adversarial network model, the generator is used to generate second type data based on first type data, the first type data is sample data to be cleaned, and the second type data is sample pseudo data that enables the discriminator to identify as the target type data; In the training stage of the generative adversarial network model, the discriminator is used to distinguish the second type data from the first type data to determine the second type data as non-target type data.
2. The method according to claim 1, characterized in that, the discriminator determines the second type data as non-target type data based on a dynamic detection threshold to determine the second type data as non-target type data.
3. The method according to claim 2, characterized in that, the discriminator adopts the following method to distinguish the second type data from the first type data based on a dynamic detection threshold to determine the second type data as non-target type data: inputting the second type data and the first type data into the discriminator; determining the recognition accuracy of the discriminator, where the recognition accuracy is the accuracy of the discriminator in distinguishing the second type data from the first type data based on a first detection threshold; adjusting the first detection threshold based on the recognition accuracy to obtain a second detection threshold; distinguishing the second type data from the first type data based on the second detection threshold; iteratively executing the above process until the discriminator determines the second type data as non-target type data.
4. The method according to claim 3, characterized in that, the determining the recognition accuracy of the discriminator includes: determining a first quantity that the discriminator determines the first type data as target type data based on the first detection threshold, and determining a second quantity that the discriminator determines the second type data as target type data based on the first detection threshold; determining the sum value of the first quantity and the second quantity, and determining the ratio of the sum value to a third quantity as the accuracy of the discriminator in distinguishing the second type data from the first type data based on the first detection threshold; the third quantity is the total quantity of the first type data and the second type data.
5. The method according to claim 3 or 4, characterized in that, the adjusting the first detection threshold based on the recognition accuracy to obtain a second detection threshold includes: if the recognition accuracy is greater than or equal to an accuracy threshold, using the first detection threshold as the second detection threshold; if the recognition accuracy is less than the accuracy threshold, determining the second detection threshold based on the first detection threshold and the recognition accuracy.
6. The method according to claim 5, wherein, the following formula is used to determine a second detection threshold based on the first detection threshold and the recognition accuracy rate: T t = T t-1 + α * (Acc - T t-1 ) where t represents the current iteration number, T is a preset initial threshold, α is a learning rate, and Acc is the current model training accuracy rate.
7. The method according to claim 1, wherein, the obtaining of the dataset to be cleaned includes: obtaining the collected original dataset, where the original dataset includes multiple pieces of original data, and the format of each piece of original data is text; performing preprocessing on the original dataset to obtain the dataset to be cleaned, where the dataset to be cleaned includes multiple pieces of data to be cleaned, and the format of each piece of data to be cleaned is a vector.
8. The method according to claim 7, wherein, the performing of preprocessing on the original dataset to obtain the dataset to be cleaned includes: performing first preprocessing on the original dataset to obtain a first dataset; the first preprocessing includes removing any one or more of stop words, special characters, and garbled characters; performing word segmentation processing on the first dataset to obtain a second dataset; vectorizing the text data in the second dataset to obtain the dataset to be cleaned.
9. A data cleaning device, wherein, it includes: an obtaining unit, configured to obtain a dataset to be cleaned; an execution unit, configured to call a preset generative adversarial network model to process the dataset to be cleaned to obtain cleaned data, where the generator in the generative adversarial network model is used to generate the cleaned data, and the discriminator in the generative adversarial network model is used to determine the probability that the cleaned data generated by the generator is target type data, and the target type data is dirty data or clean data; in the training stage of the generative adversarial network model, the generator is used to generate a second type of data based on a first type of data, the first type of data is sample data to be cleaned, and the second type of data is sample pseudo data that enables the discriminator to identify as the target type data; in the training stage of the generative adversarial network model, the discriminator is used to distinguish the second type of data and the first type of data to determine that the second type of data is not the target type data.
10. The device according to claim 9, wherein, the discriminator determines that the second type of data is not the target type data based on a dynamic detection threshold to determine that the second type of data is not the target type data.
11. The device according to claim 10, wherein, the discriminator in the execution unit uses the following method to distinguish the second type of data and the first type of data based on a dynamic detection threshold to determine that the second type of data is not the target type data: inputting the second type of data and the first type of data into the discriminator of the generative adversarial network prediction model; determining the recognition accuracy rate of the discriminator, where the recognition accuracy rate is the accuracy rate at which the discriminator distinguishes the second type of data and the first type of data based on a first detection threshold; Adjust the first detection threshold based on the recognition accuracy rate to obtain a second detection threshold; Distinguish the second type of data and the first type of data based on the second detection threshold; Iteratively execute the above process until the discriminator determines that the second type of data is not the target type of data.
12. The apparatus according to claim 11, wherein, The execution unit determines the recognition accuracy rate of the discriminator in the following manner: Determine the first quantity of the first type of data determined by the discriminator as the target type of data based on the first detection threshold, and determine the second quantity of the second type of data determined by the discriminator as the target type of data based on the first detection threshold; Determine the sum value of the first quantity and the second quantity, and determine the ratio of the sum value to the third quantity as the accuracy rate of the discriminator for distinguishing the second type of data and the first type of data based on the first detection threshold; The third quantity is the total quantity of the first type of data and the second type of data.
13. The apparatus according to claim 11 or 12, wherein, The execution unit adjusts the first detection threshold based on the accuracy rate in the following manner to obtain a second detection threshold: If the recognition accuracy rate is greater than or equal to the accuracy rate threshold, use the first detection threshold as the second detection threshold; If the recognition accuracy rate is less than the accuracy rate threshold, determine the second detection threshold based on the first detection threshold and the recognition accuracy rate.
14. The apparatus according to claim 13, wherein, The execution unit determines the second detection threshold based on the first detection threshold and the recognition accuracy rate using the following arithmetic formula: T t = T t-1 + α * (Acc - T t-1 ); where t represents the current iteration number, T is a preset initial threshold, α is a learning rate, and Acc is the current model training accuracy rate.
15. The apparatus according to claim 9, wherein, The acquisition unit acquires the dataset to be cleaned in the following manner, including: Acquire the collected original dataset, where the original dataset includes multiple pieces of original data, and the format of each piece of original data is text; Preprocess the original dataset to obtain the dataset to be cleaned, where the dataset to be cleaned includes multiple pieces of data to be cleaned, and the format of each piece of data to be cleaned is a vector.
16. The apparatus according to claim 15, wherein, The acquisition unit preprocesses the original dataset to obtain the dataset to be cleaned in the following manner, including: Perform a first preprocessing on the original dataset to obtain a first dataset; the first preprocessing includes: removing any one or more of stop words, special characters, and garbled characters; Perform word segmentation on the first dataset to obtain a second dataset; Vectorize the text data in the second dataset to obtain the dataset to be cleaned.
17. A data cleaning apparatus, wherein, Comprising: A processor; A memory for storing instructions executable by the processor; wherein, the processor is configured to: execute the data cleaning method according to any one of claims 1-8.
18. A non-transitory computer-readable storage medium, when the instructions in the storage medium are executed by a processor of a mobile terminal, enable the mobile terminal to execute the data cleaning method according to any one of claims 1-8.
Citation Information
Cited By
Intelligent data cleaning method and device, computer equipment and storage medium
CN120994653A
Data intelligent cleaning methods, devices, computer equipment and storage media
CN120994653B