Data evaluation method, device, computer equipment and storage medium
By filtering and multi-dimensional analysis of the data set, the problems of high cost, low efficiency and poor accuracy of data evaluation in the prior art are solved, and efficient and reliable data quality evaluation is achieved.
Patent Information
- Application Number
- CN202110266548.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-03-11
- Publication Date
- 2025-08-26
- Estimated Expiration
- 2041-03-11
AI Technical Summary
In the prior art, data evaluation methods rely on manual review, resulting in high cost, low efficiency and poor accuracy, especially in text semantic model training, noisy data has a significant impact.
By obtaining the data set to be evaluated, applying a preset filtering strategy to obtain the candidate data set, and analyzing it based on evaluation dimensions such as availability, consistency, anomaly detection and complexity, and obtaining evaluation parameters to determine data quality.
It improves the accuracy and reliability of data evaluation, reduces the need for manual review, and improves the efficiency and accuracy of data evaluation.
Smart Images

Figure CN113704389B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of Internet technology, and in particular to a data evaluation method, apparatus, computer equipment, and storage medium. Background Art
[0002] With the rapid development of artificial intelligence, various models are being used more and more widely. For example, in the customer service field, call centers generate massive amounts of voice calls every day. In order to monitor service quality and public opinion risks, the text semantic model constructed in the intelligent quality inspection system can be used to analyze whether the user's words have risk tendencies, whether there is dissatisfaction, and whether the customer service's statements are unreasonable.
[0003] The effectiveness of text semantic models relies on training with large amounts of high-quality labeled data. During the data annotation process, different annotators often have different understandings, so the annotated datasets often contain a lot of abnormal noise data. This noise data directly affects the effectiveness of text semantic model training. Furthermore, different data contributes differently to improving text semantic model performance: simple data is easy to distinguish and has limited impact on text semantic model performance; excessive amounts of simple data can dominate changes in text semantic model parameters, affecting the efficiency of text semantic model training. Therefore, evaluating data quality is crucial.
[0004] In the existing technology, manual evaluation of the data set is required: the sampled data is given to experts for manual review, and the accuracy of the data set is estimated based on the accuracy of the data sampling; or, knowledge points (golden set) are set in the data set, and the golden set is put into each annotator's data set to be annotated, and the accuracy of the annotator's annotated data is estimated based on the accuracy of the golden set.
[0005] Sampling verification methods rely on manually specifying the sampling ratio. A high sampling ratio results in high labor and time costs, while a low sampling ratio can lead to high limitations and bias in the results, resulting in low data assessment accuracy. The golden set must be pre-configured and added to the dataset by experts. This set requires frequent updates to ensure that each annotator's dataset contains the unannotated golden set. This is cumbersome and reduces the accuracy of data assessment. Summary of the Invention
[0006] The embodiments of the present application provide a data evaluation method, apparatus, computer device, and storage medium, which can improve the accuracy and reliability of data evaluation.
[0007] To solve the above technical problems, the embodiments of the present application provide the following technical solutions:
[0008] The present invention provides a data evaluation method, including:
[0009] Get the dataset to be evaluated;
[0010] Filter the dataset to be evaluated according to a preset filtering strategy to obtain a candidate dataset;
[0011] Analyzing the candidate data set based on the evaluation dimension;
[0012] Obtaining evaluation parameters of the data set to be evaluated according to the analysis results;
[0013] An evaluation result corresponding to the data set to be evaluated is determined according to the evaluation parameters, so as to process the data according to the evaluation result.
[0014] According to one aspect of the present application, a data evaluation device is further provided, comprising:
[0015] A first acquisition unit is used to acquire a data set to be evaluated;
[0016] A filtering unit, configured to filter the dataset to be evaluated according to a preset filtering strategy to obtain a candidate dataset;
[0017] an analyzing unit, configured to analyze the candidate data set based on an evaluation dimension;
[0018] A second acquiring unit, configured to acquire evaluation parameters of the data set to be evaluated according to the analysis result;
[0019] A determining unit is configured to determine an evaluation result corresponding to the data set to be evaluated according to the evaluation parameters, so as to process the data according to the evaluation result.
[0020] According to one aspect of the present application, a computer device is also provided, including a processor and a memory, wherein a computer program is stored in the memory, and when the processor calls the computer program in the memory, any data evaluation method provided in the embodiments of the present application is executed.
[0021] According to one aspect of the present application, a storage medium is further provided, wherein the storage medium is used to store a computer program, and the computer program is loaded by a processor to execute any one of the data evaluation methods provided in the embodiments of the present application.
[0022] The embodiment of the present application can obtain a data set to be evaluated, and filter the data set to be evaluated according to a preset filtering strategy to obtain a candidate data set. The candidate data set can then be analyzed based on the evaluation dimension, and evaluation parameters of the data set to be evaluated can be obtained based on the analysis results. At this time, the evaluation result corresponding to the data set to be evaluated can be determined based on the evaluation parameters, and the data can be processed based on the evaluation results. This solution can filter the data set to be evaluated, analyze the filtered candidate data set based on the evaluation dimension to obtain evaluation parameters, and determine the evaluation result corresponding to the data set to be evaluated based on the evaluation parameters, thereby avoiding manual review and evaluation, and improving the accuracy and reliability of data evaluation. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.
[0024] Figure 1 Schematic diagram of a scenario in which the data evaluation method provided in an embodiment of the present application is applied;
[0025] Figure 2 Schematic diagram of the data evaluation method provided in the embodiment of the present application;
[0026] Figure 3 is a schematic diagram of performing complexity evaluation on data provided in an embodiment of the present application;
[0027] Figure 4 This is another flow chart of the data evaluation method provided in an embodiment of the present application;
[0028] Figure 5 This is a flowchart of training a text-to-speech model according to an embodiment of the present application;
[0029] Figure 6 This is a flow chart of detecting a text to be detected using a trained text-to-speech model provided in an embodiment of the present application;
[0030] Figure 7 is a schematic diagram of a data evaluation device provided in an embodiment of the present application;
[0031] Figure 8 It is a structural diagram of the computer device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0032] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without making creative efforts are within the scope of protection of this application.
[0033] Embodiments of the present application provide a data evaluation method, apparatus, computer device, and storage medium.
[0034] See also Figure 1 , Figure 1 Schematic diagram of the scenario of the application of the data evaluation method provided in the embodiment of the present application, the data evaluation method application may include a data evaluation device, which may be specifically integrated in a computer device such as a server or a terminal, the server may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content distribution networks (Content Delivery Network, CDN), and big data and artificial intelligence platforms, but is not limited to this. The server and the terminal may be directly or indirectly connected via wired or wireless communication, which is not limited in this application. The terminal may be a mobile phone, tablet computer, laptop computer, desktop computer, or wearable device, etc.
[0035] Among them, the computer device can be used to obtain the data set to be evaluated, and filter the data set to be evaluated according to the preset filtering strategy to obtain a candidate data set. Then, the candidate data set can be analyzed based on the evaluation dimension, and the evaluation parameters of the data set to be evaluated can be obtained based on the analysis results. For example, the candidate data set can be evaluated for availability, consistency, anomaly detection, and complexity, and evaluation parameters such as availability rate, consistency rate, anomaly rate, and complexity rate can be obtained. At this time, the evaluation result corresponding to the data set to be evaluated (such as whether the data quality is qualified or unqualified) can be determined based on the evaluation parameters, so that the data can be processed according to the evaluation results, thereby improving the accuracy and reliability of the data evaluation.
[0036] It should be noted that Figure 1The scenario diagram of the data evaluation method application shown is only an example. The data evaluation method application and scenario described in the embodiment of the present application are intended to more clearly illustrate the technical solution of the embodiment of the present application, and do not constitute a limitation on the technical solution provided by the embodiment of the present application. Ordinary technicians in this field can know that with the evolution of the application of the data evaluation method and the emergence of new business scenarios, the technical solution provided by the embodiment of the present application is also applicable to similar technical problems.
[0037] It should be noted that the order of description of the following embodiments is not intended to limit the preferred order of the embodiments.
[0038] In this embodiment, the description will be made from the perspective of a data evaluation device, which can be integrated into a computer device such as a server or a terminal.
[0039] See also Figure 2 , Figure 2 : is a flow chart of a data evaluation method provided in one embodiment of the present application. The data evaluation method may include:
[0040] S101: Obtain a dataset to be evaluated.
[0041] The dataset to be evaluated may contain multiple sample data. The specific type of the sample data can be flexibly set according to actual needs. For example, the sample data can be text, speech, images, or other types of data. If the sample data is speech, semantic recognition can be performed on the speech to convert it into text. The sample data in the dataset to be evaluated can be labeled. The label can be used to characterize the category of the sample data or to identify characteristics of the sample data, such as whether the sample data has a risk tendency.
[0042] Methods for obtaining the dataset to be evaluated may include: obtaining multiple pre-stored sample data from a local database to obtain the dataset to be evaluated; or, downloading the dataset to be evaluated from a server; or, receiving multiple sample data sent by a terminal and generating the dataset to be evaluated based on the multiple sample data; and so on.
[0043] S102: Filter the dataset to be evaluated according to a preset filtering strategy to obtain a candidate dataset.
[0044] In order to improve the efficiency and reliability of processing the dataset to be evaluated, the dataset to be evaluated may be filtered to remove unnecessary interference data.
[0045] Among them, the preset filtering strategy can be flexibly set according to actual needs.
[0046] In one embodiment, filtering the dataset to be evaluated according to a preset filtering strategy to obtain the candidate dataset may include filtering, according to the preset filtering strategy, empty samples, samples consisting entirely of digits, or text with a word count less than a preset word count threshold in the dataset to obtain the candidate dataset. Empty samples may be blank text or garbled text.
[0047] S103. Analyze the candidate data set based on the evaluation dimension.
[0048] S104: Obtain evaluation parameters of the data set to be evaluated according to the analysis results.
[0049] The evaluation dimensions may include at least one of availability evaluation, consistency evaluation, anomaly detection evaluation, and complexity evaluation, and the evaluation parameters may include at least one of availability rate, consistency rate, anomaly rate, and complexity rate. Of course, the evaluation dimensions and evaluation parameters can be flexibly set according to actual needs, and the specific content is not limited here.
[0050] In one embodiment, the evaluation dimension includes availability evaluation, the evaluation parameter includes availability rate, and the candidate data set is analyzed based on the evaluation dimension. Obtaining the evaluation parameters of the data set to be evaluated according to the analysis results may include: obtaining a first data volume of the data set to be evaluated and a second data volume of the candidate data set; and calculating the availability rate of the data set to be evaluated based on the first data volume and the second data volume.
[0051] To accurately determine the availability of a dataset to be evaluated, an availability assessment can be performed on the dataset. The availability assessment can be performed by measuring the availability of the candidate datasets obtained after filtering the dataset to be evaluated. The availability rate can be measured by the proportion of candidate datasets in the dataset to be evaluated. A higher availability rate indicates better data quality, while a lower availability rate indicates poorer data quality.
[0052] Specifically, during the availability evaluation process, the first data volume of the dataset to be evaluated (i.e., the data size of the dataset to be evaluated) and the second data volume of the candidate dataset (i.e., the data size of the candidate dataset) can be obtained. Then, the ratio between the first data volume and the second data volume can be calculated, and the availability of the dataset to be evaluated can be determined based on the ratio between the first data volume and the second data volume. For example, the first data volume of the dataset to be evaluated is C, and the candidate dataset D is a The second data volume is |D a |, then the availability rate p of the dataset to be evaluated a For: p a =|D a | / C.
[0053] In one embodiment, the evaluation dimension includes consistency evaluation, the evaluation parameter includes consistency rate, and the candidate data set is analyzed based on the evaluation dimension. Obtaining the evaluation parameters of the data set to be evaluated based on the analysis results may include: performing duplicate detection on the candidate data set to obtain a duplicate data set; extracting data with consistent labels from the duplicate data set; obtaining the ratio of the data volume of the data with consistent labels to the data volume in the data set to be evaluated; and determining the consistency rate of the data set to be evaluated based on the ratio.
[0054] To accurately determine the data quality of the dataset to be evaluated, a consistency assessment can be performed on the dataset. This consistency assessment can be performed by determining whether the labels of the sample data in the candidate dataset obtained after filtering the dataset to be evaluated are consistent. The consistency rate can be measured by the proportion of sample data in the candidate dataset with consistent labels in the dataset to be evaluated. A higher consistency rate indicates better data quality, while a lower consistency rate indicates poorer data quality.
[0055] Specifically, in the process of consistency evaluation, the candidate data set can be tested for duplicates to obtain a duplicate data set. The duplicate data set may include multiple duplicate data samples. Since the sample data of the data set to be evaluated can be set with labels, the sample data in the duplicate data set is also set with labels. Then, the data with consistent labels (i.e., sample data) can be extracted from the duplicate data set, and the ratio of the amount of data with consistent labels to the amount of data in the data set to be evaluated can be obtained, and the consistency rate of the data set to be evaluated can be determined based on the ratio. For example, the amount of data with consistent labels in the duplicate data set is C1, the amount of data in the data set to be evaluated is C, and the consistency rate of the data set to be evaluated is p d = C1 / C. For another example, the amount of data with inconsistent labels that can be extracted from the repeated data set is C d , the amount of data in the dataset to be evaluated is C, and the consistency rate of the dataset to be evaluated is p d =1-C d / C.
[0056] In one embodiment, the evaluation dimension includes anomaly detection evaluation, the evaluation parameter includes anomaly rate, and the candidate data set is analyzed based on the evaluation dimension. Obtaining the evaluation parameters of the data set to be evaluated based on the analysis results may include: performing word segmentation processing on each sample data in the candidate data set to obtain a word set corresponding to each sample data; vectorizing the word set to obtain a word vector; projecting the word vector to a sample feature space of a preset dimension to obtain a text feature vector; clustering the text feature vector, and determining the abnormal data in the candidate data set based on the clustering results; and determining the abnormal rate of the data set to be evaluated based on the abnormal data.
[0057] To improve the accuracy of the dataset evaluation, anomaly detection can be performed on the dataset. This anomaly detection evaluation characterizes the abnormality of the dataset by determining whether the sample data in the candidate dataset, obtained after filtering the dataset, is abnormal. The anomaly rate can be characterized by the proportion of abnormal data in the candidate dataset to the dataset. A higher anomaly rate indicates lower data quality, while a lower anomaly rate indicates better data quality.
[0058] Specifically, during the anomaly detection evaluation process, when the sample data is text, each sample data in the candidate data set can be segmented separately to obtain a word set corresponding to each sample data; when the sample data is speech, semantic recognition can be performed on the sample data to convert the speech into text. In this case, each sample data in the candidate data set can be segmented separately to obtain a word set corresponding to each sample data. The word segmentation processing method can be flexibly set according to actual needs. For example, the Jieba word segmentation strategy can be used to segment each sample data in the candidate data set, or semantic recognition can be performed on the sample data and each sample data in the candidate data set can be segmented separately according to single words, words, or phrases.
[0059] Then, the word set corresponding to each sample data can be vectorized to obtain the word vector corresponding to each sample data. For example, the word set of each sample data can be vectorized using the related model word to vector (abbreviated as word2vec) for generating word vectors, and the word vector of each sample data can be projected into a sample feature space of a preset dimension (for example, N dimension, N can be 128 or flexibly set according to actual needs) to obtain a text feature vector, which can be specifically as follows:
[0060]
[0061] Among them, sen_vec can represent the text feature vector (also called sentence vector), m can represent the number of words contained in each sample data, vec i It can represent the i-th word vector.
[0062] At this point, the text feature vectors can be clustered, and abnormal data in the candidate data set can be determined based on the clustering results. In one embodiment, clustering the text feature vectors and determining abnormal data in the candidate data set based on the clustering results can include: performing dimensionality reduction processing on the text feature vectors to obtain a reduced-dimensionality feature vector; normalizing the reduced-dimensionality feature vector to obtain a normalized feature vector; clustering the normalized feature vector, and determining abnormal data in the candidate data set based on the clustering results.
[0063] In order to improve the accuracy of clustering, the text feature vector can be processed by dimensionality reduction and normalization. For example, the N-dimensional text feature vector can be reduced in dimensionality based on latent semantic analysis (LSA) to obtain a reduced feature vector. The dimension of the reduced feature vector can be D-dimensional, and D can be 15 or flexibly set according to actual needs. Among them, LSA can map the text in the high-dimensional vector space model (VSM) representation to the low-dimensional latent semantic space. The mapping can be achieved by singular value decomposition (SVD) of the text feature vector matrix.
[0064] Then, the eigenvector after dimensionality reduction can be normalized to obtain a normalized eigenvector. For example, since the eigenvector can be represented by a numerical value, in order to improve the reliability of subsequent clustering, the eigenvector after dimensionality reduction can be normalized to a numerical value in the range of 0 to 1 to obtain a normalized eigenvector. At this time, the normalized eigenvector can be clustered, and the abnormal data in the candidate data set can be determined based on the clustering results. For example, the normalized eigenvector can be clustered by a density clustering algorithm (Density-Based Spatial Clustering of Applications with Noise, DBSCAN). Among them, the abnormal data can be data with incorrect labels due to negligence. The abnormal data can be an isolated point (i.e., isolated data) in the clustered data set. Samples with similar distances in the text feature vector space can be clustered into a cluster, and points that do not belong to any cluster are isolated points.
[0065] In one embodiment, clustering the normalized feature vectors and determining abnormal data in the candidate data set based on the clustering results may include: discretizing the normalized feature vectors into multiple feature points; selecting any one feature point from the multiple feature points as a core point; assigning all feature points within a preset neighborhood range centered on the core point to the same class family; selecting another feature point from the multiple feature points as the core point, and returning to execute the operation of assigning all feature points within a preset neighborhood range centered on the core point to the same class family until the multiple feature points are traversed; and filtering out sample data corresponding to feature points that are not assigned to a class family to obtain abnormal data.
[0066] Specifically, the normalized feature vector can be discretized into multiple feature points, and any one feature point can be selected from the multiple feature points as the core point, and all feature points within a preset neighborhood with the core point as the center can be assigned to the same class family. For example, all feature points within a preset neighborhood with a radius of R as the center can be assigned to the same class family, and the value of R can be flexibly set according to actual needs. Another feature point can be selected from the multiple feature points as the core point, and the operation of assigning all feature points within the preset neighborhood with the core point as the center to the same class family is returned until multiple feature points are traversed. For example, all feature points in the neighborhood of the core point can be traversed in turn as core points {y1, y2, ..., y i}, for example, i As a new core point, i All feature points in the neighborhood are assigned to the same cluster, and traverse y in turn i All feature points in the neighborhood are respectively regarded as core points {z1,z2,...,z i}, and so on, the cluster gradually increases until there are no core points that can be expanded. Then find a core point that has not been assigned a cluster and assign it to a new cluster. Repeat the above steps to expand the core point until all core points in the data set are assigned cluster labels. The data points in the data set that are not assigned any cluster labels are outliers, and finally we can get the collection of outliers D outlier , that is, filtering out the sample data corresponding to feature points that are not assigned to a cluster, obtaining abnormal data. A core point is a point whose density exceeds a certain threshold, Threshold, and can serve as the center of a cluster. The density ρ is the number of data points whose distance from this point is less than r, where the Euclidean distance can be used. An abnormal point (also called a noise point) is a point that does not belong to any cluster.
[0067] After obtaining the abnormal data, the abnormality rate of the dataset to be evaluated can be determined based on the abnormal data. For example, the ratio between the amount of abnormal data and the amount of data in the dataset to be evaluated can be calculated, and the ratio between the amount of abnormal data and the amount of data in the dataset to be evaluated can be used as the abnormality rate: abnormality rate = amount of abnormal data / amount of data in the dataset to be evaluated.
[0068] In one embodiment, the evaluation dimension includes anomaly detection evaluation, the evaluation parameter includes anomaly rate, and the candidate data set is analyzed based on the evaluation dimension. Obtaining the evaluation parameters of the data set to be evaluated based on the analysis results may include: performing classification prediction on each sample data in the candidate data set through the trained classification model to obtain the classification probability corresponding to each sample data in the candidate data set; screening out sample data with a classification probability less than a preset threshold as abnormal data; and determining the abnormal rate of the data set to be evaluated based on the screened abnormal data with a classification probability less than the preset threshold.
[0069] To improve the accuracy of anomaly rate calculation, the anomaly rate can be calculated using anomaly data filtered out by a trained classification model. The specific type and structure of the trained classification model can be flexibly set according to actual needs. For example, the classification model can be a FastText model, a textCNN model, an RCNN model, or a BERT model. After obtaining a candidate dataset, the trained classification model can be used to perform classification predictions on each sample data in the candidate dataset, obtaining the corresponding class label for each sample data in the candidate dataset and the classification probability of each sample data. The classification probability can be the probability of the class label to which the sample data belongs. Sample data with a classification probability less than a preset threshold can then be filtered out and classified as anomaly data. The specific value of the preset threshold can be flexibly set according to actual needs. The anomaly rate of the dataset to be evaluated can then be determined based on the filtered anomaly data with a classification probability less than the preset threshold. For example, the anomaly rate = the amount of anomaly data with a classification probability less than the preset threshold / the amount of data in the dataset to be evaluated.
[0070] In one embodiment, the data evaluation method may further include: obtaining multiple training samples, the multiple training samples including labeled samples and complementary labeled samples, the labeled samples being samples labeled with true labels, and the complementary labeled samples being samples labeled with labels complementary to the true labels; performing negative learning on the initial classification model based on the labeled samples and the complementary labeled samples to train the initial classification model to obtain a candidate classification model, and predicting the sample classification probability of each training sample; screening out training samples whose sample classification probability is greater than the target probability threshold to obtain candidate training samples; and fine-tuning the candidate classification model based on the candidate training samples to obtain a trained classification model.
[0071] In order to improve the accuracy and reliability of the trained classification model in screening abnormal data, the classification model can be trained in advance. Since the sample data may be difficult to distinguish, and the understanding bias will lead to the label error of the sample data, in order to accurately detect the abnormal data with incorrect labels, the classification model can be trained by negative learning training, so that the trained classification model can quickly and accurately find the abnormal points with incorrect labels. noise Negative learning training can be performed using complementary labels. In multi-class labeling tasks, the label set is {L1, L2, ..., L N}, the labeler divides the data into T label classes, then the complementary label C=random({L1,L2,…,L N}-T), that is, randomly selecting a label from the label set excluding the annotated label as the complementary label. When training with complementary labels, for incorrectly labeled data, the probability that the complementary label is the true label is low, and for correctly labeled data, the probability that the complementary label is the true label is zero. Therefore, negative learning reduces the risk of erroneous information and avoids the problem of overfitting outliers in positive learning, making it more suitable for noisy and abnormal data.
[0072] Specifically, multiple pre-stored training samples can be obtained from a local database, or multiple training samples can be downloaded from a server, etc. The training samples can be text, voice or images, etc. The multiple training samples can include label samples and complementary label samples. Label samples are samples marked with true labels, and complementary label samples are samples marked with labels complementary to the true labels. For example, the true label marked with training sample A is category A, and the true label of training sample B is category B, but training sample B can be used as complementary label sample B, and the label marked with complementary label sample B that is complementary to the true label can be not category C, not category D, not category E, not category F, and not category G, etc.
[0073] Then, the initial classification model can be used to perform negative learning based on the labeled samples and the complementary labeled samples to train the initial classification model, obtain a candidate classification model, and predict the sample classification probability of each training sample. For example, the initial classification model can be used to predict the predicted category and sample classification probability of the training sample, and the loss value can be calculated based on the sample classification probability through the loss function to adjust the parameters of the classification model to appropriate values according to the loss value to obtain a candidate classification model. The loss function in negative learning can be:
[0074]
[0075] Among them, L can represent the loss value, N can represent the number of training samples, and C k It can represent the kth training sample, P kIt can represent the sample classification probability of the kth training sample.
[0076] Secondly, training samples whose classification probabilities are greater than a target probability threshold can be screened out to obtain candidate training samples. The target probability threshold can be flexibly set according to actual needs. At this point, the candidate classification model can be further trained based on the candidate training samples to fine-tune the parameters of the candidate classification model and obtain a trained classification model.
[0077] In one embodiment, determining the abnormality rate of the data set to be evaluated based on the screened abnormal data with a classification probability less than a preset threshold may include: obtaining a target feature vector corresponding to each sample data in the candidate data set; determining the target abnormal data in the candidate data set based on the target feature vector; and determining the abnormality rate of the data set to be evaluated based on the target abnormal data and the screened abnormal data with a classification probability less than a preset threshold.
[0078] In order to improve the accuracy of the abnormality rate determination, the abnormal data screened out by the clustering method and the abnormal data screened out by the trained classification model can be combined to determine the abnormality rate of the data set to be evaluated. Specifically, each sample data in the candidate data set can be segmented in the above manner to obtain the word set corresponding to each sample data; the word set can be vectorized to obtain the word vector; the word vector can be projected to the sample feature space of the preset dimension to obtain the target feature vector (i.e., the text feature vector); the target feature vector can be reduced in dimension to obtain the reduced-dimensional feature vector; the reduced-dimensional feature vector can be normalized to obtain the normalized feature vector; the normalized feature vector can be clustered, and the target abnormal data in the candidate data set can be determined based on the clustering results. And, in the above manner, each sample data in the candidate data set can be classified and predicted by the trained classification model to obtain the classification probability corresponding to each sample data in the candidate data set; the sample data with a classification probability less than the preset threshold can be screened out as abnormal data, and then, the target abnormal data D isolate , and the abnormal data D whose classification probability is less than the preset threshold noise , determine the anomaly rate of the dataset to be evaluated. For example, the total amount of abnormal data D outlier =D isolate UD noise , thus obtaining the abnormal rate p o :
[0079]
[0080] Here, C can represent the amount of data in the dataset to be evaluated.
[0081] In one embodiment, the evaluation dimension includes complexity evaluation, the evaluation parameter includes complexity rate, and the candidate data set is analyzed based on the evaluation dimension. Obtaining the evaluation parameters of the data set to be evaluated based on the analysis results may include: performing word segmentation processing on each sample data in the data set to obtain a word set corresponding to each sample data; vectorizing the word set to obtain a vector matrix; performing a convolution operation on the vector matrix to obtain multiple one-dimensional vectors; performing a pooling operation on the multiple one-dimensional vectors to obtain a pooled vector; splicing the pooled vector to obtain a spliced vector; performing a full connection operation on the spliced vector to obtain the label category probability of each sample data; and determining the complexity rate of the data set to be evaluated based on the label category probability.
[0082] To improve the accuracy of the dataset evaluation, a complexity assessment can be performed on the dataset. The complexity assessment can be measured by the complexity of the sample data in the candidate dataset obtained after filtering the dataset. The complexity rate can be represented by the proportion of complex sample data in the candidate dataset to the dataset under evaluation. A higher complexity rate indicates better data quality, while a lower complexity rate indicates poorer data quality.
[0083] Specifically, the complexity of each sample data can be evaluated based on the TEXTCNN model, such as Figure 3 As shown in the figure, the Jieba word segmentation strategy can be used to segment each sample data in the dataset to obtain the word set corresponding to each sample data. Then, Word2vec is used to vectorize the word set to obtain a vector matrix. The vector matrix can be a two-dimensional matrix. The size of the word vector can be D, and the maximum length of the sample data can be M. Here, M can be 20 and D can be 128, or it can be flexibly set according to actual needs.
[0084] After obtaining the vector matrix, three convolution kernels of sizes 2, 3, and 4 can be used to perform convolution operations on the vector matrix to obtain multiple one-dimensional vectors; the multiple one-dimensional vectors are pooled through the pooling layer to obtain the pooled vector corresponding to each sample data. For example, the maximum value of the one-dimensional vector obtained by convolution through the pooling layer is taken to obtain the pooled vector. The pooled vectors corresponding to each sample data are spliced into a vector to obtain the spliced vector, and the spliced vector is fully connected through the fully connected layer to obtain the label category probability of each sample data. The label category with the largest probability is the probability value p of the first category. max , the second largest label category probability is the second category probability value p second , if p max Less than P th1 , and p secondGreater than P th2 , it means that the sample data is near the dividing line, the prediction is uncertain, and it is a relatively complex and difficult to predict sample data. th1 and P th2 It can be flexibly set according to actual needs. For example, here P th1 It can be taken as 0.65, P th2 It can be set to 0.25. max Less than P th1 , and p second Greater than P th2 Sample data of complex sample D complex , so that the complexity rate p of the data set to be evaluated can be calculated c :
[0085]
[0086] Here, C can represent the amount of data in the dataset to be evaluated.
[0087] S105: Determine an evaluation result corresponding to the data set to be evaluated according to the evaluation parameters, and process the data according to the evaluation result.
[0088] In one embodiment, the evaluation parameters may include at least any one of availability rate, consistency rate, abnormality rate, and complexity rate. Determining the evaluation result corresponding to the data set to be evaluated based on the evaluation parameters may include: obtaining the cumulative value of the availability rate, consistency rate, abnormality rate, and complexity rate, and determining the evaluation result corresponding to the data set to be evaluated based on the cumulative value; or, setting weight values for the availability rate, consistency rate, abnormality rate, and complexity rate respectively, performing weighted operations based on the availability rate, consistency rate, abnormality rate, complexity rate and their corresponding weight values to obtain a target value, and the target value determines the evaluation result corresponding to the data set to be evaluated.
[0089] Among them, the evaluation results may include qualified data quality and unqualified data quality, etc. For example, the cumulative value = availability rate + consistency rate + anomaly rate + complexity rate. The higher the cumulative value, the better the data quality. Conversely, the lower the cumulative value, the worse the data quality. It can be set that when the cumulative value is greater than or equal to the target threshold, the evaluation result is qualified data quality, and when the cumulative value is less than the target threshold, the evaluation result is unqualified data quality. The target threshold can be flexibly set according to actual needs, and the specific value is not limited here.
[0090] For another example, the target value = availability rate * weight value of availability rate + consistency rate * weight value of consistency rate + anomaly rate * weight value of anomaly rate + complexity rate * weight value of complexity rate. The higher the target value, the better the data quality. Conversely, the lower the target value, the worse the data quality. It can be set that when the target value is greater than or equal to the target threshold, the evaluation result is that the data quality is qualified. When the target value is less than the target threshold, the evaluation result is that the data quality is unqualified. Among them, the weight value of the availability rate, the weight value of the consistency rate, the weight value of the anomaly rate, the weight value of the complexity rate, and the target threshold can be flexibly set according to actual needs, and the specific values are not limited here.
[0091] It should be noted that, depending on actual needs, the dataset to be evaluated can be evaluated based on only one of the following dimensions: availability, consistency, anomaly, or complexity. Alternatively, any two or three of these dimensions can be considered. Alternatively, all four dimensions can be considered together to evaluate the annotation quality and determine whether the dataset to be evaluated is qualified. Evaluating the quality of annotated data from multiple dimensions improves evaluation efficiency and accuracy.
[0092] In one embodiment, after determining the evaluation result corresponding to the data set to be evaluated based on the evaluation parameters, the data evaluation method may further include: when the evaluation result corresponding to the data set to be evaluated is that the data quality is qualified, using the data set to be evaluated to train the training model to obtain a trained model, and processing the data through the trained model; when the evaluation result corresponding to the data set to be evaluated is that the data quality is unqualified, analyzing the factors affecting the quality of the data set to be evaluated.
[0093] The model structure and specific type of the model to be trained can be flexibly set according to actual needs. For example, the model to be trained can be a text semantic model. When the evaluation result corresponding to the dataset to be evaluated is that the data quality is qualified, the dataset to be evaluated can be used to train the model to obtain a trained model. The trained model can then be used to process the data. For example, in the customer service field, call centers generate a large amount of voice calls every day. To monitor service quality and public opinion risks, one or more trained models can be set up in the intelligent quality inspection system. The trained models can be used to analyze whether the user's words have risk tendencies, whether there is dissatisfaction, and whether the customer service staff's statements are unreasonable. For another example, the trained model can be used to analyze the category to which the data belongs.
[0094] When the evaluation result corresponding to the data set to be evaluated is that the data quality is unqualified, the factors affecting the quality of the data set to be evaluated can be analyzed. For example, based on the availability rate, it can be analyzed whether the factor affecting the quality of the data set to be evaluated is low availability; based on the consistency rate, it can be analyzed whether the factor affecting the quality of the data set to be evaluated is low consistency; based on the anomaly rate, it can be analyzed whether the factor affecting the quality of the data set to be evaluated is caused by abnormal data; based on the complexity rate, it can be analyzed whether the factor affecting the quality of the data set to be evaluated is low complexity, and so on.
[0095] It should be noted that the data evaluation method in the embodiments of this application can also be used to manage the quality of labeled data. By fully evaluating the labeled data sets of annotators, the work of annotators can be supervised and evaluated. Comparative experiments have shown that the method significantly reduces the number of qualified labeled data within the same labeling time, significantly improving the accuracy of model training compared to models trained on raw data.
[0096] The embodiment of the present application can obtain a data set to be evaluated, and filter the data set to be evaluated according to a preset filtering strategy to obtain a candidate data set. The candidate data set can then be analyzed based on the evaluation dimension, and evaluation parameters of the data set to be evaluated can be obtained based on the analysis results. At this time, the evaluation result corresponding to the data set to be evaluated can be determined based on the evaluation parameters, and the data can be processed based on the evaluation results. This solution can filter the data set to be evaluated, analyze the filtered candidate data set based on the evaluation dimension to obtain evaluation parameters, and determine the evaluation result corresponding to the data set to be evaluated based on the evaluation parameters, thereby avoiding manual review and evaluation, and improving the efficiency, accuracy, and reliability of data evaluation.
[0097] The method described in the above embodiment is further described in detail below with examples.
[0098] This embodiment takes the data evaluation device integrated in the server as an example. The process of the data evaluation method provided in the embodiment of the present application may include text screening, model training, and model application, etc., which can be specifically as follows:
[0099] 1. Text Screening
[0100] like Figure 4 As shown, the text screening process provided by the embodiment of the present application may include:
[0101] S201: Obtain a dataset to be evaluated containing multiple texts.
[0102] Taking the sample data as text as an example, for example, the server can obtain multiple pre-stored texts from the database to obtain the data set to be evaluated, or the server can receive multiple texts sent by the terminal to obtain the data set to be evaluated, and so on.
[0103] S202: Filter the dataset to be evaluated according to a preset filtering strategy to obtain a candidate dataset.
[0104] For example, the server can filter out empty text, garbled text, text consisting entirely of numbers, or text with a word count less than a preset word count threshold in the dataset to be evaluated according to a preset filtering strategy to obtain a candidate dataset, thereby removing unnecessary interfering data and improving the efficiency and reliability of processing the dataset to be evaluated.
[0105] S203: Perform availability evaluation on the candidate dataset to obtain the availability rate of the dataset to be evaluated.
[0106] For example, the server may obtain a first data volume of the dataset to be evaluated (i.e., the data size of the dataset to be evaluated) and a second data volume of the candidate dataset (i.e., the data size of the candidate dataset). Then, the server may calculate a ratio between the first data volume and the second data volume, and determine the availability of the dataset to be evaluated based on the ratio between the first data volume and the second data volume: availability ratio p a =Second data amount |D a | / First data volume C.
[0107] S204: Perform consistency evaluation on the candidate data set to obtain the consistency rate of the data set to be evaluated.
[0108] For example, the server can perform duplicate detection on the candidate data set to obtain a duplicate data set, which may include multiple duplicate texts. Then, the server can extract texts with inconsistent labels from the duplicate data set, obtain the ratio of the amount of texts with inconsistent labels to the amount of data in the data set to be evaluated, and determine the consistency rate of the data set to be evaluated based on the ratio: the consistency rate is p d = 1-the amount of text with inconsistent labels C d / The amount of data in the dataset to be evaluated C.
[0109] S205: Perform anomaly detection evaluation on the candidate data set to obtain the anomaly rate of the data set to be evaluated.
[0110] For example, the server can use the Jieba word segmentation strategy to perform word segmentation on each text in the candidate data set to obtain the word set corresponding to each text; use word2vec to vectorize the word set to obtain the word vector; project the word vector to the sample feature space of a preset dimension to obtain the text feature vector; use LSA to reduce the dimension of the text feature vector to obtain the reduced dimension feature vector; normalize the reduced dimension feature vector to obtain the normalized feature vector; cluster the normalized feature vector through the DBSCAN algorithm, and determine the abnormal data in the candidate data set based on the clustering results. The ratio between the data volume of the abnormal data and the data volume of the data set to be evaluated can be calculated, and the ratio between the data volume of the abnormal data and the data volume of the data set to be evaluated is used as the abnormality rate: abnormality rate = data volume of abnormal data / data volume of the data set to be evaluated.
[0111] For another example, the server can use a trained classification model (such as the FastText model) to perform classification prediction on each text in the candidate data set to obtain the classification probability corresponding to each text in the candidate data set; filter out texts with a classification probability less than a preset threshold as abnormal data; and determine the abnormal rate of the data set to be evaluated based on the filtered abnormal data with a classification probability less than the preset threshold: abnormal rate = the amount of abnormal data with a classification probability less than the preset threshold / the amount of data in the data set to be evaluated.
[0112] For another example, the abnormal data screened out by clustering and the abnormal data screened out by the trained classification model can be combined to determine the abnormal rate p of the data set to be evaluated. o :p o =|D isolate UD noise | / C, where D isolate It can represent the abnormal data filtered out by clustering, D noise It can represent the abnormal data filtered out by the trained classification model, and C can represent the data volume of the data set to be evaluated.
[0113] S206: Perform complexity evaluation on the candidate data set to obtain the complexity rate of the data set to be evaluated.
[0114] For example, the server can use the TEXTCNN model to perform a complexity assessment on the candidate data set. Specifically, the Jieba word segmentation strategy can be used to segment each text in the data set to obtain a word set corresponding to each text. Then use Word2vec to vectorize the word set to obtain a vector matrix. The vector matrix can be a two-dimensional matrix. Three convolution kernels of sizes 2, 3, and 4 are used to perform convolution operations on the vector matrix to obtain multiple one-dimensional vectors. The multiple one-dimensional vectors are subjected to maximum pooling operations through the pooling layer to obtain the pooled vector corresponding to each text. The pooled vectors corresponding to each text are spliced into a vector to obtain a spliced vector. The spliced vectors are fully connected through the fully connected layer to obtain the label category probability of each sample data. The label category with the largest probability is the probability value p of the first category. max , the second largest label category probability is the second category probability value p second , you can filter out p max Less than P th1 , and p second Greater than P th2 Text, get complex sample D complex , so that the complexity rate p of the data set to be evaluated can be calculated c :p c =D complex / C, where C can represent the amount of data in the dataset to be evaluated.
[0115] S207: Determine an evaluation result corresponding to the dataset to be evaluated according to the availability rate, consistency rate, anomaly rate, and complexity rate.
[0116] For example, the server can calculate the cumulative value of availability rate, consistency rate, anomaly rate, and complexity rate: cumulative value = availability rate + consistency rate + anomaly rate + complexity rate. When the cumulative value is greater than or equal to the target threshold, the evaluation result is that the data quality is qualified. When the cumulative value is less than the target threshold, the evaluation result is that the data quality is unqualified.
[0117] For another example, the server can perform weighted operations based on the availability rate, consistency rate, anomaly rate, complexity rate and their corresponding weight values to obtain the target value: target value = availability rate * weight value of availability rate + consistency rate * weight value of consistency rate + anomaly rate * weight value of anomaly rate + complexity rate * weight value of complexity rate. When the target value is greater than or equal to the target threshold, the evaluation result is that the data quality is qualified; when the target value is less than the target threshold, the evaluation result is that the data quality is unqualified.
[0118] It should be noted that the execution order between steps S203 to S206 can be flexibly set according to actual needs. For example, steps S203 to S206 can be executed serially, or in parallel, or in any order.
[0119] The embodiment of the present application can filter the data set to be evaluated, analyze the filtered candidate data set based on multiple different evaluation dimensions to obtain evaluation parameters such as availability rate, consistency rate, anomaly rate, complexity rate, and determine the evaluation result corresponding to the data set to be evaluated based on the evaluation parameters, thereby realizing the evaluation of data quality from multiple dimensions and improving the efficiency and accuracy of data evaluation.
[0120] 2. Model Training
[0121] like Figure 5 As shown, the model training process provided in the embodiment of the present application may include:
[0122] S301: Obtain a target data set whose data quality is qualified according to an evaluation result.
[0123] S302: Predict the target data set using the initial text semantic model to obtain a prediction result.
[0124] S303: Adjust the parameters of the initial text semantic model according to the prediction results to obtain a trained text semantic model.
[0125] When the evaluation result corresponding to the data set to be evaluated is that the data quality is qualified, the server can use the data set to be evaluated with the evaluation result that the data quality is qualified as the target data set, and can use the target data set to train the initial text semantic model to obtain a trained text semantic model. For example, the text in the target data set can be predicted by the initial text semantic model, and the parameters of the initial text semantic model can be adjusted according to the prediction result to obtain a trained text semantic model. The embodiment of the present application can train the text semantic model using the target data set with qualified data quality, thereby improving the accuracy of model training and enhancing the performance of the model.
[0126] 3. Application of the Model
[0127] like Figure 6 As shown, the process of model application provided in the embodiment of the present application may include:
[0128] S401: Obtain the text to be detected generated by the source end.
[0129] The source end may be a terminal such as a mobile phone, a computer or other smart device, and the text to be detected may be text input by the user through the source end, or text obtained by converting the voice input by the user through the source end.
[0130] S402 : Perform semantic analysis on the text to be detected using the trained text semantic model to obtain an analysis result.
[0131] For example, to monitor service quality and public opinion risks by monitoring call center audio, a trained text semantic model can be used to analyze whether the corresponding text to be monitored, such as the audio, exhibits risky tendencies, contains dissatisfaction, or contains unreasonable customer service statements. Another example is the trained text semantic model, which can be used to analyze the category to which the text to be monitored belongs.
[0132] S403: Determine the risk level of the text to be detected based on the analysis results.
[0133] The risk level can be flexibly set based on actual needs. The higher the risk level, the greater the risk. For example, if the analysis results determine that the text to be tested has a high probability of risk, the risk level will be higher. When the risk level is 0, it means that the text to be tested has no risk tendency.
[0134] S404: The source end that generates the text to be detected performs corresponding processing according to the risk level.
[0135] When the risk level is high (for example, the risk level is greater than a preset level threshold, and the preset level threshold can be flexibly set according to actual needs), the server can output a risk warning and the identification of the source end, etc. For example, the risk warning and the identification of the source end can be sent to a designated mailbox or device, etc. for maintenance personnel to view, and take corresponding measures to process the source end that generates the text to be detected, for example, punishing or warning the user corresponding to the source end. When the risk level is low, the server can output a prompt message of no risk or a low risk level for maintenance personnel to view. The embodiment of the present application can analyze the risk level of the text to be detected by the trained text semantic model to take corresponding processing measures, thereby improving the timeliness and reliability of the corresponding processing of the source end that generates the text to be detected.
[0136] To facilitate better implementation of the data evaluation method provided in the embodiment of the present application, the embodiment of the present application also provides a device based on the above data evaluation method. The meanings of the terms are the same as those in the above data evaluation method, and the specific implementation details can be referred to the description in the method embodiment.
[0137] See also Figure 7 , Figure 7This is a structural diagram of a data evaluation device provided in an embodiment of the present application, wherein the data evaluation device may include a first acquisition unit 501, a filtering unit 502, an analysis unit 503, a second acquisition unit 504, and a determination unit 505, etc.
[0138] The first acquisition unit 501 is used to acquire a data set to be evaluated.
[0139] The filtering unit 502 is configured to filter the dataset to be evaluated according to a preset filtering strategy to obtain a candidate dataset.
[0140] The analyzing unit 503 is configured to analyze the candidate data set based on the evaluation dimension.
[0141] The second acquiring unit 504 is configured to acquire evaluation parameters of the dataset to be evaluated according to the analysis result.
[0142] The determining unit 505 is configured to determine an evaluation result corresponding to the data set to be evaluated according to the evaluation parameters, so as to process the data according to the evaluation result.
[0143] In one embodiment, the evaluation dimension includes anomaly detection evaluation, and the evaluation parameter includes anomaly rate. The analysis unit 503 can be specifically used to: perform word segmentation processing on each sample data in the candidate data set to obtain a word set corresponding to each sample data; perform vectorization processing on the word set to obtain a word vector; project the word vector to a sample feature space of a preset dimension to obtain a text feature vector; cluster the text feature vector, and determine the abnormal data in the candidate data set based on the clustering results; the second acquisition unit 504 can be specifically used to: determine the abnormal rate of the data set to be evaluated based on the abnormal data.
[0144] In one embodiment, the analysis unit 503 can be specifically used to: perform dimensionality reduction processing on the text feature vector to obtain a reduced-dimensionality feature vector; perform normalization processing on the reduced-dimensionality feature vector to obtain a normalized feature vector; cluster the normalized feature vector, and determine abnormal data in the candidate data set based on the clustering results.
[0145] In one embodiment, the analysis unit 503 can be specifically used to: discretize the normalized feature vector into multiple feature points; select any one feature point from the multiple feature points as the core point; assign all feature points within a preset neighborhood range centered on the core point to the same class family; select another feature point from the multiple feature points as the core point, and return to execute the operation of assigning all feature points within a preset neighborhood range centered on the core point to the same class family until multiple feature points are traversed; filter out sample data corresponding to feature points that are not assigned to a class family to obtain abnormal data.
[0146] In one embodiment, the evaluation dimension includes anomaly detection evaluation, and the evaluation parameter includes anomaly rate. The analysis unit 503 can be specifically used to: perform classification prediction on each sample data in the candidate data set through the trained classification model to obtain the classification probability corresponding to each sample data in the candidate data set; screen out sample data with a classification probability less than a preset threshold as abnormal data; the second acquisition unit 504 can be specifically used to: determine the abnormal rate of the data set to be evaluated based on the screened abnormal data with a classification probability less than a preset threshold.
[0147] In one embodiment, the data evaluation device may further include:
[0148] a third acquiring unit, configured to acquire a plurality of training samples, the plurality of training samples including labeled samples and complementary labeled samples, wherein the labeled samples are samples labeled with true labels, and the complementary labeled samples are samples labeled with labels complementary to the true labels;
[0149] A training unit is used to perform negative learning on the initial classification model based on the labeled samples and the complementary labeled samples to train the initial classification model, obtain a candidate classification model, and predict the sample classification probability of each training sample;
[0150] A screening unit is used to screen out training samples whose sample classification probability is greater than a target probability threshold to obtain candidate training samples;
[0151] The fine-tuning unit is used to fine-tune the candidate classification model based on the candidate training samples to obtain a trained classification model.
[0152] In one embodiment, the second acquisition unit 504 can be specifically used to: obtain the target feature vector corresponding to each sample data in the candidate data set; determine the target abnormal data in the candidate data set based on the target feature vector; determine the abnormal rate of the data set to be evaluated based on the target abnormal data and the screened abnormal data with a classification probability less than a preset threshold.
[0153] In one embodiment, the evaluation dimension includes complexity evaluation, and the evaluation parameter includes complexity rate. The analysis unit 503 can be specifically used to: perform word segmentation processing on each sample data in the data set to obtain a word set corresponding to each sample data; vectorize the word set to obtain a vector matrix; perform convolution operation on the vector matrix to obtain multiple one-dimensional vectors; perform pooling operation on the multiple one-dimensional vectors to obtain pooled vectors; splice the pooled vectors to obtain spliced vectors; perform full connection operation on the spliced vectors to obtain the label category probability of each sample data; the second acquisition unit 504 can be specifically used to: determine the complexity rate of the data set to be evaluated based on the label category probability.
[0154] In one embodiment, the evaluation dimension includes availability evaluation, and the evaluation parameter includes availability rate. The analysis unit 503 can be specifically used to: obtain a first data volume of the data set to be evaluated and a second data volume of the candidate data set; the second acquisition unit 504 can be specifically used to: calculate the availability rate of the data set to be evaluated based on the first data volume and the second data volume.
[0155] In one embodiment, the evaluation dimension includes consistency evaluation, and the evaluation parameter includes consistency rate. The analysis unit 503 can be specifically used to: perform duplicate detection on the candidate data set to obtain a duplicate data set; extract data with consistent labels from the duplicate data set;
[0156] The second obtaining unit 504 may be specifically configured to: obtain a ratio of the amount of data with consistent labels to the amount of data in the dataset to be evaluated; and determine a consistency rate of the dataset to be evaluated according to the ratio.
[0157] In one embodiment, the evaluation parameters include at least one of an availability rate, a consistency rate, an abnormality rate, and a complexity rate.
[0158] In one embodiment, the determination unit 505 can be specifically used to: obtain the cumulative values of the availability rate, consistency rate, abnormality rate, and complexity rate, and determine the evaluation result corresponding to the data set to be evaluated based on the cumulative values; or, set weight values for the availability rate, consistency rate, abnormality rate, and complexity rate respectively, perform weighted operations based on the availability rate, consistency rate, abnormality rate, complexity rate and their corresponding weight values to obtain a target value, and the target value determines the evaluation result corresponding to the data set to be evaluated.
[0159] In one embodiment, the data evaluation device may further include:
[0160] The processing unit is used to train the training model using the data set to be evaluated to obtain a trained model when the evaluation result corresponding to the data set to be evaluated is that the data quality is qualified; when the evaluation result corresponding to the data set to be evaluated is that the data quality is unqualified, analyze the factors affecting the quality of the data set to be evaluated.
[0161] In this embodiment of the present application, a first acquisition unit 501 acquires a dataset to be evaluated, and a filtering unit 502 filters the dataset to be evaluated according to a preset filtering strategy to obtain a candidate dataset. An analysis unit 503 then analyzes the candidate dataset based on evaluation dimensions, and a second acquisition unit 504 obtains evaluation parameters for the dataset to be evaluated based on the analysis results. At this point, a determination unit 505 determines the evaluation result corresponding to the dataset to be evaluated based on the evaluation parameters, and processes the data based on the evaluation results. This solution can filter the dataset to be evaluated, analyze the filtered candidate dataset based on the evaluation dimensions to obtain evaluation parameters, and determine the evaluation result corresponding to the dataset to be evaluated based on the evaluation parameters, thereby avoiding manual review and evaluation and improving the efficiency, accuracy, and reliability of data evaluation.
[0162] The present application also provides a computer device, which may be a server or a terminal, such as Figure 8 , which shows a schematic diagram of the structure of the computer device involved in the embodiment of the present application, specifically:
[0163] The computer device may include one or more processing core processors 601, one or more computer readable storage media memories 602, a power supply 603, an input unit 604 and other components. Those skilled in the art will understand that Figure 8 The computer device structure shown in the figure does not constitute a limitation on the computer device, and may include more or fewer components than shown in the figure, or combine certain components, or arrange components differently.
[0164] Processor 601 is the control center of the computer device. It connects the various components of the entire computer device using various interfaces and lines. By running or executing software programs and / or modules stored in memory 602 and accessing data stored in memory 602, it performs various functions of the computer device and processes data, thereby performing overall testing of the computer device. Optionally, processor 601 may include one or more processing cores; preferably, processor 601 may integrate an application processor and a modem processor, wherein the application processor primarily processes the operating system, user interface, and application programs, while the modem processor primarily handles wireless communications. It is understood that the modem processor may not be integrated into processor 601.
[0165] The memory 602 can be used to store software programs and modules. The processor 601 executes various functional applications and data processing by running the software programs and modules stored in the memory 602. The memory 602 may mainly include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function (such as a sound playback function, an image playback function, etc.), etc.; the data storage area may store data created according to the use of the computer device, etc. In addition, the memory 602 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other volatile solid-state storage device. Accordingly, the memory 602 may also include a memory controller to provide the processor 601 with access to the memory 602.
[0166] The computer device also includes a power supply 603 for supplying power to various components. Preferably, the power supply 603 can be logically connected to the processor 601 via a power management system, thereby enabling the power management system to manage charging, discharging, and power consumption. The power supply 603 can also include one or more DC or AC power supplies, a recharging system, a power failure detection circuit, a power converter or inverter, a power status indicator, and other arbitrary components.
[0167] The computer device may further include an input unit 604, which may be configured to receive input digital or character information and generate keyboard, mouse, joystick, optical or trackball signal input related to user settings and function control.
[0168] Although not shown, the computer device may further include a display unit, etc., which will not be described in detail here. Specifically, in this embodiment, the processor 601 in the computer device will load the executable files corresponding to one or more application processes into the memory 602 according to the following instructions, and the processor 601 will run the application stored in the memory 602 to implement various functions as follows:
[0169] Obtain the data set to be evaluated; filter the data set to be evaluated according to the preset filtering strategy to obtain a candidate data set; analyze the candidate data set based on the evaluation dimension; obtain the evaluation parameters of the data set to be evaluated based on the analysis results; determine the evaluation result corresponding to the data set to be evaluated based on the evaluation parameters, and process the data according to the evaluation results.
[0170] In the above embodiments, the description of each embodiment has its own focus. For the part that is not described in detail in a certain embodiment, please refer to the detailed description of the data evaluation method above, and will not be repeated here.
[0171] According to one aspect of the present application, a computer program product or computer program is provided, the computer program product or computer program including computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the methods provided in various optional implementations of the above embodiments.
[0172] Those skilled in the art will appreciate that all or part of the steps in the various methods of the above embodiments can be accomplished by computer instructions, or by controlling related hardware via computer instructions. The computer instructions can be stored in a computer-readable storage medium and loaded and executed by a processor. To this end, an embodiment of the present application provides a storage medium having a computer program stored therein. The computer program can include computer instructions, and the computer program can be loaded by a processor to execute any of the data evaluation methods provided in the embodiments of the present application.
[0173] The specific implementation of the above operations can be found in the previous embodiments and will not be repeated here.
[0174] The storage medium may include a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.
[0175] Since the instructions stored in the storage medium can execute the steps in any data evaluation method provided in the embodiments of the present application, the beneficial effects that can be achieved by any data evaluation method provided in the embodiments of the present application can be achieved. Please refer to the previous embodiments for details and will not be repeated here.
[0176] The above is a detailed introduction to a data evaluation method, device, computer equipment and storage medium provided in the embodiments of the present application. Specific examples are used herein to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core ideas. At the same time, for technical personnel in this field, based on the ideas of the present application, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.
Claims
1. A data evaluation method, characterized in that: include: Get the dataset to be evaluated; Filter the dataset to be evaluated according to a preset filtering strategy to obtain a candidate dataset; Analyzing the candidate data set based on the evaluation dimension; Obtaining evaluation parameters of the data set to be evaluated according to the analysis results; determining an evaluation result corresponding to the dataset to be evaluated according to the evaluation parameters, so as to process the dataset to be evaluated according to the evaluation result; When the evaluation dimension includes complexity evaluation, the evaluation parameter includes complexity rate, and analyzing the candidate data set based on the evaluation dimension and obtaining the evaluation parameter of the data set to be evaluated according to the analysis result includes: Each sample data in the data set is segmented to obtain a word set corresponding to each sample data; the word set is vectorized to obtain a vector matrix; a convolution operation is performed on the vector matrix to obtain multiple one-dimensional vectors; a pooling operation is performed on the multiple one-dimensional vectors to obtain a pooled vector; the pooled vectors are spliced to obtain a spliced vector; a full connection operation is performed on the spliced vector to obtain a label category probability for each sample data; and the complexity rate of the data set to be evaluated is determined based on the label category probability.
2. The data evaluation method according to claim 1, characterized in that The evaluation dimension includes anomaly detection evaluation, the evaluation parameter includes anomaly rate, and the candidate data set is analyzed based on the evaluation dimension, and the evaluation parameter of the data set to be evaluated is obtained according to the analysis result, including: Perform word segmentation on each sample data in the candidate data set to obtain a word set corresponding to each sample data; Performing vectorization processing on the word set to obtain word vectors; Projecting the word vector to a sample feature space of a preset dimension to obtain a text feature vector; Clustering the text feature vectors, and determining abnormal data in the candidate data set based on the clustering results; The abnormality rate of the dataset to be evaluated is determined according to the abnormal data.
3. The data evaluation method according to claim 2, characterized in that: Clustering the text feature vectors and determining abnormal data in the candidate data set based on the clustering results includes: Performing dimensionality reduction processing on the text feature vector to obtain a feature vector after dimensionality reduction; Normalize the eigenvector after dimensionality reduction to obtain a normalized eigenvector; Clustering is performed on the normalized feature vectors, and abnormal data in the candidate data set is determined based on the clustering result.
4. The data evaluation method according to claim 3, characterized in that: Clustering the normalized feature vectors and determining abnormal data in the candidate data set based on the clustering results includes: discretizing the normalized feature vector into a plurality of feature points; Select any one feature point from the plurality of feature points as a core point; Allocating all feature points within a preset neighborhood centered on the core point to the same class family; Select another feature point from the plurality of feature points as a core point, and return to perform an operation of assigning all feature points within a preset neighborhood range centered on the core point to the same cluster until all the feature points are traversed; The sample data corresponding to the feature points that are not assigned to the class family are filtered out to obtain abnormal data.
5. The data evaluation method according to claim 1, characterized in that The evaluation dimension includes anomaly detection evaluation, the evaluation parameter includes anomaly rate, and the candidate data set is analyzed based on the evaluation dimension, and the evaluation parameter of the data set to be evaluated is obtained according to the analysis result, including: Perform classification prediction on each sample data in the candidate data set using the trained classification model to obtain the classification probability corresponding to each sample data in the candidate data set; Filter out sample data with a classification probability less than a preset threshold as abnormal data; The abnormal rate of the data set to be evaluated is determined based on the screened abnormal data with a classification probability less than a preset threshold.
6. The data evaluation method according to claim 5, characterized in that: The data evaluation method further comprises: Acquire multiple training samples, wherein the multiple training samples include labeled samples and complementary labeled samples, wherein the labeled samples are samples labeled with true labels, and the complementary labeled samples are samples labeled with labels complementary to the true labels; Performing negative learning based on the labeled sample and the complementary labeled sample by the initial classification model to train the initial classification model, obtain a candidate classification model, and predict a sample classification probability of each training sample; Filter out training samples whose sample classification probability is greater than the target probability threshold to obtain candidate training samples; Fine-tune the candidate classification model based on the candidate training samples to obtain a trained classification model.
7. The data evaluation method according to claim 5, characterized in that: Determining the abnormality rate of the dataset to be evaluated based on the abnormal data whose classification probability is less than a preset threshold after being screened out includes: Obtaining a target feature vector corresponding to each sample data in the candidate data set; determining target abnormal data in the candidate data set based on the target feature vector; The abnormality rate of the data set to be evaluated is determined based on the target abnormal data and the screened abnormal data with a classification probability less than a preset threshold.
8. The data evaluation method according to claim 1, characterized in that The evaluation dimension includes availability evaluation, the evaluation parameter includes availability, and the analyzing the candidate dataset based on the evaluation dimension and obtaining the evaluation parameter of the dataset to be evaluated according to the analysis result includes: Acquiring a first data volume of the dataset to be evaluated and a second data volume of the candidate dataset; The availability of the to-be-evaluated data set is calculated according to the first data volume and the second data volume.
9. The data evaluation method according to claim 1, characterized in that: The evaluation dimension includes consistency evaluation, the evaluation parameter includes consistency rate, and the analyzing the candidate data set based on the evaluation dimension and obtaining the evaluation parameter of the data set to be evaluated according to the analysis result includes: Performing duplicate detection on the candidate data set to obtain a duplicate data set; Extracting data with consistent labels from the repeated dataset; Obtaining a ratio of the amount of data with consistent labels to the amount of data in the dataset to be evaluated; The consistency rate of the dataset to be evaluated is determined according to the ratio.
10. The data evaluation method according to claim 1, characterized in that: The evaluation parameters include at least one of an availability rate, a consistency rate, an anomaly rate, and a complexity rate. Determining the evaluation result corresponding to the to-be-evaluated data set according to the evaluation parameters includes: Obtaining the accumulated values of the availability rate, consistency rate, anomaly rate, and complexity rate, and determining the evaluation result corresponding to the data set to be evaluated according to the accumulated values; or, Weight values are set for the availability rate, consistency rate, anomaly rate, and complexity rate respectively, and a weighted operation is performed based on the availability rate, consistency rate, anomaly rate, complexity rate and their corresponding weight values to obtain a target value, which determines the evaluation result corresponding to the data set to be evaluated.
11. The data evaluation method according to any one of claims 1 to 10, characterized in that: After determining the evaluation result corresponding to the to-be-evaluated data set according to the evaluation parameters, the data evaluation method further includes: When the evaluation result corresponding to the data set to be evaluated is that the data quality is qualified, the training model is trained using the data set to be evaluated to obtain a trained model, and the data is processed by the trained model; When the evaluation result corresponding to the dataset to be evaluated is that the data quality is unqualified, factors affecting the quality of the dataset to be evaluated are analyzed.
12. A data evaluation device, characterized in that: include: A first acquisition unit is used to acquire a data set to be evaluated; A filtering unit, configured to filter the dataset to be evaluated according to a preset filtering strategy to obtain a candidate dataset; an analyzing unit, configured to analyze the candidate data set based on an evaluation dimension; A second acquiring unit, configured to acquire evaluation parameters of the data set to be evaluated according to the analysis result; a determining unit, configured to determine an evaluation result corresponding to the dataset to be evaluated according to the evaluation parameters, so as to process the dataset to be evaluated according to the evaluation result; When the evaluation dimension includes complexity evaluation, the evaluation parameter includes complexity rate, and analyzing the candidate data set based on the evaluation dimension and obtaining the evaluation parameter of the data set to be evaluated according to the analysis result includes: Each sample data in the data set is segmented to obtain a word set corresponding to each sample data; the word set is vectorized to obtain a vector matrix; a convolution operation is performed on the vector matrix to obtain multiple one-dimensional vectors; a pooling operation is performed on the multiple one-dimensional vectors to obtain a pooled vector; the pooled vectors are spliced to obtain a spliced vector; a full connection operation is performed on the spliced vector to obtain a label category probability for each sample data; and the complexity rate of the data set to be evaluated is determined based on the label category probability.
13. A computer device, characterized in that: The method comprises a processor and a memory, wherein a computer program is stored in the memory, and when the processor calls the computer program in the memory, the data evaluation method according to any one of claims 1 to 11 is executed.
14. A storage medium, characterized in that The storage medium is used to store a computer program, and the computer program is loaded by a processor to execute the data evaluation method according to any one of claims 1 to 11.
15. A computer program product, characterized in that The computer program product includes computer instructions, which are stored in a computer-readable storage medium; a processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the data evaluation method described in any one of claims 1 to 11.
Citation Information
Patent Citations
Mining performance evaluation method based on industrial equipment data
CN108197280A
Text classification method and device, electronic equipment and computer readable storage medium
CN110717039A
Abnormal data determination method and device, storage medium and electronic equipment
CN111831704A