Detection Method and Device for Abnormal Samples, Electronic Device, Storage Medium

By using the verification set to detect abnormal samples in machine learning models and using similarity calculation to filter out the sample with labeling errors, the poor prediction effect caused by the labeling errors in the training set is solved, and efficient abnormal sample detection and cost reduction are achieved.

CN114077859BActive Publication Date: 2025-07-11ALIBABA GROUP HOLDING LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202010827028.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-08-17
Publication Date
2025-07-11
Estimated Expiration
2040-08-17

AI Technical Summary

Technical Problem

In the prior art, sample data labeling errors in the training set lead to poor prediction results in machine learning models, making it difficult to effectively detect and correct abnormal samples.

Method used

By obtaining the sample library and verification set to be detected, using machine learning models to predict the verification set, determining that the sample data whose prediction results are inconsistent with the standard category is an abnormal sample, and using similarity calculations to filter out sample data that may have labeling errors from similar samples.

Benefits of technology

Accurately screening out abnormal samples in the training set reduces detection costs, improves the prediction accuracy and generalization capabilities of the model, and reduces the workload of manual verification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114077859B_ABST
    Figure CN114077859B_ABST
Patent Text Reader

Abstract

One or more embodiments of this specification provide a method and apparatus for detecting abnormal samples, an electronic device, and a storage medium. The method may include: obtaining a sample library to be detected and a validation set corresponding to the sample library to be detected, where the sample data in the sample library to be detected is labeled with a data category, and the sample data in the validation set is labeled with a standard category; training a machine learning model using the sample library to be detected, and using the machine learning model to predict the sample data in the validation set; and determining the target sample data in the validation set whose prediction result is inconsistent with the standard category as an abnormal sample.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] One or more embodiments of this specification relate to the field of artificial intelligence technology, and in particular, to a method and device for detecting abnormal samples, an electronic device, and a storage medium. Background Art

[0002] In related technologies, machine learning technology can use algorithms to learn from existing data and make judgments and decisions on real-world situations. Machine learning technology includes supervised learning, unsupervised learning, semi-supervised learning, reinforcement learning, and so on.

[0003] For the training process of supervised learning, the input sample data is called a "training set", and the sample data in the training set has a clear identification or result (i.e., a sample label). When using a supervised learning algorithm to establish a prediction model, the supervised learning algorithm establishes a learning process to compare the prediction result with the actual result of the "training set" and continuously adjusts the prediction model until the prediction result of the model reaches an expected accuracy rate.

[0004] Common application scenarios of supervised learning include classification problems, regression problems, etc. Common algorithms include logistic regression, neural networks, decision trees, support vector machines, Bayesian classifiers, and so on. Summary of the Invention

[0005] In view of this, one or more embodiments of this specification provide a method and device for detecting abnormal samples, an electronic device, and a storage medium.

[0006] To achieve the above object, one or more embodiments of this specification provide the following technical solutions:

[0007] According to the first aspect of one or more embodiments of this specification, a method for detecting abnormal samples is proposed, including:

[0008] Obtain a sample library to be detected and a validation set corresponding to the sample library to be detected. The sample data in the sample library to be detected is labeled with data categories, and the sample data in the validation set is labeled with standard categories;

[0009] Use the sample library to be detected to train a machine learning model, and use the machine learning model to predict the sample data in the validation set;

[0010] Determine the target sample data in the validation set whose prediction result is inconsistent with the standard category as abnormal samples.

[0011] Optionally, it further includes:

[0012] Obtain first sample data with a predicted result of the target category in the sample library to be detected, where the target category is the predicted result of second sample data in the target sample data;

[0013] Calculate the similarity between the first sample data and the second sample data;

[0014] Determine whether the first sample data is an abnormal sample according to the similarity;

[0015] Optionally, it further includes:

[0016] Divide the sample data in the sample library to be detected into N parts, where N is an integer greater than or equal to 2;

[0017] Use M of them as a training set to train the machine learning model, and use the machine learning model to predict the sample data of N - M parts;

[0018] If the predicted result of the third sample data in the sample library to be detected is inconsistent with the labeled data category of the third sample data, regard the third sample data as an abnormal sample.

[0019] Optionally, the using the sample library to be detected to train the machine learning model includes:

[0020] Adopt a neural network algorithm to train the sample library to be detected to obtain a neural network model;

[0021] Among them, when the error between the predicted result of the sample data in the validation set and the corresponding standard category during the iteration process meets the preset error condition, use the model parameters at this time as the model parameters of the neural network model.

[0022] Optionally, the calculating the similarity between the first sample data and the second sample data includes:

[0023] Determine the first vector data calculated by the first sample data in the preset intermediate layer of the neural network model;

[0024] Determine the second vector data calculated by the second sample data in the preset intermediate layer;

[0025] Calculate the similarity between the first vector data and the second vector data.

[0026] Optionally, the preset intermediate layer at least includes the layer before the output layer of the neural network model.

[0027] Optionally, the determining whether the first sample data is an abnormal sample according to the similarity includes:

[0028] Sort multiple first sample data according to the magnitude of the similarity;

[0029] Select at least one first sample data in sequence according to the ranking to obtain the detection result of the abnormal sample in the selected first sample data until the obtained detection result meets the preset detection condition.

[0030] Optionally, the number of first sample data selected each time is positively correlated with the number of selections.

[0031] According to the second aspect of one or more embodiments of the present specification, a detection device for abnormal samples is proposed, including:

[0032] A first acquisition unit that acquires a sample library to be detected and a validation set corresponding to the sample library to be detected, wherein the sample data in the sample library to be detected is labeled with a data category, and the sample data in the validation set is labeled with a standard category;

[0033] A first training unit that trains a machine learning model using the sample library to be detected and uses the machine learning model to predict the sample data in the validation set;

[0034] A first detection unit that determines the target sample data with inconsistent prediction results and standard categories in the validation set as abnormal samples.

[0035] Optionally, it further includes:

[0036] A second acquisition unit that obtains first sample data with a predicted result of a target category in the sample library to be detected, where the target category is the predicted result of the second sample data in the target sample data;

[0037] A calculation unit that calculates the similarity between the first sample data and the second sample data;

[0038] A second detection unit that determines whether the first sample data is an abnormal sample according to the similarity.

[0039] Optionally, it further includes:

[0040] A division unit that divides the sample data in the sample library to be detected into N parts, where N is an integer greater than or equal to 2;

[0041] A second training unit that uses M of them as a training set to train the machine learning model and uses the machine learning model to predict the sample data of N - M parts;

[0042] A third detection unit that, if the predicted result of the third sample data in the sample library to be detected is inconsistent with the data category labeled for the third sample data, regards the third sample data as an abnormal sample.

[0043] Optionally, the first training unit is specifically configured to:

[0044] Train the sample library to be detected by using a neural network algorithm to obtain a neural network model;

[0045] Among them, the model parameters of the neural network model are the model parameters when the error between the prediction result of the sample data in the verification set and the corresponding standard category meets the preset error condition during the iteration process.

[0046] Optionally, the first training unit is further configured to:

[0047] Determine first vector data calculated by the first sample data in a preset intermediate layer of the neural network model;

[0048] Determine second vector data calculated by the second sample data in the preset intermediate layer;

[0049] Calculate the similarity between the first vector data and the second vector data.

[0050] Optionally, the preset intermediate layer at least includes the layer before the output layer of the neural network model.

[0051] Optionally, the first detection unit is specifically configured to:

[0052] Sort multiple first sample data according to the magnitude of the similarity;

[0053] Select at least one first sample data in sequence according to the ranking to obtain the detection result of abnormal samples in the selected first sample data until the obtained detection result meets the preset detection condition.

[0054] Optionally, the number of first sample data selected each time is positively correlated with the number of selections.

[0055] According to a third aspect of one or more embodiments of the present specification, an electronic device is proposed, including:

[0056] A processor;

[0057] A memory for storing instructions executable by the processor;

[0058] Among them, the processor realizes the method as described in any one of the above embodiments by running the executable instructions.

[0059] According to a fourth aspect of one or more embodiments of the present specification, a computer-readable storage medium is proposed, on which computer instructions are stored, and when the instructions are executed by a processor, the steps of the method as described in any one of the above embodiments are realized.

[0060] As can be seen from the above embodiments, in the technical solution provided in this specification, a validation set corresponding to the sample library is configured, and the sample data in the validation set is labeled with standard categories (the standard labels are all correct by default, and there are no labeling errors). Then, the sample data in the validation set can be used to verify the machine learning model trained based on the sample library, and thus the sample data with incorrect predictions in the validation set (i.e., the target sample data) is used as the abnormal sample.

[0061] Furthermore, the similar sample data in the sample library similar to it may also be abnormal samples with incorrect labels. Then, sample data can be selected from these similar sample data for further verification to determine whether it is an abnormal sample.

[0062] On the one hand, the above machine learning model is trained by the sample library. By selecting sample data from the sample data similar to the target sample data with incorrect predictions for verification, the non-outlier systematic error samples in the sample library can be accurately screened for detection. On the other hand, by the above method of selecting sample data for anomaly detection, the amount of sample data to be detected can be reduced, thereby effectively reducing the detection cost of abnormal samples. Description of the Drawings

[0063] Figure 1 is a flowchart of a method for detecting abnormal samples provided by an exemplary embodiment.

[0064] Figure 2 is a schematic diagram of a neural network model provided by an exemplary embodiment.

[0065] Figure 3 is a flowchart of a method for detecting abnormal texts provided by an exemplary embodiment.

[0066] Figure 4 is a flowchart of another method for detecting abnormal texts provided by an exemplary embodiment.

[0067] Figure 5 is a schematic diagram of the structure of a device provided by an exemplary embodiment.

[0068] Figure 6 is a block diagram of a device for detecting abnormal texts provided by an exemplary embodiment. Detailed Embodiments

[0069] Exemplary embodiments will be described in detail herein, and examples thereof are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with one or more embodiments of this specification. On the contrary, they are merely examples of devices and methods consistent with some aspects of one or more embodiments of this specification as detailed in the appended claims.

[0070] It should be noted that: in other embodiments, the steps of the corresponding methods are not necessarily executed in the order shown and described in this specification. In some other embodiments, the steps included in the method may be more or less than those described in this specification. In addition, a single step described in this specification may be decomposed into multiple steps for description in other embodiments; and multiple steps described in this specification may also be combined into a single step for description in other embodiments.

[0071] When training a machine learning model using a supervised learning algorithm, the sample data (i.e., the training set) used for training needs to be labeled with labels, that is, to establish the mapping relationship between the sample data and the labels, and construct "input-output" pairs, so as to input them into the supervised learning algorithm for training. Among them, the training set is used as the correct information for the supervised learning algorithm to learn the law reflected by the mapping relationship from input to output. Then, after the training is completed, other data can be predicted. It can be seen that the training process of the supervised learning algorithm depends on the labeling of the sample data. If there are many sample data with mislabeled labels (hereinafter referred to as abnormal samples) in the training set, it will greatly affect the prediction effect of the trained machine learning model.

[0072] This specification aims to provide a detection scheme for abnormal samples, which can reduce the screening cost while effectively screening out abnormal samples in the sample library used for training the model.

[0073] Please refer to Figure 1 , Figure 1 which is a flowchart of a method for detecting abnormal samples provided by an exemplary embodiment. As Figure 1 shown, this method is applied to any electronic device that can be used to train a machine learning model, and may include the following steps:

[0074] Step 102, obtain a sample library to be detected and a validation set corresponding to the sample library to be detected. The sample data in the sample library to be detected is labeled with data categories, and the sample data in the validation set is labeled with standard categories.

[0075] In this embodiment, the sample library to be detected can be the sample library used for training a machine learning model subsequently, that is, the sample library serves as the training set of the machine learning model. There is a corresponding label for the sample data in the sample library, and what this label annotates can be the data category of the sample data; of course, the specific content of the label can be flexibly set according to actual needs, and this specification does not limit this. The technical solution provided in this specification is precisely to detect mislabeled labels (that is, actually there is no corresponding relationship between this incorrect label and the sample data), so the label of the sample data in the sample library is the label to be detected. For the validation set corresponding to the sample library to be detected, it is used to verify the prediction effect of the trained machine learning model. Therefore, the label annotated for the sample data in this validation set is defaulted to be the correct label (hereinafter referred to as the standard label), that is, there are no abnormal samples. For example, when detecting the data category annotated for the sample data in the sample library, the sample data in the corresponding validation set is annotated with the standard category.

[0076] For example, the sample data in the sample library to be detected and the validation set can be text data; in this case, the label annotated for the text data can be the text category of the text data, that is, the sample library to be detected is used to train a machine learning model that can recognize text categories. Of course, the sample data and the corresponding label can also be any other type of data, and one or more embodiments of this specification do not limit this. For example, the sample data can also be image data, audio data, etc. Taking face recognition as an example, the image data is a face image, and the corresponding label is the user identifier (ID number, username, user account, etc.) corresponding to this face image. For the convenience of description, subsequent descriptions will all be based on the sample data being text as an example, and the detection principle for other types of sample data is similar.

[0077] Illustrated in combination with the customer service scenario, the customer service platform can train a machine learning model for text classification to predict the user's input question and obtain the knowledge point to which the user's question belongs, and then determine the answering method for the user's question according to this knowledge point; for example, a detailed explanation for the knowledge point to which the user's question belongs can be returned. For the above application scenario, it is necessary to construct a text library (i.e., the sample library) for training the above model. The content of the text data in this text library is the user's question and is annotated with a label indicating the knowledge point to which the question belongs. However, in actual situations, due to reasons such as the complexity of the business itself, the difficulty of the annotation work, the carelessness of the annotators, the inconsistency before and after due to changes in the business scope, and the non-uniformity of the knowledge base maintenance standards, there are many mislabeled labels in the text library.

[0078] Step 104, use the sample library to be detected to train a machine learning model, and use the machine learning model to predict the sample data in the validation set.

[0079] In this embodiment, the sample library to be detected can be used as a training set, and a supervised machine learning model can be obtained by training the training set using a supervised learning algorithm. Based on the trained machine learning model, the above-mentioned validation set can be used as a probe to detect abnormal samples with annotation errors in the sample library to be detected. Among them, based on the characteristic of using the validation set as a probe, the data scale of the validation set can be smaller than that of the sample library to be detected. For example, the number of sample data included in the validation set is 1 / 100 of the number of sample data included in the sample library to be detected; of course, the data scale of the validation set can be flexibly set according to the actual situation, and this specification does not limit this.

[0080] Step 106, determine the target sample data with inconsistent prediction results and standard categories in the validation set as abnormal samples.

[0081] In this embodiment, the validation set can be used as a whole to detect the sample library to be detected, and at the same time, random annotation errors (outliers) and some systematic annotation errors (non-outliers) can be detected. Specifically, after obtaining the target sample data with incorrect predictions of the trained machine learning model on the validation set, the detection results for the abnormal samples in the target sample data can be obtained, and then the detection results corresponding to the above-mentioned target sample data can be used as the detection results for the abnormal samples in the sample library to be detected. Among them, the above detection results corresponding to the target sample data can be output by the above electronic device for manual detection by the annotator; then, the annotator can modify the detected annotation errors and input them into the electronic device.

[0082] In this embodiment, in addition to using the entire validation set as a probe to detect the sample library to be detected as described above, each sample in the validation set can also be used as a probe to detect abnormal samples, so as to accurately screen out systematic error samples belonging to non-outliers.

[0083] Specifically, input any sample data in the validation set into the machine learning model. If the prediction result output by the machine learning model is inconsistent with the standard label of the sample data, it is determined that the prediction for the sample data is incorrect; among them, each sample data with incorrect predictions in the validation set constitutes the target sample data. Then, the process of using each sample in the validation set as a probe for detection can include: obtaining the first sample data with a predicted result of the target category in the sample library to be detected, where the target category is the predicted result of the second sample data in the target sample data (the second sample data can be any sample data in the target sample data, that is, the second sample data is the sample data with incorrect predictions in the validation set); calculating the similarity between the first sample data and the second sample data, and thus determining whether the first sample data is an abnormal sample according to the calculated similarity.

[0084] The principle of using the validation set as a probe to detect abnormal samples is as follows: Use the machine learning model trained based on the sample library to be detected to predict the sample data in the validation set. For each sample data predicted incorrectly in the validation set (i.e., the above-mentioned second sample data), search for the similar sample data in the sample library to be detected (i.e., the above-mentioned first sample data), and the detected sample data may include abnormal samples with incorrect label annotations. Those skilled in the art should understand that: for any sample data predicted incorrectly by the machine learning model in the validation set, since the machine learning model is trained by the sample library to be detected, that is, the rules learned by the machine learning model can reflect the mapping relationship between the sample data and the labels to be detected in the sample library to be detected, and if this sample data is predicted incorrectly, then the similar sample data in the sample library to be detected is very likely to have a similar error. Therefore, sample data can be selected from these similar sample data for further verification to determine whether it is an abnormal sample. For example, similar sample data can be selected based on the calculated similarity, and the detection results for the abnormal samples in the selected similar sample data can be obtained. Similarly to the above, the detection results of the above-mentioned similar sample data can be output by the above-mentioned electronic device for manual detection by the annotator; then, the annotator can modify the detected incorrect annotation and input it into the electronic device.

[0085] Regarding abnormal samples, there are two types: non-outlier systematic error samples (non-outliers) and outlier error samples (outliers). Taking text as an example for illustration. For example, if a certain text and other texts with the same text category are all mislabeled as a wrong category, then these texts belong to non-outlier systematic error samples. And for the case of accidental random mislabeling (such as only one text's text category is mislabeled while the annotations of other texts with the same text category are correct), it belongs to outlier error samples.

[0086] On the one hand, the above-mentioned machine learning model is trained by this sample library. By selecting sample data from the sample data similar to the target sample data predicted incorrectly for verification, non-outlier systematic error samples in the sample library can be accurately screened for detection. On the other hand, through the above method of selecting sample data for abnormal detection, the amount of sample data to be detected can be reduced, thereby effectively reducing the detection cost of abnormal samples.

[0087] As an exemplary embodiment, a neural network algorithm can be used to train a sample library to be detected to obtain a neural network model, so as to obtain better generalization performance, that is, the sample data can be better fitted during the training process. To prevent the neural network from overfitting (that is, the error rate of the neural network on the training set becomes lower and lower, while the actual prediction effect decreases instead), a validation set can be used to evaluate the generalization ability of the neural network model. For example, the model parameters of the neural network model are the model parameters when the error between the prediction result of the sample data in the validation set and the corresponding standard category during the iteration process meets the preset error condition. Among them, the preset error condition can be flexibly adjusted according to the actual situation, and this specification does not limit this.

[0088] For example, the model parameters corresponding to the case where the above error is the smallest during the iteration process can be used as the model parameters of the neural network model. For example, during the iteration process of training, in each iteration cycle, the prediction effect is evaluated on the validation set using the current model, and then the model parameters when the prediction effect on the validation set is the best in each iteration cycle are used as the final model parameters of the neural network model.

[0089] Alternatively, the model parameters corresponding to the case where the above error does not exceed the preset error threshold are used as the model parameters of the neural network model. For example, due to the strong fitting ability of the neural network, the prediction effect on the validation set will first increase and then gradually decrease during the iteration process. Therefore, the error of the current model on the validation set can be calculated in each iteration cycle until the error of the current model on the validation set is larger than the error of the previous iteration cycle, then the training is stopped, and the model parameters in the previous iteration cycle are used as the final model parameters of the neural network model.

[0090] Continuing with the above embodiment of training a neural network model using a neural network algorithm, for the operation of calculating similarity in step 106, since a neural network is used for training, the feature encoding ability of the neural network can be fully utilized, and the vector calculated by the intermediate layer (hidden layer) of the neural network is used to represent the corresponding sample data, so as to participate in the similarity calculation instead of the sample data, which can effectively improve the accuracy of calculating similarity. Specifically, for the similarity calculation between the first sample data and the second sample data, the first vector data calculated by the first sample data in the preset intermediate layer of the neural network model can be determined, and the second vector data calculated by the second sample data in the preset intermediate layer can be determined, and then the similarity between the first vector data and the second vector data is calculated.

[0091] For example, such as Figure 2As shown, the neural network includes an input layer, hidden layers (also called hidden layers), and an output layer. Among them, for the number of hidden layers, developers can flexibly select according to the actual situation and experience. Figure 2 Taking the case where the hidden layer contains two layers as an example for illustration. This specification can make full use of the feature encoding ability of the neural network, and use the vector calculated on the hidden layer from the input sample data to replace the sample data for similarity calculation, so as to provide a reason for the sample data that is predicted incorrectly on the validation set. Among them, in the neural network, the features extracted by the hidden layer closer to the output layer are the features that can better reflect the mapping relationship law between the sample data and the label; therefore, the selected preset hidden layer can at least include the layer before the output layer of the neural network model, so as to effectively improve the accuracy of calculating similarity.

[0092] In one case, only the hidden layer located before the output layer can be selected as the preset hidden layer, so as to reduce the amount of calculation data while ensuring the accuracy of similarity calculation. In another case, the vectors of the layer before the output layer and other hidden layers can be selected to participate in the similarity calculation together. Taking Figure 2 as an example, for the sample data in the to-be-detected sample library and the validation set, the vectors on the first layer and the second layer of the hidden layer can be selected simultaneously, and then the selected vectors are concatenated to participate in the similarity calculation together.

[0093] In this embodiment, after calculating the similarity between the first sample data and the second sample data, some sample data can be selected based on the size of the similarity for the detection of abnormal samples, so as to reduce the amount of data for detection. For example, multiple first sample data can be sorted according to the size of the similarity (sorted from large to small according to the similarity), and then at least one first sample data is selected in turn according to the ranking to obtain the detection result for the abnormal samples in the selected first sample data until the obtained detection result meets the preset detection conditions. In other words, similar sample data with relatively high similarity is preferentially detected.

[0094] In order to further reduce the amount of data for detection and detect as many abnormal samples as possible, it can be set that: the number of first sample data selected each time is positively correlated with the number of selection times.

[0095] In this embodiment, the sample data in the sample library to be detected can be divided into N parts, where N is an integer greater than or equal to 2. Based on dividing the sample library to be detected into N parts, the sample data in the sample library to be detected that conflicts with the overall distribution can be detected in the form of cross-validation. Specifically, M of them are used as the training set to train the machine learning model, and the machine learning model is used to predict the N-M parts of the sample data; if the prediction result of the third sample data in the sample library to be detected is inconsistent with the labeled data category of the third sample data, the third sample data is used as an abnormal sample.

[0096] For ease of understanding, the following uses text as the sample data and text category as the label to illustrate in detail the detection scheme for abnormal samples in this specification.

[0097] In the detection scheme for abnormal texts in this specification, the detection of abnormal texts can be divided into three stages. For the first stage and the second stage, it will be described in combination with Figure 3 as follows. Please refer to Figure 3 , Figure 3 which is a flowchart of a method for detecting abnormal texts provided by an exemplary embodiment. As Figure 3 shown, the method may include the following steps:

[0098] Step 302, divide the corpus into N parts.

[0099] In this embodiment, the corpus can be used as the training set for training the neural network model. In this corpus, the text is the question input by the user, and the labeled text category is the knowledge point to which the text belongs, that is, a mapping relationship of "text -> category" is established.

[0100] Taking the scenario of a network operator as an example, the above "text -> category" mapping relationship can be shown in Table 1 as follows:

[0101] User Questions (Sample Data) Text Category (Label) Can domestic remaining data be used to turn on the hotspot? Personal Hotspot Settings Instructions MMS cannot be sent MMS Usage Instructions How to use the phone bill redeemed with points Points Redemption Introduction How to replace a lost SIM card when in a different location Cross-region SIM Card Replacement Introduction What is targeted data? Targeted Data Introduction …… ……

[0102] Table 1

[0103] Step 304, perform cross-validation on the N parts of the corpus obtained by the division.

[0104] Step 306, output the abnormal text.

[0105] In this embodiment, based on dividing the corpus into N parts, the texts in the corpus that conflict with the overall distribution can be detected in the form of cross-validation. Specifically, the corpus is divided into N parts, with any N - 1 parts used as the training set to train a deep neural network model for text classification, and predictions are made on the other 1 part. After N cycles, a prediction result will be obtained for each text in the corpus. If the prediction result of the deep neural network model for a certain text is inconsistent with the label of the text, it indicates that there is a conflict between this text and the overall distribution of the corpus. Then, they can be sorted according to the degree of error (i.e., the prediction score output by the deep neural network model), and then the suspicious samples that may have errors are output according to the ranking for the annotators to detect the suspicious samples and modify the labels of the abnormal samples among them. Thus, it can be seen that in the first stage, most of the outliers in the corpus can be detected.

[0106] Furthermore, in addition to outliers, there may also be non-outliers in the corpus. For example, texts such as "How many points do I have", "Point balance", "I want to check my points" belong to the knowledge point "Point query method"; texts such as "Where can I view the items I redeemed", "Point redemption details", "Was the point redemption successful just now" belong to the knowledge point "Point redemption record query method". Then, for the above multiple texts that should have been labeled as "Point redemption record query method" but were actually labeled as "Point query method", they are systematic error samples, that is, non-outliers.

[0107] For the above non-outliers, the second and third stages are further used for detection.

[0108] Step 308, train a neural network model using the corpus.

[0109] In this embodiment, after the above first stage, the annotators have modified the incorrect labels of the above outliers, thus updating the corpus. Then, in the second stage, the updated corpus can be used to train a deep neural network model to further detect the remaining outliers and some non-outliers.

[0110] Step 310, evaluate the prediction effect on the validation set.

[0111] Step 312, if the prediction effect reaches the best, then go to step 314; otherwise, return to step 308.

[0112] In this embodiment, in order to prevent the deep neural network from overfitting, a validation set can be used to evaluate the generalization ability of the neural network model. Among them, the labels of the texts in the validation set can be considered to be completely correct. For example, the annotator can use the texts modified in the above first stage as part of the validation set, and can also randomly select a certain number of texts from the online logs of the online customer service system for annotation as the sample data in the validation set.

[0113] Based on the configured validation set, the model parameters corresponding to the case with the minimum error (i.e., the best prediction effect) during the iterative process can be used as the model parameters of the deep neural network model. For example, during the iterative training process, in each iteration cycle, the current model is used to evaluate the prediction effect on the validation set, and then the model parameters with the best prediction effect on the validation set in each iteration cycle are selected as the final model parameters of the neural network model.

[0114] Alternatively, the model parameters corresponding to the case where the error does not exceed the preset error threshold can be used as the model parameters of the deep neural network model. For example, due to the strong fitting ability of the neural network, the prediction effect on the validation set will first increase and then gradually decrease during the iterative process. Therefore, the error of the current model on the validation set can be calculated in each iteration cycle until the error of the current model on the validation set is larger than the error of the previous iteration cycle, at which point the training is stopped, and the model parameters in the previous iteration cycle are used as the final model parameters of the neural network model.

[0115] Step 314, make predictions on the corpus.

[0116] Step 316, output abnormal texts.

[0117] Similarly, the predictions can be sorted according to the degree of prediction error (i.e., the prediction scores output by the deep neural network model), and then the suspicious samples with possible labeling errors can be output according to the ranking for the annotator to detect the suspicious samples and modify the labels of the abnormal samples among them.

[0118] In the above second stage, the characteristic that the neural network "usually fits simple samples first and then difficult samples" is utilized. By using the validation set to monitor the iterative process, the fitting process can be stopped at an appropriate position. At this time, the samples that cannot be predicted correctly in the training set are difficult-to-fit samples. It should be noted that the difficult-to-fit samples contain both random errors and systematic errors. Since not all samples with systematic errors have a high fitting difficulty, further detection is required in the third stage.

[0119] Please refer to Figure 4 , Figure 4is a flowchart of another method for detecting abnormal text provided by an exemplary embodiment. As Figure 4 shown, the method may include the following steps:

[0120] Step 402, training a neural network model using a corpus.

[0121] After the above second stage, the annotator has modified the misannotations in the abnormal samples output above, thereby updating the corpus. Then, in the third stage, the updated corpus can be used to train a deep neural network model to further detect the remaining non-outliers. Figure 3

[0122] Step 404, evaluating the prediction effect on the validation set.

[0123] Step 406, if the prediction effect reaches the best, go to step 408; otherwise, return to step 402.

[0124] The implementation process of the above steps 402-406 is similar to the above steps 308-312 and will not be elaborated here.

[0125] Step 408, extracting the vector calculated by the neural network model on the intermediate layer to represent the text.

[0126] Step 410, determining the target text.

[0127] Step 412, determining the similarity with similar texts.

[0128] Since a deep neural network is used for training, the text feature encoding ability of the neural network can be fully utilized, and the vector calculated by the intermediate layer (hidden layer) of the deep neural network is used to represent the corresponding sample data, so as to participate in the similarity calculation instead of the sample data, which can effectively improve the accuracy of calculating the similarity. For the specific process of this similarity calculation, reference can be made to the description in the above Figure 2 part and will not be elaborated here.

[0129] By the above method of matching similar texts using the vector of the network intermediate layer, error reasons can be provided for the prediction error samples in the validation set. If the error reason is misannotation, the neighboring text in the matched corpus is used as an abnormal sample to further check whether there is misannotation in the labels of these abnormal samples.

[0130] Step 414, outputting the neighboring text.

[0131] Step 416, if the detection condition is met, go to step 418; otherwise, return to step 414.

[0132] ​In this embodiment, the similarity between the similar sample data for which the tag to be detected is the prediction result for any target sample data and the any target sample data can be calculated, and the similar sample data can be selected based on the calculated similarity, and the detection result for the abnormal samples in the selected similar sample data can be obtained.

[0133] For example, for a piece of text predicted incorrectly in the validation set, the predicted category is c. Calculate the cosine similarity between its vector and the vectors of each piece of text in the training set (i.e., the text library to be detected) labeled with the category c, and find the nearest neighbor sample with the closest similarity. First, a small number of nearest neighbor samples (e.g., 5) can be output, and the annotator can detect whether there are mislabeled texts among them. When there are no mislabeled texts among them, it indicates that the prediction error for the text in the validation set is not caused by the mislabeling of the text library to be detected. Then, the search for similar sample data for the text predicted incorrectly in the validation set can be stopped, and the above operations can be further performed for other texts predicted incorrectly. When there are mislabeled texts among them, it can be determined that the prediction error is caused by the mislabeling of the text library to be detected, and then the annotator can modify the mislabeling of these texts. And at this time, the number of nearest neighbor samples can be continuously increased (e.g., 20); similarly, the annotator can detect whether there are mislabeled texts among them. In the process of repeatedly executing the above operation of outputting nearest neighbor samples for detection, it can be set to stop the loop until the detection result meets the preset detection conditions. For example, the proportion of the number of mislabeled nearest neighbor samples in the total number of nearest neighbor samples output this time is lower than the preset threshold. Of course, the preset detection conditions within each loop period can be set to be the same or different, and this specification does not limit this.

[0134] Step 418, if all target texts have been traversed, go to step 420; otherwise, return to step 410.

[0135] Step 420, end the third stage.

[0136] In the third stage, each piece of text in the validation set is used as a probe in a fine-grained manner to detect the remaining systematic mislabeling (non-outliers).

[0137] It should be noted that, in addition to being executed in the above order, the above three stages can also be executed independently of each other, and this specification does not limit this.

[0138] Corresponding to the above method embodiment, this specification also provides an embodiment of a detection device for abnormal samples.

[0139] Figure 5 It is a schematic structural diagram of a device provided by an exemplary embodiment. Please refer to Figure 5, at the hardware level, the device includes a processor 502, an internal bus 504, a network interface 506, a memory 508, and a non-volatile memory 510. Of course, it may also include other hardware required for other services. The processor 502 reads the corresponding computer program from the non-volatile memory 510 into the memory 508 and then runs it, forming a detection device for abnormal samples at the logical level. Of course, in addition to the software implementation method, one or more embodiments of this specification do not exclude other implementation methods, such as logic devices or a combination of software and hardware, etc. That is to say, the execution subject of the following processing flow is not limited to each logical unit, and can also be hardware or a logic device.

[0140] Please refer to Figure 6 , in a software implementation, the detection device for abnormal samples may include:

[0141] A first acquisition unit 61, which acquires a sample library to be detected and a validation set corresponding to the sample library to be detected. The sample data in the sample library to be detected is labeled with a data category, and the sample data in the validation set is labeled with a standard category;

[0142] A first training unit 62, which trains a machine learning model using the sample library to be detected and uses the machine learning model to predict the sample data in the validation set;

[0143] A first detection unit 63, which determines the target sample data in the validation set whose prediction result is inconsistent with the standard category as an abnormal sample.

[0144] Optionally, it further includes:

[0145] A second acquisition unit 64, which obtains first sample data with a prediction result of a target category in the sample library to be detected, where the target category is the prediction result of second sample data in the target sample data;

[0146] A calculation unit 65, which calculates the similarity between the first sample data and the second sample data;

[0147] A second detection unit 66, which determines whether the first sample data is an abnormal sample according to the similarity.

[0148] Optionally, it further includes:

[0149] A division unit 67, which divides the sample data in the sample library to be detected into N parts, where N is an integer greater than or equal to 2;

[0150] A second training unit 68, which uses M of them as a training set to train the machine learning model and uses the machine learning model to predict the sample data of N - M parts;

[0151] The third detection unit 69, if the prediction result of the third sample data in the to-be-detected sample library is inconsistent with the labeled data category of the third sample data, uses the third sample data as an abnormal sample.

[0152] Optionally, the first training unit 62 is specifically configured to:

[0153] Train the to-be-detected sample library using a neural network algorithm to obtain a neural network model;

[0154] Among them, the model parameters of the neural network model are the model parameters when the error between the prediction result of the sample data in the validation set and the corresponding standard category during the iteration process meets the preset error condition.

[0155] Optionally, the first training unit 62 is further configured to:

[0156] Determine the first vector data calculated by the first sample data in the preset intermediate layer of the neural network model;

[0157] Determine the second vector data calculated by the second sample data in the preset intermediate layer;

[0158] Calculate the similarity between the first vector data and the second vector data.

[0159] Optionally, the preset intermediate layer includes at least the layer before the output layer of the neural network model.

[0160] Optionally, the first detection unit 63 is specifically configured to:

[0161] Sort multiple first sample data according to the magnitude of the similarity;

[0162] Select at least one first sample data in sequence according to the ranking to obtain the detection result of abnormal samples among the selected first sample data until the obtained detection result meets the preset detection condition.

[0163] Optionally, the number of first sample data selected each time is positively correlated with the number of selection times.

[0164] The systems, devices, modules, or units illustrated in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer, and the specific form of the computer can be a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email transceiver device, a game console, a tablet computer, a wearable device, or a combination of any several of these devices.

[0165] In a typical configuration, a computer includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0166] The memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.

[0167] Computer-readable media include permanent and non-permanent, removable and non-removable media that can be used to store information by any method or technology. Information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, disk storage, quantum memory, graphene-based storage media or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include temporary computer-readable media (transitory media), such as modulated data signals and carrier waves.

[0168] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.

[0169] The above is a description of a specific embodiment of the specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0170] The terms used in one or more embodiments of this specification are for the purpose of describing particular embodiments only and are not intended to limit one or more embodiments of this specification. The singular forms "a", "the", and "said" used in one or more embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly dictates otherwise. It should also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items.

[0171] It should be understood that although the terms first, second, third, etc. may be used in one or more embodiments of this specification to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of one or more embodiments of this specification, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to determining".

[0172] The above description is only the preferred embodiment of one or more embodiments of this specification and is not intended to limit one or more embodiments of this specification. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of one or more embodiments of this specification shall be included within the scope of protection of one or more embodiments of this specification.

Claims

1. A detection method for abnormal samples, characterized in that, Including: Obtain a sample library to be detected and a validation set corresponding to the sample library to be detected. The sample data in the sample library to be detected is labeled with a data category, and the sample data in the validation set is labeled with a standard category; Use the sample library to be detected to train a machine learning model, and use the machine learning model to predict the sample data in the validation set; Determine the target sample data in the validation set whose prediction result is inconsistent with the standard category as an abnormal sample; Also including: Obtain first sample data with a prediction result of a target category in the sample library to be detected, where the target category is the prediction result of second sample data in the target sample data; Calculate the similarity between the first sample data and the second sample data; Determine whether the first sample data is an abnormal sample according to the similarity; The determining whether the first sample data is an abnormal sample according to the similarity includes: Sort multiple first sample data according to the magnitude of the similarity; Select at least one first sample data in sequence according to the ranking to obtain a detection result for the abnormal sample among the selected first sample data until the obtained detection result meets a preset detection condition.

2. The method according to claim 1, wherein Also including: Divide the sample data in the sample library to be detected into N parts, where N is an integer greater than or equal to 2; Use M of them as a training set to train the machine learning model, and use the machine learning model to predict the sample data of N - M parts; If the prediction result of the third sample data in the sample library to be detected is inconsistent with the data category labeled for the third sample data, regard the third sample data as an abnormal sample.

3. The method according to claim 1, wherein The using the sample library to be detected to train a machine learning model includes: Adopt a neural network algorithm to train the sample library to be detected to obtain a neural network model; Among them, use the model parameters when the error between the prediction result of the sample data in the validation set and the corresponding standard category in the iterative process meets a preset error condition as the model parameters of the neural network model.

4. The method according to claim 3, wherein The calculating the similarity between the first sample data and the second sample data includes: Determine the first vector data calculated by the first sample data in a preset intermediate layer of the neural network model; Determine the second vector data calculated by the second sample data in the preset intermediate layer; Calculate the similarity between the first vector data and the second vector data.

5. The method according to claim 4, characterized in that, The preset intermediate layer includes at least the layer before the output layer of the neural network model.

6. The method according to claim 1, characterized in that, The number of first sample data selected each time is positively correlated with the number of selections.

7. A detection device for abnormal samples, characterized in that, Including: A first obtaining unit that obtains a sample library to be detected and a validation set corresponding to the sample library to be detected. The sample data in the sample library to be detected is labeled with a data category, and the sample data in the validation set is labeled with a standard category; A first training unit that uses the sample library to be detected to train a machine learning model and uses the machine learning model to predict the sample data in the validation set; A first detecting unit that determines the target sample data in the validation set whose prediction result is inconsistent with the standard category as an abnormal sample; A second acquisition unit that acquires first sample data with a predicted result of a target category in the sample library to be detected, where the target category is the predicted result of second sample data in the target sample data; A calculation unit that calculates the similarity between the first sample data and the second sample data; A second detection unit that determines whether the first sample data is an abnormal sample according to the similarity; The first detection unit is specifically configured to: Sort multiple first sample data according to the magnitude of the similarity; Select at least one first sample data in sequence according to the ranking to obtain a detection result for abnormal samples among the selected first sample data until the obtained detection result meets a preset detection condition.

8. The device according to claim 7, characterized in that, It further includes: A division unit that divides the sample data in the sample library to be detected into N parts, where N is an integer greater than or equal to 2; A second training unit that uses M of them as a training set to train the machine learning model and uses the machine learning model to predict the sample data of N - M parts; A third detection unit that, if the predicted result of the third sample data in the sample library to be detected is inconsistent with the labeled data category of the third sample data, regards the third sample data as an abnormal sample.

9. The device according to claim 7, characterized in that The first training unit is specifically configured to: Train the sample library to be detected using a neural network algorithm to obtain a neural network model; Among them, the model parameters of the neural network model are the model parameters when the error between the predicted result of the sample data in the validation set and the corresponding standard category during the iteration process meets a preset error condition.

10. The device according to claim 9, characterized in that, The first training unit is further configured to: Determine the first vector data calculated by the first sample data in a preset intermediate layer of the neural network model; Determine the second vector data calculated by the second sample data in the preset intermediate layer; Calculate the similarity between the first vector data and the second vector data.

11. The device according to claim 10, wherein The preset intermediate layer includes at least the layer before the output layer of the neural network model.

12. The device according to claim 7, characterized in that The number of first sample data selected each time is positively correlated with the number of selections.

13. An electronic device, characterized in that, It includes: A processor; A memory for storing instructions executable by the processor; Among them, the processor realizes the method according to any one of claims 1 - 6 by running the executable instructions.

14. A computer-readable storage medium, characterized in that, A computer instruction is stored thereon, and when the instruction is executed by the processor, the steps of the method according to any one of claims 1 - 6 are realized.

Citation Information

Patent Citations

  • Face recognition method and face recognition device

    CN104899579A

  • Scene investigation image retrieval method based on low-level image features and CNN features

    CN108985346A

  • Optimization method and device of training set for text classification

    CN110580290A