A method, device, medium and equipment for detecting out-of-distribution samples in text classification
By using the default training data set of categories to train the text classification model and using the trained sub-model to detect external distributed samples, the problem of error detection in the existing technology of Chinese and foreign distributed samples is solved, and more accurate text classification prediction is achieved.
Patent Information
- Application Number
- CN202111211129.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-18
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2041-10-18
AI Technical Summary
The prior art cannot obtain correct prediction results when facing out-of-distribution of samples, resulting in incorrect classification prediction.
The text classification model is trained through the default training data set of categories, and multiple out-distributed samples test sub-models corresponding to each text category are obtained, and these sub-models are used to test the new input text to determine whether it is an out-distributed sample.
It effectively avoids incorrect classification predictions for externally distributed samples. It is better not to make predictions than to make extremely likely mistakes.
Smart Images

Figure CN114020905B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of text editing, and particularly to a method, device, medium, and equipment for detecting out-of-distribution samples in text classification. Background Art
[0002] Text classification is to give an input sentence, and the model needs to determine what its classification is. An out-of-distribution sample is that the input sentence and the data used by the model during training are inconsistent in data distribution. That is, the input sentence does not belong to the categories included in the data for training the model. For example, the data used to train the text classification model are all in the categories of "society" and "sports", but there is also the category of "politics" during testing. That is, the model has never seen data in the "politics" category during training, and for the model, the data in the "politics" category is an out-of-distribution sample. The prior art cannot obtain correct prediction results when facing out-of-distribution samples. Summary of the Invention
[0003] Aiming at the problems existing in the prior art, this application mainly provides a method, device, medium, and equipment for detecting out-of-distribution samples in text classification. By training a text classification model with a category-default training data set and using the trained model to detect and judge the current input, it can avoid making wrong classification predictions for the current input.
[0004] To achieve the above object, a technical solution adopted by this application is: to provide a method for detecting out-of-distribution samples in text classification, which includes:
[0005] Successively using a category-default training data set that lacks one of the text categories in the training data set containing multiple text categories to train the text classification model, obtaining multiple out-of-distribution sample test sub-models corresponding to each text category; and using each out-of-distribution sample test sub-model to test the new input text, and judging whether the new input text is an out-of-distribution sample according to each test result.
[0006] Another technical solution adopted by this application is: to provide a device for detecting out-of-distribution samples in text classification, which includes:
[0007] A model training module, which is used to successively use a category-default training data set that lacks one of the text categories in the training data set containing multiple text categories to train the text classification model, obtaining multiple out-of-distribution sample test sub-models corresponding to each text category; and a new text test module, which is used to use each out-of-distribution sample test sub-model to test the new input text, and judging whether the new input text is an out-of-distribution sample according to each test result.
[0008] Another technical solution adopted in this application is: a computer-readable storage medium storing computer instructions, characterized in that the computer instructions are operated to execute the out-of-distribution sample detection method for text classification in the above solution.
[0009] Another technical solution adopted in this application is: a computer device including a processor and a memory, the memory storing computer instructions, and the computer instructions are operated to execute the out-of-distribution sample detection method for text classification in the above solution.
[0010] The beneficial effects that can be achieved by the technical solution of this application are: This application designs an out-of-distribution sample detection method, device, medium and equipment for text classification. The method uses class-default training data to train a text classification model, and uses the trained model to detect the current input. If the model does not have enough confidence to make a correct judgment on it, it is judged as an out-of-distribution sample, thus avoiding making a wrong classification prediction when the current input is an out-of-distribution sample. Description of the Drawings
[0011] In order to more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of this application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0012] Figure 1 It is a flowchart showing a specific implementation of an out-of-distribution sample detection method for text classification in this application;
[0013] Figure 2 It is a flowchart showing the process of training a text classification model in a specific embodiment of an out-of-distribution sample detection method for text classification in this application;
[0014] Figure 3 It is a flowchart showing the process of testing a new input text in a specific embodiment of an out-of-distribution sample detection method for text classification in this application;
[0015] Figure 4 It is a schematic diagram showing a specific implementation of an out-of-distribution sample detection device for text classification in this application;
[0016] Figure 5 It is a schematic diagram showing a specific embodiment of an out-of-distribution sample detection device for text classification in this application.
[0017] Through the above-mentioned drawings, specific embodiments of the present disclosure have been shown, and will be described in more detail hereinafter. These drawings and the written description are not intended to limit the scope of the concept of the present disclosure in any way, but to illustrate the concept of the present disclosure to those skilled in the art by reference to specific embodiments. Detailed Description of the Embodiments
[0018] The following elaborates on the preferred embodiments of the present application in conjunction with the accompanying drawings, so that the advantages and features of the present application can be more easily understood by those skilled in the art, thereby making the scope of protection of the present application more clearly defined.
[0019] It should be noted that, in this text, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variation thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, elements defined by the statement "comprising..." do not exclude the presence of additional identical elements in the process, method, article or device comprising the said elements.
[0020] Text classification is to give an input sentence, and the model needs to determine what its category is, such as "politics", "economy", "society", "sports", etc. The so-called out-of-distribution means that the sentence input during testing is inconsistent with the data used to train the model in terms of data distribution. For example, the data used to train the text classification model are all in the categories of "society" and "sports", but there is also the category of "politics" during testing. For such a category, the model will never be able to obtain the correct prediction result because the model has never seen data in the "politics" category during training. Out-of-distribution detection is that the model needs to determine during testing whether the current input text is an out-of-distribution category. If so, find it out; if not, make a category prediction for it.
[0021] This task is actually a problem of "model confidence". For the current input, does the model have enough confidence to make a correct judgment on it? When the model is not confident enough, this sample may be an out-of-distribution sample. In other words, the significance of out-of-distribution detection is: it is better not to make a prediction than to make a prediction that is very likely to be wrong.
[0022] The present application innovatively proposes a method based on k-fold model integration. Using training data including k text categories, first train k sub-models. Each sub-model treats one of the k categories as the "out-of-distribution" category. After training the k sub-models, during testing, use the k sub-models to test the newly input text separately, and integrate the respective test results to form a final predicted probability distribution. If the probability of one class is significantly higher than that of other classes in this distribution, it is considered that the model is very confident in predicting the category of the current input sample, which also means that this sample is not an out-of-distribution sample; if the probabilities of all classes are similar, it is considered that the model is not confident in predicting this sample, so it is an out-of-distribution data. With this method, out-of-distribution samples can be well detected.
[0023] The following will use specific embodiments and in combination with the accompanying drawings to elaborate in detail on the technical solutions of the present application. These several specific embodiments can be combined with each other, and for the same or similar concepts or processes, they may not be repeated in some embodiments.
[0024] Figure 1 It shows a specific implementation manner of a method for detecting out-of-distribution samples in text classification of the present application.
[0025] In this specific implementation manner, the method for detecting out-of-distribution samples in text classification of the present application mainly includes a model training process S101, which successively uses the category-default training data sets lacking one of the text categories in the training data set including multiple text categories to train the text classification model, and obtains multiple out-of-distribution sample test sub-models corresponding to each text category; and a new text testing process S102, which uses each out-of-distribution sample test sub-model to test the newly input text, and determines whether the newly input text is an out-of-distribution sample according to each test result.
[0026] By using the category-default training data to train the text classification model and using the trained model to detect the current input, if the model does not have enough confidence to make a correct judgment on it, it is judged as an out-of-distribution sample, so as to avoid making a wrong classification prediction for the current input that is an out-of-distribution sample. It is better not to make a prediction than to make a prediction that is very likely to be wrong.
[0027] The model training process S101 represents the process of successively using the category-default training data sets lacking one of the text categories in the training data set including multiple text categories to train the text classification model, and obtaining multiple out-of-distribution sample test sub-models corresponding to each text category, which is beneficial to subsequent use of each out-of-distribution sample test sub-model to test the newly input text.
[0028] In a specific embodiment of the present application, the process of training the text classification model using the class-default training dataset includes performing uncertainty training on the text classification model, such that the out-of-distribution sample test sub-model obtained by training the text classification model predicts a text class missing from the corresponding class-default training dataset, and the probabilities of each class in the class-default training dataset tend to be evenly distributed.
[0029] In a specific embodiment of the present application, the process of training the text classification model using the class-default training dataset includes making the out-of-distribution sample test sub-model obtained by training the text model predict a text class missing from the corresponding class-default training dataset, and the probabilities of each class in the class-default training dataset all tend to 0.
[0030] In a specific embodiment of the present application, the process of training the text classification model using the class-default training dataset includes training the text classification model using the class-default training dataset with a conventional training method.
[0031] Preferably, the text classification model is trained using the cross-entropy loss function with the class-default training dataset.
[0032] In a specific example of the present application, the text classification model is trained using the cross-entropy loss function with the class-default training dataset, and uncertainty training is performed on the text classification model, such that the out-of-distribution sample test sub-model obtained by training the text classification model predicts a text class missing from the corresponding class-default training dataset, and the probabilities of each class in the class-default training dataset tend to be evenly distributed, thereby obtaining an out-of-distribution sample test sub-model corresponding to each text class.
[0033] In a specific example of the present application, as Figure 2 shown, the above training data total set includes data of K classes. Each time, 1 class is drawn from the above training data total set, and the model is trained on the data of the remaining K - 1 classes according to the normal cross-entropy loss function. When the model encounters the drawn class, the model treats the data of this class as an out-of-distribution sample for training, and makes the model's prediction result tend to be "evenly distributed", which is expressed by the formula
[0034] min KL(f(x), u)
[0035] Among them, KL is the Kullback-Leibler divergence, which is used to measure the similarity between two probability distributions. f(x) is the probability distribution predicted by the model, and u is the uniform distribution. The meaning of this formula is to make the model prediction distribution approximate the uniform distribution, so as to indicate the uncertainty of the model about the external distribution samples.
[0036] The probability of predicting the drawn category as one of the known K1 categories is 1 / (K - 1), which indicates that the model encounters an out-of-distribution sample and is extremely unsure about the prediction result of this sample. For example, when K = 5, the model should predict the probability of each visible category as 1 / 4 when encountering an out-of-distribution sample. The above process is for one category. For each of the K categories, a separate model is trained according to the above idea. Each model focuses on a specific out-of-distribution category, so as to judge whether a newly input sample is an out-of-distribution sample based on the prediction results of each model during subsequent testing. Since the prediction results of K models need to be combined, the method proposed by this scheme is also called k-fold model ensemble.
[0037] In a specific example of this application, the above training data set includes three categories A, B, and C. First, category A is taken out and made invisible to model 1. In other words, the model only knows about categories B and C now and does not know the existence of category A. At this time, category A is the "out-of-distribution data". Model 1 is trained on the data of categories B and C according to the conventional method, and can then accurately predict the samples of categories B and C. On the data of category A, the model prediction is made to tend to the uniform distribution (0.5, 0.5), indicating that the probability of the current sample being category B and the probability of being category C are both 0.5, and further indicating that the model considers category A to be out-of-distribution. Similarly, category B is taken out, and model 2 is conventionally trained on categories A and C, and the prediction of model 2 is made to tend to the uniform distribution (0.5, 0.5) on category B. The same goes for model 3.
[0038] In the new text testing process S102, each out-of-distribution sample testing sub-model is used to test the newly input text, and based on each test result, it is judged whether the newly input text is an out-of-distribution sample. The already trained sub-model can be used to test the newly input text. If the newly input text is an out-of-distribution sample, then it is not predicted, that is, it is not classified. It is better not to make a prediction than to make a possibly wrong prediction.
[0039] In a specific embodiment of the present application, the process of using each out-of-distribution sample testing sub-model to test the new input text and determining whether the new input text is an out-of-distribution sample according to each test result includes using each out-of-distribution sample testing sub-model to test the new input text respectively to obtain the corresponding predicted probability distribution results of each model, integrating the predicted probability distribution results corresponding to each model to obtain a final predicted probability distribution, and then judging whether the new input text is an out-of-distribution sample based on the final predicted probability distribution.
[0040] In a specific embodiment of the present application, the process of integrating the predicted probability distribution results corresponding to each model to obtain a final predicted probability distribution includes judging whether the new input text is an out-of-distribution sample according to the average value of each test result. That is, in the predicted probability distribution results corresponding to each model, the probability values of each category are averaged.
[0041] In a specific embodiment of the present application, the process of integrating the predicted probability distribution results corresponding to each model to obtain a final predicted probability distribution includes adding up the probability values of each category in the predicted probability distribution results corresponding to each model to obtain a final predicted probability distribution.
[0042] In a specific embodiment of the present application, the final probability distribution is obtained according to the average value of each test result.
[0043] In a specific embodiment of the present application, if the entropy of the final probability distribution is greater than the preset probability distribution entropy threshold, then the new input text is determined to be an out-of-distribution sample.
[0044] In a specific embodiment of the present application, if the entropy of the final probability distribution is not greater than the preset probability distribution entropy threshold, then the new input text is determined to be an in-distribution sample.
[0045] In a specific example of the present application, as Figure 3 shown, after training K sub-models, they can be used in the test stage. For the input test sample, first send it into the K sub-models respectively to obtain their respective probability distributions, and then integrate their prediction results to obtain the final probability distribution. If the probability of one category in this distribution is significantly higher than other categories, it indicates that the currently input test sample should belong to this category and is an "in-distribution" sample; if the probabilities of all categories are close and tend to be evenly distributed, it indicates that the currently input test sample is an "out-of-distribution" sample. This can be measured by the "entropy" of the probability distribution:
[0046]
[0047] The greater the entropy, the greater the uncertainty of the probability distribution, and the greater the likelihood that the model predicts the current sample as an "out-of-distribution" sample. Therefore, in this solution, a threshold a is set. When the entropy H(P) of the final probability distribution is greater than a, the sample is considered an "out-of-distribution" sample; otherwise, it is an in-distribution sample. This finally realizes the detection of out-of-distribution samples.
[0048] In a specific example of this application, a text classification model is trained using a total data set including 3 categories A, B, and C to obtain 3 corresponding out-of-distribution sample test sub-models. When a new text is input, the 3 models make predictions respectively, and their prediction results are integrated to obtain a probability distribution for categories A, B, and C. For example, it is (0.1, 0.1, 0.8). The probability of category C is 0.8, which indicates that this sample belongs to category C and is not an out-of-distribution sample. However, if the probability distribution is (0.3, 0.3, 0.4), it indicates that this sample is very likely to be an out-of-distribution sample because this probability distribution is very close to the uniform distribution (0.33, 0.33, 0.33). Therefore, by judging whether the probability distribution obtained by integrating the final model is close to the uniform distribution, we can determine whether the newly input text is an out-of-distribution sample.
[0049] In a specific example of this application, if the newly input text is not an out-of-distribution sample, then normal prediction classification is performed on this sample.
[0050] Figure 4 A specific implementation manner of a text classification out-of-distribution sample detection device of this application is shown.
[0051] In this specific implementation manner, the text classification out-of-distribution sample detection device of this application mainly includes a model training module 401, which is used to sequentially train a text classification model using a category-default training data set lacking one of the text categories in a training data set including multiple text categories, to obtain multiple out-of-distribution sample test sub-models corresponding to each text category; and a new text test module 402, which is used to test the newly input text using each out-of-distribution sample test sub-model and determine whether the newly input text is an out-of-distribution sample according to each test result.
[0052] By training a text classification model using category-default training data and using the trained model to detect the current input, if the model does not have enough confidence to make a correct judgment on it, it is judged as an out-of-distribution sample, thus avoiding making a wrong classification prediction for the current input that is an out-of-distribution sample. It is better not to make a prediction than to make a very likely wrong prediction.
[0053] The model training module 401 is used to train the text classification model in turn by using the category - default training data sets in the training data set containing multiple text categories, each lacking one of the text categories, to obtain multiple out - of - distribution sample test sub - models corresponding to each text category, which is beneficial for subsequent testing of newly input texts using each out - of - distribution sample test sub - model.
[0054] In a specific embodiment of the present application, the above - mentioned model training module 401 includes an uncertainty training sub - module as Figure 5 shown, which is used to perform uncertainty training on the text classification model, so that the out - of - distribution sample test sub - model formed by training the text classification model predicts one text category missing in the corresponding category - default training data set, and the probabilities of each category in the category - default training data set tend to be evenly distributed.
[0055] In a specific embodiment of the present application, the above - mentioned uncertainty training sub - module can make the out - of - distribution sample test sub - model formed by training the text model predict one text category missing in the corresponding category - default training data set, and the probabilities of each category in the category - default training data set all tend to 0.
[0056] In a specific embodiment of the present application, the above - mentioned model training module 401 includes a conventional training sub - module as Figure 5 shown, which can use the category - default training data set to train the text classification model using conventional training methods.
[0057] In a specific embodiment of the present application, the above - mentioned conventional training sub - module can use the category - default training data set to train the text classification model using the cross - entropy loss function.
[0058] The above - mentioned new text testing module 402 can use the trained sub - models to test newly input texts. If the newly input text is an out - of - distribution sample, then it will not be predicted, that is, it will not be classified. It is better not to make a prediction than to make a possibly wrong prediction.
[0059] In a specific embodiment of the present application, the above - mentioned new text testing module 402 uses each out - of - distribution sample test sub - model to test the newly input text respectively, and obtains the corresponding prediction probability distribution results of each model.
[0060] In a specific embodiment of the present application, the above - mentioned new text testing module 402 includes a result integration sub - module as Figure 5 shown, which is used to obtain the average value of each test result and judge whether the newly input text is an out - of - distribution sample according to the average value.
[0061] In a specific embodiment of the present application, the above result integration sub-module can superimpose the probability values of each category in the corresponding prediction probability distribution result of each model to obtain a final prediction probability distribution.
[0062] In a specific embodiment of the present application, the above result integration sub-module can obtain the final probability distribution according to the average value of each test result.
[0063] In a specific example of the present application, the above result integration sub-module can, and when the entropy of the final probability distribution is greater than a preset probability distribution entropy threshold, determine the newly input text as an out-of-distribution sample.
[0064] The text classification out-of-distribution sample detection device provided by the present application can be used to execute the text classification out-of-distribution sample detection method described in any of the above embodiments, and its implementation principle and technical effects are similar, which will not be elaborated here.
[0065] In a specific embodiment of the present application, each functional module in a text classification out-of-distribution sample detection device of the present application can be directly in hardware, in a software module executed by a processor, or in a combination of the two.
[0066] The software module can reside in a RAM memory, a flash memory, a ROM memory, an EPROM memory, an EEPROM memory, a register, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art. The exemplary storage medium is coupled to the processor such that the processor can read information from and write information to the storage medium.
[0067] The processor can be a Central Processing Unit (CPU), or it can be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic, discrete hardware components, or any combination thereof. The general-purpose processor can be a microprocessor, but in an alternative, the processor can be any conventional processor, controller, microcontroller, or state machine. The processor can also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors in combination with a DSP core, or any other such configuration. In an alternative, the storage medium can be integrated with the processor. The processor and the storage medium can reside in an ASIC. The ASIC can reside in a user terminal. In an alternative, the processor and the storage medium can reside in the user terminal as discrete components.
[0068] In another specific embodiment of the present application, a computer-readable storage medium stores computer instructions, and the computer instructions are operated to detect out-of-distribution samples for text classification in the above-mentioned solution.
[0069] In another specific embodiment of the present application, a computer device includes a processor and a memory, and the memory stores computer instructions, and the computer instructions are operated to execute the method for detecting out-of-distribution samples for text classification in the above-mentioned solution.
[0070] In several embodiments provided by the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.
[0071] The unit described as a separation component may or may not be physically separated. The component shown as a unit may or may not be a physical unit, that is, it may be located in one place or may be distributed over multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0072] The above are only the embodiments of the present application, and do not limit the patent scope of the present application accordingly. Any equivalent structural transformation made by using the content of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall be included in the patent protection scope of the present application by the same token.
Claims
1. A method for detecting out-of-distribution samples in text classification, characterized in that, it includes, successively using a class-default training data set lacking one of the text classes in a total training data set including multiple text classes to train a text classification model, and obtaining multiple out-of-distribution sample test sub-models corresponding to each of the text classes; and, using each of the out-of-distribution sample test sub-models to test a newly input text, and judging whether the newly input text is an out-of-distribution sample according to each test result; wherein, judging whether the newly input text is an out-of-distribution sample according to the average value of each test result, and obtaining a final probability distribution according to the average value of each test result. If the entropy of the final probability distribution is greater than a preset probability distribution entropy threshold, then the newly input text is determined to be an out-of-distribution sample.
2. The method for detecting out-of-distribution samples in text classification according to claim 1, characterized in that, the process of successively using a class-default training data set lacking one of the text classes in a total training data set including multiple text classes to train the text classification model includes, making each of the out-of-distribution sample test sub-models predict one of the text classes missing in the corresponding class-default training data set, so that the probabilities of each class in the class-default training data set tend to be evenly distributed.
3. The method for detecting out-of-distribution samples in text classification according to claim 1 or 2, characterized in that, the process of successively using a class-default training data set lacking one of the text classes in a total training data set including multiple text classes to train the text classification model includes, using the class-default training data set to train the text classification model with a cross-entropy loss function.
4. A device for detecting out-of-distribution samples in text classification, characterized in that, it includes, a model training module, configured to successively use a class-default training data set lacking one of the text classes in a total training data set including multiple text classes to train a text classification model, and obtain multiple out-of-distribution sample test sub-models corresponding to each of the text classes; and, a new text test module, configured to use each of the out-of-distribution sample test sub-models to test a newly input text, and judge whether the newly input text is an out-of-distribution sample according to each test result; wherein, judging whether the newly input text is an out-of-distribution sample according to the average value of each test result, and obtaining a final probability distribution according to the average value of each test result. If the entropy of the final probability distribution is greater than a preset probability distribution entropy threshold, then the newly input text is determined to be an out-of-distribution sample.
5. The device for detecting out-of-distribution samples in text classification according to claim 4, characterized in that, the model training module includes an uncertainty training sub-module and a conventional training sub-module; The uncertain training sub-module is configured to make each of the outlier sample test sub-models predict one of the text categories missing in the corresponding category-default training dataset, so that the probabilities of each category in the category-default training dataset tend to be evenly distributed; The conventional training sub-module is configured to train the text classification model by using the category-default training dataset and the cross-entropy loss function.
6. A computer-readable storage medium storing computer instructions, characterized in that, the computer instructions are operated to execute the outlier sample detection method for text classification according to any one of claims 1-3.
7. A computer device comprising a processor and a memory, the memory storing computer instructions, wherein the processor operates the computer instructions to execute the outlier sample detection method for text classification according to any one of claims 1-3.
Citation Information
Patent Citations
Data screening method and device
CN109599096A
Text automatic classification method and device and storage medium
CN109739985A