Training Method of Text Classification Model, Text Processing Method, Device and Medium

Through the mutually heterogeneous text classification model, the noise sample filtering and parameter optimization are solved, and the problem of noise samples in deep learning text classification is achieved, achieving more efficient training and reducing costs.

CN115130538BActive Publication Date: 2025-07-29TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210417059.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-20
Publication Date
2025-07-29
Estimated Expiration
2042-04-20

AI Technical Summary

Technical Problem

Traditional deep learning text classification methods are prone to memory errors when processing noise samples, resulting in increased training time and cost. Especially when there are many categories, the learning of the noise transfer matrix is difficult.

Method used

By using the first and second text classification models that are heterogeneous, the text category probability of the sample sentence is output separately, noise sample filtering is performed, clean sample sets are obtained, and model parameters are optimized based on the clean sample sets to achieve collaborative learning.

Benefits of technology

It reduces training time and cost, avoids the model from falling into self-enclosure during training, reduces the error accumulation of noise samples during screening, and improves the classification ability on clean data sets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115130538B_ABST
    Figure CN115130538B_ABST
Patent Text Reader

Abstract

Embodiments of the present application disclose a method for training a text classification model, a method for text processing, a device, and a medium, which are used to reduce mislabeled samples. The method includes: obtaining a plurality of batch sample sets corresponding to a target scenario, inputting a sample sentence into a first text classification model to output M first text category probabilities corresponding to the sample sentence, inputting the sample sentence into a second text classification model to output M second text category probabilities corresponding to the sample sentence, filtering noise samples from each batch sample set based on the M first text category probabilities to obtain a first clean sample set, filtering noise samples from each batch sample set based on the M second text category probabilities to obtain a second clean sample set, adjusting the parameters of the first text classification model based on the second clean sample set to obtain a first target text classification model, and adjusting the parameters of the second text classification model based on the first clean sample set to obtain a second target text classification model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present application relate to the technical field of natural language processing, and in particular, to a method for training a text classification model, a method for text processing, a device, and a medium. Background Art

[0002] With the rapid development of information technology, deep learning methods have attracted extensive research interest in the field of text classification. However, traditional deep learning-based text classification methods have high requirements for the scale and quality of the dataset.

[0003] Since many errors often occur during the annotation process of many datasets, resulting in many noisy samples with incorrect annotations in the dataset, and deep learning models are very likely to memorize the noisy samples in the dataset due to their huge number of parameters. Therefore, in the case of obvious misannotations in the dataset, how to avoid the model from memorizing the noisy samples is an important challenge.

[0004] Currently, the more commonly used method is to simulate the situation of noise in the dataset during the training process by using a noise transfer matrix. First, train the classification model, and then simultaneously train the noise transfer matrix and the classification model, and then train the classification model alone. After repeating the alternating training several times, the classification model can obtain the ability to classify samples on a clean dataset. However, the alternating training requires more training time than traditional models, and if there are more categories in text classification, the learning difficulty of the noise transfer matrix will be higher, resulting in an increase in time cost and training cost. Summary of the Invention

[0005] Embodiments of the present application provide a method for training a text classification model, a method for text processing, a device, and a medium, which are used to optimize a second text classification model based on a first clean sample set and optimize a first text classification model based on a second clean sample set. It can not only prevent the first text classification model and the second text classification model from falling into self-closure during training, reduce the accumulation of errors caused by noisy samples during the screening process, but also obtain the ability to classify samples on a clean dataset through collaborative learning of two classification models with different model architectures, reducing the training time, thereby reducing the time cost and training cost.

[0006] One aspect of the embodiments of the present application provides a method for training a text classification model, including:

[0007] Obtain multiple batch sample sets corresponding to the target scenario from the original sample dataset, where each batch sample set includes N sample sentences, and N is an integer greater than or equal to 1;

[0008] For each of the N sample sentences, input the sample sentence into the first text classification model, and output M first text category probabilities corresponding to the sample sentence through the first text classification model, where M is an integer greater than or equal to 1;

[0009] For each of the N sample sentences, input the sample sentence into the second text classification model, and output M second text category probabilities corresponding to the sample sentence through the second text classification model, where the second text classification model and the first text classification model are heterogeneous models to each other;

[0010] Based on the M first text category probabilities, perform noise sample filtering on each batch of sample sets to obtain a first clean sample set;

[0011] Based on the M second text category probabilities, perform noise sample filtering on each batch of sample sets to obtain a second clean sample set;

[0012] Based on the second clean sample set, adjust the parameters of the first text classification model to obtain a first target text classification model;

[0013] Based on the first clean sample set, adjust the parameters of the second text classification model to obtain a second target text classification model.

[0014] On the other hand, the present application provides a training device for a text classification model, including:

[0015] An acquisition unit, configured to acquire multiple batches of sample sets corresponding to a target scenario from an original sample dataset, where each batch of sample sets includes N sample sentences, and N is an integer greater than or equal to 1;

[0016] A processing unit, configured to, for each of the N sample sentences, input the sample sentence into the first text classification model, and output M first text category probabilities corresponding to the sample sentence through the first text classification model, where M is an integer greater than or equal to 1;

[0017] The processing unit is further configured to calculate losses according to the M first text category probabilities corresponding to the sample sentence and the number of sample sentences to obtain N first loss values corresponding to each batch of sample sets;

[0018] A determination unit, configured to perform noise sample filtering on each batch of sample sets according to the N first loss values to obtain a first clean sample set;

[0019] The processing unit is further configured to, for each of the N sample sentences, input the sample sentence into the second text classification model, and output M second text category probabilities corresponding to the sample sentence through the second text classification model;

[0020] The processing unit is further configured to calculate a loss according to the M second text category probabilities corresponding to the sample sentences and the number of sample sentences, so as to obtain N second loss values corresponding to each batch of sample sets;

[0021] The determining unit is further configured to perform noise sample filtering on each batch of sample sets according to the N second loss values to obtain a second clean sample set, where K is an integer greater than or equal to 1;

[0022] The processing unit is further configured to adjust the parameters of the first text classification model according to the second clean sample set, the M first text category probabilities corresponding to the sample sentences in the second clean sample set, and the number of sample sentences in the second clean sample set, so as to obtain a first target text classification model;

[0023] The processing unit is further configured to adjust the parameters of the second text classification model according to the first clean sample set, the M second text category probabilities corresponding to the sample sentences in the first clean sample set, and the number of sample sentences in the first clean sample set, so as to obtain a second target text classification model.

[0024] In a possible design, in an implementation manner of another aspect of the embodiments of the present application, the determining unit may specifically be configured to:

[0025] Calculate a loss according to the M first text category probabilities corresponding to the sample sentences and the number of sample sentences, so as to obtain N first loss values corresponding to each batch of sample sets;

[0026] Perform noise sample filtering on each batch of sample sets according to the N first loss values to obtain a first clean sample set;

[0027] Performing noise sample filtering on each batch of sample sets based on the M second text category probabilities includes:

[0028] Calculate a loss according to the M second text category probabilities corresponding to the sample sentences and the number of sample sentences, so as to obtain N second loss values corresponding to each batch of sample sets;

[0029] Perform noise sample filtering on each batch of sample sets according to the N second loss values to obtain a second clean sample set.

[0030] In a possible design, in an implementation manner of another aspect of the embodiments of the present application,

[0031] The processing unit is further configured to calculate the noise sample screening rate corresponding to each batch of sample sets according to the batch number corresponding to each batch of sample sets, the filtering rate corresponding to each batch of sample sets, and the total number of batches;

[0032] The determining unit can specifically be used to: determine the clean sample sentences corresponding to each batch of sample sets according to the noise sample screening rate and the order of the N first loss values from small to large, so as to obtain the first clean sample set;

[0033] The determining unit can specifically be used to: determine the clean sample sentences corresponding to each batch of sample sets according to the noise sample screening rate and the order of the N second loss values from small to large, so as to obtain the second clean sample set.

[0034] In a possible design, in an implementation manner of another aspect of the embodiments of the present application,

[0035] The processing unit is further used to perform word-level encoding processing on each batch of sample sets to obtain word vector encodings corresponding to each word;

[0036] The processing unit can specifically be used to:

[0037] Input the word vector encoding into the first text classification model, and perform sentence vector conversion on the word vector encoding through the first text classification model to obtain the first sample sentence vectors corresponding to each sample sentence;

[0038] Perform category probability prediction on the first sample sentence vectors corresponding to each sample sentence to obtain M first text category probabilities corresponding to the sample sentences;

[0039] The processing unit can specifically be used to:

[0040] Input the word vector encoding into the second text classification model, and perform sentence vector conversion on the word vector encoding through the second text classification model to obtain the second sample sentence vectors corresponding to each sample sentence;

[0041] Perform category probability prediction on the second sample sentence vectors corresponding to each sample sentence to obtain M second text category probabilities corresponding to the sample sentences.

[0042] In a possible design, in an implementation manner of another aspect of the embodiments of the present application,

[0043] The processing unit is further used to perform data augmentation processing on the original sample data set to obtain a strongly data-augmented sample set and a weakly data-augmented sample set respectively corresponding to each batch of sample sets;

[0044] The processing unit is further used to perform word-level encoding processing on the strongly data-augmented sample set and the weakly data-augmented sample set respectively to obtain strongly data word vector encodings and weakly data word vector encodings corresponding to each word;

[0045] The processing unit is further configured to input the strong data word vector encoding and the weak data word vector encoding into the first text classification model respectively, perform sentence vector transformation through the first text classification model, and obtain the first strong data augmented sample sentence vector corresponding to each strong data augmented sample sentence, and the first weak data augmented sample sentence vector corresponding to each weak data augmented sample sentence;

[0046] The processing unit is further configured to input the strong data word vector encoding and the weak data word vector encoding into the second text classification model respectively, perform sentence vector transformation through the second text classification model, and obtain the second strong data augmented sample sentence vector corresponding to each strong data augmented sample sentence, and the second weak data augmented sample sentence vector corresponding to each weak data augmented sample sentence.

[0047] In a possible design, in an implementation manner of another aspect of the embodiments of the present application,

[0048] The processing unit is further configured to calculate losses according to the first strong data augmented sample sentence vector, the first weak data augmented sample sentence vector, and the number of sample sentences, and obtain N third loss values corresponding to each batch of sample sets;

[0049] The processing unit is further configured to calculate losses according to the second strong data augmented sample sentence vector, the second weak data augmented sample sentence vector, and the number of sample sentences, and obtain N fourth loss values corresponding to each batch of sample sets;

[0050] The processing unit may specifically be configured to: adjust the parameters of the first text classification model according to the second clean sample set and the N third loss values to obtain a first target text classification model;

[0051] The processing unit may specifically be configured to: adjust the parameters of the second text classification model according to the first clean sample set and the N fourth loss values to obtain a second target text classification model.

[0052] In a possible design, in an implementation manner of another aspect of the embodiments of the present application,

[0053] The processing unit is further configured to calculate the loss weight corresponding to each batch of sample sets according to the number of batches corresponding to each batch of sample sets;

[0054] The processing unit is further configured to calculate a total loss value based on the first loss value, the second loss value, the third loss value, the fourth loss value, and the loss weight;

[0055] The processing unit may specifically be configured to: adjust the parameters of the first text classification model according to the second clean sample set and the total loss value to obtain a first target text classification model;

[0056] The processing unit can specifically be used to: adjust the parameters of the second text classification model according to the first clean sample set and the total loss value to obtain the second target text classification model.

[0057] In a possible design, in an implementation manner of another aspect of this application, the processing unit can specifically be used to:

[0058] Perform back translation on the original sample data set;

[0059] Perform lexical substitution on the original sample data set;

[0060] Perform random noise injection on the original sample data set;

[0061] Perform literal surface transformation on the original sample data set.

[0062] Another aspect of this application provides a text processing method, including:

[0063] Perform sentence segmentation on the text to be processed to obtain the sentences to be processed;

[0064] Convert the sentences to be processed into vectors to obtain the sentence vectors to be processed;

[0065] Input the sentence vectors to be processed into the aforementioned text classification model, and output the M category probabilities corresponding to the sentence vectors to be processed through the text classification model, where the text classification model is the first target text classification model or the second target text classification model, and M is an integer greater than or equal to 1;

[0066] Determine the target text category of the sentences to be processed according to the M category probabilities.

[0067] Another aspect of this application provides a text processing device, including:

[0068] A processing unit, used to perform sentence segmentation on the text to be processed to obtain the sentences to be processed;

[0069] The processing unit is further used to convert the sentences to be processed into vectors to obtain the sentence vectors to be processed;

[0070] The processing unit is further used to input the sentence vectors to be processed into the above text classification model, and output the M category probabilities corresponding to the sentence vectors to be processed through the text classification model, where the text classification model is the first target text classification model or the second target text classification model, and M is an integer greater than or equal to 1;

[0071] A determination unit, used to determine the target text category of the sentences to be processed according to the M category probabilities.

[0072] On the other hand, this application provides a computer device, including: a memory, a processor, and a bus system;

[0073] The memory is used to store programs;

[0074] The processor is used to implement the methods in the above aspects when executing the programs in the memory;

[0075] The bus system is used to connect the memory and the processor, so that the memory and the processor can communicate.

[0076] On the other hand, this application provides a computer-readable storage medium. Instructions are stored in the computer-readable storage medium. When it runs on a computer, it enables the computer to execute the methods in the above aspects.

[0077] It can be seen from the above technical solutions that the embodiments of this application have the following beneficial effects:

[0078] By outputting M first text category probabilities corresponding to a sample sentence through a first text classification model, and outputting M second text category probabilities corresponding to the sample sentence through a second text classification model that is heterogeneous to the first text classification model, it is possible to filter out noise samples in each batch of sample sets based on the M first text category probabilities to obtain a first clean sample set, and filter out noise samples in each batch of sample sets based on the M second text category probabilities to obtain a second clean sample set. Then, the parameters of the first text classification model can be adjusted based on the second clean sample set to obtain a first target text classification model, and the parameters of the second text classification model can be adjusted based on the first clean sample set to obtain a second target text classification model. In the above manner, it is possible to obtain M first text category probabilities corresponding to each sample sentence through the first text classification model to filter out noise samples that may have annotation errors in the sample set to obtain a first clean sample set. Similarly, M second text category probabilities can be obtained through the second text classification model that is heterogeneous to the first text classification model to filter out noise samples that may have annotation errors in the sample set to obtain a second clean sample set. Then, the second text classification model can be optimized based on the first clean sample set, and the first text classification model can be optimized based on the second clean sample set. This can not only prevent the first text classification model and the second text classification model that is heterogeneous to the first text classification model from falling into self-closure during training, but also filter out noise samples, reduce the accumulation of errors caused by noise samples during the screening process, and can also perform collaborative learning on two classification models that are heterogeneous to each other, so as to obtain the ability to classify samples on a clean data set, without the need to perform multiple alternating trainings on multiple models, reducing the training time, and thus reducing the time cost and training cost. Description of the Drawings

[0079] Figure 1 It is a schematic architecture diagram of the text object control system in the embodiment of the present application;

[0080] Figure 2 It is a flowchart of an embodiment of the training method of the text classification model in the embodiment of the present application;

[0081] Figure 3 It is a flowchart of another embodiment of the training method of the text classification model in the embodiment of the present application;

[0082] Figure 4 It is a flowchart of another embodiment of the training method of the text classification model in the embodiment of the present application;

[0083] Figure 5 It is a flowchart of another embodiment of the training method of the text classification model in the embodiment of the present application;

[0084] Figure 6 It is a flowchart of another embodiment of the training method of the text classification model in the embodiment of the present application;

[0085] Figure 7 It is a flowchart of another embodiment of the training method of the text classification model in the embodiment of the present application;

[0086] Figure 8 It is a flowchart of another embodiment of the training method of the text classification model in the embodiment of the present application;

[0087] Figure 9 It is a schematic diagram of the principle process of the training method of the text classification model in the embodiment of the present application;

[0088] Figure 10 It is a schematic diagram of the principle process of the collaborative learning of the training method of the text classification model in the embodiment of the present application;

[0089] Figure 11 It is a schematic diagram of the principle process of the data augmentation processing of the training method of the text classification model in the embodiment of the present application;

[0090] Figure 12 It is a schematic diagram of the principle process of the consensus learning of the training method of the text classification model in the embodiment of the present application;

[0091] Figure 13 It is a flowchart of an embodiment of the method for text processing in the embodiment of the present application;

[0092] Figure 14 It is a schematic diagram of an embodiment of the training device of the text classification model in the embodiment of the present application;

[0093] Figure 15 It is a schematic diagram of an embodiment of the text processing device in the embodiment of the present application;

[0094] Figure 16 It is a schematic diagram of an embodiment of the computer device in the embodiment of the present application. Detailed implementation manners

[0095] The embodiment of the present application provides a method for training a text classification model, a method for text processing, a device and a medium, which are used to optimize a second text classification model based on a first clean sample set and optimize a first text classification model based on a second clean sample set. It can not only prevent the first text classification model and the second text classification model from getting stuck in self-closure during training, reduce the accumulation of errors caused by noise samples during the screening process, but also obtain the ability to classify samples on a clean data set by co-learning two classification models with different model architectures, reduce the training time, and thus can reduce the time cost and training cost.

[0096] The terms "first", "second", "third", "fourth", etc. (if any) in the specification, claims and drawings of the present application are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "corresponding to" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily limit to those clearly listed steps or units, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0097] For the convenience of understanding, some terms or concepts related to the embodiments of the present application are first explained.

[0098] 1. Convolutional Neural Network (CNN)

[0099] A convolutional neural network is a type of feedforward neural network that contains convolutional calculations and has a deep structure. The convolutional neural network has the ability of feature learning and can perform translation-invariant classification on the input information according to its hierarchical structure.

[0100] 2. Long Short-Term Memory (LSTM)

[0101] The long short-term memory network is a type of recurrent neural network designed to address the long-term dependency problem in general RNNs (recurrent neural networks). All RNNs have a chain-like form of repeating neural network modules. In a standard RNN, this repeating structural module has a very simple structure, such as a tanh layer.

[0102] 3. Batch

[0103] Batch refers to a batch of data with a certain size.

[0104] 4. Back translation

[0105] Back translation means translating some sentences (e.g., Chinese) into sentences in another language (e.g., English), and then translating them back into Chinese. The new samples obtained will have certain differences from the original sentences, but the semantics remain the same and can be used for training.

[0106] 5. Strong data augmentation

[0107] Strong data augmentation means that the samples obtained by the data augmentation method have relatively large differences from the original samples.

[0108] 6. Weak data augmentation

[0109] Weak data augmentation means that the samples obtained by the data augmentation method have relatively small differences from the original samples.

[0110] 7. Robustness

[0111] Robustness can be used to evaluate the ability of a text classification model to generate text representations and classify datasets after being subjected to various interferences, such as mislabeling of the dataset. Among them, high robustness means that after the text classification model is subjected to various interference behaviors (such as noise or encoding), its ability to classify the dataset and generate text representations remains stable.

[0112] With the rapid development of technology, artificial intelligence (AI) has gradually penetrated into all aspects of people's lives. Artificial intelligence has extensive practical significance in text translation, intelligent question answering, sentiment analysis, and other aspects. The emergence of artificial intelligence has also greatly facilitated people's lives. First, a brief introduction to artificial intelligence is given. Artificial intelligence is the theory, method, technology, and application system that uses machines controlled by mathematical computers or digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence is also the study of the design principles and implementation methods of various intelligent machines, enabling machines to have the functions of perception, reasoning, and decision-making.

[0113] Artificial intelligence is an interdisciplinary subject that involves a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning. In natural language processing technology, artificial intelligence can be used to process text and reasonably interpret the words in the text. The training method of the text classification model and the text processing method provided in the embodiments of this application belong to the field of natural language processing technology.

[0114] Natural language processing (NLP) is an important direction in the fields of computer science and artificial intelligence. It studies various theories and methods that can enable effective communication between humans and computers in natural language. Natural language processing is a science that combines linguistics, computer science, and mathematics. Therefore, the research in this field will involve natural language, that is, the language people use in daily life, so it has some close connections with linguistics. Natural language processing technology usually includes technologies such as text processing, semantic understanding, machine translation, robot question answering, and pointing maps.

[0115] It should be understood that the training method of the text classification model provided in this application can be applied to fields such as artificial intelligence, cloud technology, and intelligent transportation, and is used to achieve scenarios such as public opinion discovery or domain classification by training the text classification model.

[0116] It is understandable that in the specific embodiments of the present application, relevant data such as training samples and texts to be processed are involved. When the above embodiments of the present application are applied to specific products or technologies, user permission or consent needs to be obtained, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions.

[0117] To solve the above problems, the present application proposes a method for training a text classification model. This method is applied to Figure 1 the text object control system shown in Figure 1 , Figure 1 which is an architecture schematic diagram of the text object control system in the embodiments of the present application. As shown in Figure 1 , the server obtains the original sample data set provided by the terminal device, and outputs M second text category probabilities corresponding to the sample sentences through the second text classification model that is heterogeneous to the first text classification model. Then, based on the M first text category probabilities, noise sample filtering can be performed on each batch of sample sets to obtain the first clean sample set, and based on the M second text category probabilities, noise sample filtering can be performed on each batch of sample sets to obtain the second clean sample set. Then, the parameters of the first text classification model can be adjusted based on the second clean sample set to obtain the first target text classification model, and the parameters of the second text classification model can be adjusted based on the first clean sample set to obtain the second target text classification model. Through the above method, M first text category probabilities corresponding to each sample sentence can be obtained through the first text classification model to filter out the noise samples that may have annotation errors in the sample set to obtain the first clean sample set. Similarly, M second text category probabilities can be obtained through the second text classification model that is heterogeneous to the first text classification model to filter out the noise samples that may have annotation errors in the sample set to obtain the second clean sample set. Then, the second text classification model can be optimized based on the first clean sample set, and the first text classification model can be optimized based on the second clean sample set. This not only can prevent the first text classification model and the second text classification model that is heterogeneous to the first text classification model from falling into self-closure during training, filter out noise samples, and reduce the accumulation of errors caused by noise samples during the screening process, but also can perform collaborative learning on the two classification models that are heterogeneous to each other, so as to obtain the ability to classify samples on a clean data set, without the need to perform multiple alternating trainings on multiple models, reducing the training time, thereby reducing the time cost and training cost.

[0118] It is understandable that Figure 1Only one type of terminal device is shown herein. In actual scenarios, more types of terminal devices can participate in the data processing process. Terminal devices include, but are not limited to, mobile phones, computers, intelligent voice interaction devices, smart home appliances, vehicle-mounted terminals, etc. The specific quantity and types depend on the actual scenario and are not specifically limited herein. Additionally, Figure 1 one server is shown herein. However, in actual scenarios, multiple servers can also participate. Especially in scenarios of multi-model training interaction, the number of servers depends on the actual scenario and is not specifically limited herein.

[0119] It should be noted that in this embodiment, the server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery network (CDN), and big data and artificial intelligence platforms. The terminal device and the server can be directly or indirectly connected through wired or wireless communication methods. The terminal device and the server can be connected to form a blockchain network, which is not limited in this application.

[0120] To solve the above problems, this application proposes a training method for a text classification model. This method is generally executed by a server or a terminal device, or can also be jointly executed by a server and a terminal device.

[0121] The following will introduce the training method for the text classification model in this application. Please refer to Figure 2 One embodiment of the training method for the text classification model in the embodiment of this application includes:

[0122] In step S101, multiple batch sample sets corresponding to the target scenario are obtained from the original sample dataset. Each batch sample set includes N sample sentences, where N is an integer greater than or equal to 1.

[0123] In this embodiment, in actual scenarios, for example, in various scenarios such as public opinion discovery or domain classification, the target object often generates some comment data for some target scenarios or target products. For example, object A has some comment data for virtual game B. The rating of virtual game B can generally be rated by stars, but the sentiment (positive, negative, neutral) of the comment content may not match the actual star rating, which will constitute mislabeled samples in the comment dataset, that is, a noise sample. However, if manual relabeling or auditing is used, it often wastes a lot of manpower. Therefore, to avoid wasting human costs, this embodiment can perform batch processing on these original sample data that may have mislabeling to obtain multiple batch sample sets corresponding to the target scenario.

[0124] Batch refers to a batch size of data. Each batch sample set can include one or more texts, and each text can include one or more sample sentences, that is, N sample sentences.

[0125] Specifically, if Figure 10 As shown, for a target scenario such as commenting on a virtual game, in each training step, a certain amount of data (such as the same as the hyperparameter batch_size) can be selected from the original sample data set with incorrect annotations, i.e., noise, as a batch sample set to obtain multiple batch sample sets, that is, one training step corresponds to a batch sample set. It can be understood that the first batch sample set, the second batch sample set, and finally the Fth batch sample set can be extracted in sequence according to the order of each training step.

[0126] For example, a batch sample set corresponding to virtual game B is selected from an original sample dataset with incorrect annotations, i.e., noise, such as the review classification dataset. If there is a sample A in the batch sample set, and it is assumed that sample A is a review text with a one-star rating for virtual game B by target object A, such as "Virtual game B was launched much earlier than many games on the market, but the game itself still has many problems and needs to be further optimized", it can be considered that the emotion of the review content is relatively neutral and may not be consistent with the actual rating star, then sample A may constitute an incorrectly labeled sample in the review classification dataset, i.e., a noise sample.

[0127] In step S102, for each of the N sample sentences, the sample sentence is input into a first text classification model, and the first text classification model outputs M first text category probabilities corresponding to the sample sentence, where M is an integer greater than or equal to 1;

[0128] In this embodiment, after obtaining multiple batch sample sets, for each training step, each sample sentence in each batch sample set can be input into the first text classification model respectively, and then, the M first text category probabilities corresponding to each sample sentence can be output by the first text classification model, so that the loss value corresponding to each sample sentence can be calculated based on the M first text category probabilities corresponding to each sample sentence obtained, so that noise samples can be better filtered based on the loss value.

[0129] Among them, the first text classification model can specifically be a convolutional neural network (CNN), and can also be represented as other text classification models, such as the transformer model, etc., without specific limitations here. Among them, a convolutional neural network is a type of feedforward neural network that contains convolutional calculations and has a deep structure. A convolutional neural network has the ability of feature learning and can perform translation-invariant classification on input information according to its hierarchical structure. A convolutional neural network pays more attention to the local semantic information of the text.

[0130] Specifically, as Figure 10 shown, in the training step, in order to facilitate the first classification model to recognize or read the sample sentences, thereby improving the training efficiency of the text classification model to a certain extent, specifically, it can be by first performing character-level encoding processing on the sample sentences in the first batch of sample sets to obtain the vector encodings at the character level corresponding to each sample sentence, that is, the word vector encodings corresponding to each word. Furthermore, input the word vector encodings corresponding to each sample sentence into the first text classification model such as a convolutional neural network. The convolutional neural network can perform sentence encoding on these word vector encodings to obtain the first sample sentence vectors corresponding to each sample sentence. Then, pass the first sample sentence vectors through the fully connected layer and the softmax layer of the convolutional neural network to calculate the prediction probability of the convolutional neural network for the sample sentences, that is, the M first text category probabilities corresponding to each sample sentence. Similarly, in each subsequent training step, the same method of obtaining the M first text category probabilities corresponding to the sample sentences in the first batch of sample sets can be used to input the second batch of sample sets, the third batch of sample sets, etc. of sample sets into the first text classification model in turn, so as to output the M first text category probabilities corresponding to each sample sentence through the first text classification model. Details are not repeated here. Among them, one text category corresponds to one first text category probability. The text categories are set according to the actual application scenarios. For example, in the scenario of commenting on virtual games, the text categories can specifically be positive emotion category, negative emotion category, and neutral emotion category, etc., and can also be other categories, without specific limitations here.

[0131] In step S103, for each of the N sample sentences, input the sample sentence into the second text classification model, and output the M second text category probabilities corresponding to the sample sentence through the second text classification model, where the second text classification model and the first text classification model are heterogeneous models to each other.

[0132] In this embodiment, after obtaining multiple batches of sample sets, for each training step, each sample sentence in each batch of sample sets can be input into the second text classification model respectively. Then, the second text classification model can output the M second text category probabilities corresponding to each sample sentence, so that the loss value corresponding to each sample sentence can be calculated based on the obtained M second text category probabilities of each sample sentence in the subsequent process, thereby better filtering out noisy samples based on the loss value.

[0133] Among them, the second text classification model can adopt a model architecture with a relatively large structural difference from the first text classification model. Since different model architectures have different focuses on text during training, different model architectures will have significant differences in making classification decisions. During training, it is possible to prevent the text classification model from falling into self-enclosed learning, thereby reducing the accumulation of errors caused by mislabeled noisy samples during the screening process. Among them, the second text classification model can specifically be a long short-term memory network (LSTM), and can also be represented as other text classification models, such as the transformer model, etc., which are not specifically limited here. The long short-term memory network is a type of recurrent neural network in time, which is specifically designed to solve the long-term dependence problem existing in general RNNs (recurrent neural networks). All RNNs have a chained form of repeating neural network modules. In a standard RNN, this repeated structural module has a very simple structure, such as a tanh layer.

[0134] It can be understood that in each step of training, the first text classification model and the second text classification model will use exactly the same data for training and updating the model to maintain the learning ability of the first text classification model and the second text classification model under the same data, avoid interference caused by using different data, enable the first text classification model and the second text classification model to better perform collaborative learning, save training time to a certain extent, and thus reduce the time cost and training cost to a certain extent.

[0135] Specifically, such as Figure 9As shown, in a training step, to facilitate the second classification model to recognize or read sample sentences, thereby improving the training efficiency of the text classification model to a certain extent. Specifically, it can be achieved by first performing character-level encoding on the sample sentences in the first batch of sample sets to obtain the vector encodings at the character level corresponding to each sample sentence, that is, the character vector encodings corresponding to each character. Furthermore, inputting the character vector encodings corresponding to each sample sentence into a second text classification model such as a long short-term memory network, the long short-term memory network can perform sentence encoding on these character vector encodings to obtain the second sample sentence vectors corresponding to each sample sentence. Then, passing the second sample sentence vectors through the fully connected layer and softmax layer of the long short-term memory network to calculate the prediction probabilities of the long short-term memory network for the sample sentences, that is, the M second text category probabilities corresponding to each sample sentence. Similarly, in each subsequent training step, in the same way as obtaining the M first text category probabilities corresponding to the sample sentences in the first batch of sample sets, the second batch of sample sets, the third batch of sample sets, etc. can be successively input into the second text classification model to output the M second text category probabilities corresponding to each sample sentence through the second text classification model, which will not be elaborated here. Among them, one text category corresponds to one second text category probability.

[0136] In step S104, based on the M first text category probabilities, noise sample filtering is performed on each batch of sample sets to obtain the first clean sample set;

[0137] In this embodiment, after obtaining the M first text category probabilities, noise sample filtering can be performed on each batch of sample sets based on the M first text category probabilities to obtain the first clean sample set, so that the second text classification model can be optimized using the K first clean sample sets subsequently, realizing collaborative learning of the second text classification model based on the K first clean sample sets, which can avoid the second text classification model getting stuck in self-closure during training, reduce the accumulation of errors caused by noise samples during the screening process, enable the second text classification model to learn the ability to classify samples on the clean sample set faster, save the training time of the second text classification model, and thus reduce the time cost and training cost to a certain extent.

[0138] Specifically, after obtaining the M first text category probabilities, based on the M first text category probabilities, noise sample filtering is performed on each batch of sample sets. Specifically, it can be to calculate the loss according to the M first text category probabilities corresponding to the sample sentences and the number of sample sentences to obtain the N first loss values corresponding to each batch of sample sets. Then, noise sample filtering can be performed on each batch of sample sets according to the N first loss values to obtain the first clean sample set.

[0139] In step S105, based on the M second text category probabilities, noise sample filtering is performed on each batch of sample sets to obtain a second clean sample set;

[0140] In this embodiment, after obtaining the M second text category probabilities, noise sample filtering can be performed on each batch of sample sets based on the M second text category probabilities to obtain a second clean sample set, so that the second clean sample set can be used to optimize the first text classification model subsequently, realizing collaborative learning of the first text classification model based on the second clean sample set, being able to avoid the first text classification model falling into self - enclosure during training, reducing the accumulation of errors caused by noise samples during the screening process, enabling the first text classification model to learn the ability to classify samples on the clean sample set faster, saving the training time of the first text classification model, and thus reducing the time cost and training cost to a certain extent.

[0141] Specifically, after obtaining the M second text category probabilities, based on the M second text category probabilities, noise sample filtering is performed on each batch of sample sets. Specifically, loss calculation can be performed according to the M second text category probabilities corresponding to the sample sentences and the number of sample sentences to obtain N second loss values corresponding to each batch of sample sets. Then, noise sample filtering can be performed on each batch of sample sets according to the N second loss values to obtain a second clean sample set.

[0142] In step S106, parameter adjustment is performed on the first text classification model based on the second clean sample set to obtain a first target text classification model;

[0143] In this embodiment, after obtaining the second clean sample set, for better collaborative learning of the first text classification model, parameter adjustment can be performed on the first text classification model according to the second clean sample set, the M first text category probabilities corresponding to the sample sentences in the second clean sample set, and the number of sample sentences in the second clean sample set to obtain a first target text classification model, being able to optimize the first text classification model by using the second clean sample set, realizing collaborative learning of the first text classification model based on the second clean sample set, being able to avoid the first text classification model falling into self - enclosure during training, reducing the accumulation of errors caused by noise samples during the screening process, enabling the first text classification model to learn the ability to classify samples on the clean sample set faster, saving the training time of the first text classification model, and thus reducing the time cost and training cost to a certain extent.

[0144] Specifically, as Figure 10As shown, after obtaining the second clean sample set, the parameters of the first text classification model can be adjusted according to the second clean sample set, the M first text category probabilities corresponding to the sample sentences in the second clean sample set, and the number of sample sentences in the second clean sample set. Specifically, the second clean sample set corresponding to the first batch of sample sets, the M first text category probabilities corresponding to the sample sentences in this second clean sample set, and the number of sample sentences in this second clean sample set can be substituted into formula (1) of the update loss function corresponding to the first text classification model for loss calculation, so as to obtain the first update loss value corresponding to each sample sentence in the second clean sample set corresponding to the first batch of sample sets. Among them, formula (1) of the update loss function corresponding to the first text classification model is as follows:

[0145]

[0146] Among them, represents the clean data set used in the current training step to train the first text classification model (such as a CNN model), that is, the second clean sample set corresponding to the first batch of sample sets. C represents the total number of samples in the clean data set used in the current training step to train the first text classification model, that is, the number of sample sentences in the second clean sample set. M represents all text categories {1, 2,..., M}, and P iM represents the probability that the i-th sample sentence of the first text classification model belongs to text category M, that is, the first text category probability. y iM ∈{0, 1} can be understood as follows: if the text category of the i-th sample sentence belongs to category M, then the value of y iM is 1, otherwise it is 0.

[0147] Furthermore, after obtaining the first update loss value of the second clean sample set corresponding to the first batch of sample sets, a parameter adjustment operation can be performed on the first text classification model. Specifically, the backpropagation gradient descent algorithm can be used to update the model parameters in the CNN model until convergence, and a first intermediate text classification model can be obtained.

[0148] Furthermore, in each subsequent training step, in the same way as using the second clean sample set corresponding to the first batch of sample sets to adjust the parameters of the first text classification model, the second clean sample sets corresponding to the second batch of sample sets, the third batch of sample sets, etc. can be obtained in sequence to update the model parameters of the first intermediate text classification model updated in the previous training step. Details are not elaborated here until convergence to obtain a first target text classification model.

[0149] In step S107, the parameters of the second text classification model are adjusted based on the first clean sample set to obtain a second target text classification model.

[0150] In this embodiment, after obtaining the first clean sample set, in order to better conduct collaborative learning on the second text classification model, the parameters of the second text classification model can be adjusted according to the first clean sample set, the M second text category probabilities corresponding to the sample sentences in the first clean sample set, and the number of sample sentences in the first clean sample set, so as to obtain the second target text classification model. The second text classification model can be optimized by using the first clean sample set, realizing collaborative learning on the second text classification model based on the first clean sample set, avoiding the second text classification model from getting stuck in self-closure during training, reducing the accumulation of errors caused by noise samples during the screening process, enabling the second text classification model to learn the ability to classify samples on the clean sample set faster, saving the training time of the second text classification model, and thus reducing the time cost and training cost to a certain extent.

[0151] Specifically, as Figure 10 shown, after obtaining the first clean sample set, the parameters of the second text classification model can be adjusted according to the first clean sample set, the M second text category probabilities corresponding to the sample sentences in the first clean sample set, and the number of sample sentences in the first clean sample set. Specifically, the first clean sample set corresponding to the first batch of sample sets, the M second text category probabilities corresponding to the sample sentences in this first clean sample set, and the number of sample sentences in this first clean sample set can be substituted into formula (2) of the update loss function corresponding to the second text classification model for loss calculation, so as to obtain the second update loss value corresponding to each sample sentence in the first clean sample set corresponding to the first batch of sample sets. Among them, formula (2) of the update loss function corresponding to the second text classification model is as follows:

[0152]

[0153] Among them, B cnn represents the clean data set used in the current training step to train the second text classification model (such as the LSTM model), that is, the first clean sample set corresponding to the first batch of sample sets. C represents the total number of samples in the clean data set used in the current training step to train the second text classification model, that is, the number of sample sentences in the first clean sample set. M represents all text categories {1, 2,..., M}, and P iM represents the probability that the i-th sample sentence of the second text classification model belongs to text category M, that is, the second text category probability. y iM ∈{0, 1} can be understood as follows: if the text category of the i-th sample sentence belongs to category M, then the value of y iM is 1, otherwise it is 0.

[0154] Further, after obtaining the second update loss value of the first clean sample set corresponding to the first batch of sample sets, parameter adjustment operations can be performed on the second text classification model. Specifically, the backpropagation gradient descent algorithm can be used to update the model parameters in the LSTM model until convergence, and a second intermediate text classification model can be obtained.

[0155] Further, in each subsequent training step, in the same manner as parameter adjustment of the second text classification model using the first clean sample set corresponding to the first batch of sample sets, the first clean sample sets corresponding to sample sets such as the second batch of sample sets and the third batch of sample sets can be sequentially obtained to update the model parameters of the second intermediate text classification model updated in the previous training step. Details are not described herein again until convergence to obtain a second target text classification model.

[0156] It should be noted that there is no necessary order between step S102 and step S103. Step S102 can be executed first, step S103 can be executed first, or step S102 and step S103 can be executed simultaneously, as long as it is executed after step S101. Specifically, it is not limited here.

[0157] In the embodiments of the present application, a training method for a text classification model is provided. Through the above method, M first text category probabilities corresponding to each sample sentence can be obtained through the first text classification model to filter out noise samples that may have annotation errors in the sample set to obtain a first clean sample set. Similarly, M second text category probabilities can be obtained through the second text classification model that is heterogeneous to the first text classification model to filter out noise samples that may have annotation errors in the sample set to obtain a second clean sample set. Then, the second text classification model can be optimized based on the first clean sample set, and the first text classification model can be optimized based on the second clean sample set. This can not only prevent the first text classification model and the second text classification model that is heterogeneous to the first text classification model from falling into self-closure during training, filter out noise samples, reduce the accumulation of errors caused by noise samples during the screening process, but also obtain the ability to classify samples on a clean data set through collaborative learning of two classification models that are heterogeneous to each other, without the need for multiple alternative trainings of multiple models, reducing the training time, and thus reducing the time cost and training cost.

[0158] Optionally, based on the corresponding embodiments above, in another optional embodiment of the training method for the text classification model provided by the embodiments of the present application, as Figure 2 described above, Figure 3As shown, step S104 filters out noisy samples from each batch of sample sets based on the M first text category probabilities to obtain a first clean sample set, including: step S301 to step S302; step S105 includes: step S303 to step S304;

[0159] In step S301, loss calculation is performed according to the M first text category probabilities corresponding to the sample sentence and the number of sample sentences to obtain N first loss values corresponding to each batch of sample sets;

[0160] In this embodiment, after obtaining the M first text category probabilities corresponding to each sample sentence in each batch of sample sets, loss calculation can be performed according to the M first text category probabilities corresponding to the sample sentence and the number of sample sentences to obtain the loss value corresponding to each sample sentence, that is, the first loss value, so as to obtain N first loss values corresponding to each batch of sample sets, enabling subsequent better filtering of noisy samples based on the N first loss values to obtain a clean sample set for optimizing the text classification model, so that the first text classification model can learn the ability to classify samples on the clean sample set.

[0161] Specifically, as Figure 10 shown, in the training step, after obtaining the M first text category probabilities corresponding to each sample sentence in the first batch of sample sets, the loss of the sample sentences in the first batch of sample sets for the first text classification model such as a convolutional neural network can be calculated through the cross-entropy loss function. Specifically, the M first text category probabilities corresponding to each sample sentence in the first batch of sample sets and the number of sample sentences can be substituted into the formula of the cross-entropy loss function for loss calculation to obtain the first loss value corresponding to each sample sentence in the first batch of sample sets. Among them, the calculation formula (3) of the cross-entropy loss function is as follows:

[0162]

[0163] Among them, B represents the data set used to train the first text classification model in the current training step, that is, the first batch of sample sets, N represents the total number of samples in the data set used to train the first text classification model in the current training step, that is, the number of sample sentences in the first batch of sample sets, M represents all text categories {1, 2,..., M}, P iM represents the probability that the i-th sample sentence belongs to text category M for the first text classification model, that is, the first text category probability, y iM ∈{0, 1} can be understood as follows: if the text category of the i-th sample sentence belongs to category M, then the value of y iM is 1, otherwise it is 0.

[0164] Similarly, in each subsequent training step, in the same way as obtaining the N first loss values corresponding to the first batch of sample sets, the M first text category probabilities and the number of sample sentences corresponding to the sample sentences in the second batch of sample sets, the third batch of sample sets, etc. can be successively substituted into formula (3) of the cross-entropy loss function for loss calculation, so as to obtain the first loss value corresponding to each sample sentence in each batch of sample sets, which will not be elaborated here.

[0165] In step S302, noise sample filtering is performed on each batch of sample sets according to the N first loss values to obtain a first clean sample set;

[0166] In this embodiment, after obtaining the N first loss values corresponding to each batch of sample sets, noise sample filtering can be performed on each batch of sample sets. Then, the set of the remaining sample sentences after filtering can be used as the first clean sample set corresponding to this batch of sample sets, so as to obtain K first clean sample sets, enabling subsequent use of the K first clean sample sets to optimize the second text classification model, realizing collaborative learning of the second text classification model based on the K first clean sample sets, being able to avoid the second text classification model falling into self-closure during training, reducing the accumulation of errors caused by noise samples during the screening process, enabling the second text classification model to learn the ability to classify samples on the clean sample set faster, so as to save the training time of the second text classification model, thereby reducing the time cost and training cost to a certain extent.

[0167] Specifically, as Figure 9 shown, for a sample sentence with a relatively large loss value, it can be considered that this sample sentence is a sample that is relatively difficult to learn, that is, a noise sample. On the contrary, for a sample sentence with a relatively small loss, it can be considered that this sample sentence is a simple sample, that is, a clean sample. Therefore, in the training step, after obtaining the N first loss values corresponding to the first batch of sample sets, when performing noise sample filtering on the first batch of sample sets, this embodiment can filter out low-loss samples. Specifically, the N first loss values corresponding to each sample sentence in the first batch of sample sets can be respectively compared with a preset loss threshold. Then, the sample sentences corresponding to the first loss values greater than or equal to this loss threshold can be determined as noise samples and filtered out; on the contrary, the sample sentences corresponding to the first loss values less than this loss threshold can be determined as clean samples, and the set of these clean samples can be used as the first clean sample set corresponding to the first batch of sample sets. Other methods can also be used to filter noise samples, such as using a screening rate function, etc., which are not specifically limited here.

[0168] Similarly, in each subsequent training step, in the same way as obtaining the first clean sample set corresponding to the first batch of sample sets, the corresponding first clean sample sets of the second batch of sample sets, the third batch of sample sets, etc. can be sequentially obtained, which will not be elaborated here, so as to obtain K first clean sample sets.

[0169] In step S303, according to the M second text category probabilities corresponding to the sample sentences and the number of sample sentences, loss calculation is performed to obtain N second loss values corresponding to each batch of sample sets;

[0170] In this embodiment, after obtaining the M second text category probabilities corresponding to each sample sentence in each batch of sample sets, loss calculation can be performed according to the M second text category probabilities corresponding to the sample sentences and the number of sample sentences, so as to obtain the loss value corresponding to each sample sentence, that is, the second loss value, thereby obtaining N second loss values corresponding to each batch of sample sets, so that subsequent noise sample filtering can be better performed based on the N second loss values, so as to obtain a clean sample set to optimize the text classification model, so that the second text classification model can learn the ability to classify samples on the clean sample set.

[0171] Specifically, as Figure 10 shown, in the training step, after obtaining the M second text category probabilities corresponding to each sample sentence in the first batch of sample sets, the loss of the sample sentences in the first batch of sample sets for the second text classification model such as a long short-term memory network can be calculated through the cross-entropy loss function. Specifically, the M second text category probabilities corresponding to each sample sentence in the first batch of sample sets and the number of sample sentences can be substituted into formula (3) of the above cross-entropy loss function for loss calculation, so as to obtain the second loss value corresponding to each sample sentence in the first batch of sample sets.

[0172] It can be understood that the first text classification model and the second text classification model will use exactly the same data for training. Therefore, in formula (3) of the cross-entropy loss function adopted by the second text classification model such as a long short-term memory network, B represents the data set used to train the second text classification model in the current training step, that is, the first batch of sample sets, N represents the total number of samples in the data set used to train the second text classification model in the current training step, that is, the number of sample sentences in the first batch of sample sets, M represents all text categories {1, 2,..., M}, P iM represents the probability that the i-th sample sentence of the second text classification model belongs to text category M, that is, the second text category probability, y iM ∈{0, 1} can be understood as that if the text category of the i-th sample sentence belongs to category M, then the value of y iM is 1, otherwise it is 0.

[0173] Similarly, in each subsequent training step, in the same way as obtaining the N second loss values corresponding to the first batch of sample sets, the M second text category probabilities and the number of sample sentences corresponding to the sample sentences in the second batch of sample sets, the third batch of sample sets, and other sample sets can be substituted into the formula (3) of the cross-entropy loss function for loss calculation in turn, so as to obtain the second loss value corresponding to each sample sentence in each batch of sample sets, which will not be elaborated here.

[0174] In step S304, noise sample filtering is performed on each batch of sample sets according to the N second loss values to obtain a second clean sample set.

[0175] In this embodiment, after obtaining the N second loss values corresponding to each batch of sample sets, noise sample filtering can be performed on each batch of sample sets. Then, the set of the remaining sample sentences after filtering can be used as the second clean sample set corresponding to the batch of sample sets to obtain the second clean sample set, so that the second clean sample set can be used to optimize the first text classification model in the subsequent steps, realizing collaborative learning of the first text classification model based on the second clean sample set, which can avoid the first text classification model from falling into self-closure during training, reduce the accumulation of errors caused by noise samples in the screening process, enable the first text classification model to learn the ability to classify samples on the clean sample set faster, save the training time of the first text classification model, and thus reduce the time cost and training cost to a certain extent.

[0176] Specifically, as Figure 10 shown, for a sample sentence with a relatively large loss value, it can be considered that the sample sentence is a sample that is difficult to learn, that is, a noise sample. On the contrary, for a sample sentence with a relatively small loss, it can be considered that the sample sentence is a simple sample, that is, a clean sample. Therefore, in the training step, after obtaining the N second loss values corresponding to the first batch of sample sets, noise sample filtering is performed on the first batch of sample sets. In this embodiment, low-loss samples can be screened. Specifically, the N second loss values corresponding to each sample sentence in the first batch of sample sets can be compared with a preset loss threshold respectively. Then, the sample sentences corresponding to the second loss values greater than or equal to the loss threshold can be determined as noise samples and filtered out; on the contrary, the sample sentences corresponding to the second loss values less than the loss threshold can be determined as clean samples, and the set of these clean samples can be used as the second clean sample set corresponding to the first batch of sample sets. Other methods can also be used to filter noise samples, such as using a screening rate function, etc., which are not specifically limited here.

[0177] Similarly, in each subsequent training step, the same method of obtaining the second clean sample set corresponding to the first batch of sample sets can be used to obtain the second clean sample sets corresponding to the second batch of sample sets, the third batch of sample sets, and other sample sets in turn. No further details will be given here to obtain the second clean sample set.

[0178] Optionally, in the above Figure 3 On the basis of the corresponding embodiment, in another optional embodiment of the training method of the text classification model provided in the embodiment of the present application, as Figure 4 As shown, before step S302 filters the noise samples of each batch sample set according to N first loss values to obtain the first clean sample set, the method further includes: step S301; step S302 includes: step S402; and step S304 includes: step S403;

[0179] In step S401, the noise sample screening rate corresponding to each batch of sample sets is calculated according to the batch number corresponding to each batch of sample sets, the filtering rate corresponding to each batch of sample sets, and the total number of batches;

[0180] In this embodiment, before obtaining the N first loss values corresponding to the sample sentences in each batch sample set, the noise sample screening rate corresponding to each batch sample set can be calculated based on the number of batches corresponding to each batch sample set, the filtering rate corresponding to each batch sample set, and the total number of batches, so that the noise samples in each batch sample set can be subsequently screened by the noise sample screening rate to more accurately obtain a set of clean samples, that is, the first clean sample set.

[0181] Among them, the batch number corresponding to each batch sample set refers to the round or batch used to train the first text classification model or the second text classification model in the current training step, such as the batch number of the first batch sample set is expressed as 1, or the batch number of the second batch sample set is expressed as 2, etc. The total number of batches refers to the total rounds or total batches used to train the first text classification model or the second text classification model, that is, the total number of batches has a corresponding relationship with the acquired batch sample sets, such as the total number of batches is expressed as K. The filtering rate corresponding to each batch sample set is set according to the actual application requirements, and is used to indicate the ratio of the maximum number of noise samples to be filtered in the current training step to the total number of samples. In this embodiment, the filtering rate can be set to 0.4.

[0182] Specifically, for a sample sentence with a large loss value, it can be considered that the sample sentence is a difficult-to-learn sample, that is, a noisy sample. On the contrary, for a sample sentence with a small loss, it can be considered that the sample sentence is a simple sample, that is, a clean sample. Therefore, in the training step, in order to better filter the noisy samples in the first batch of sample sets, in this embodiment, the noise sample screening rate corresponding to each batch of sample sets can be calculated according to the batch number corresponding to each batch of sample sets, the filtering rate corresponding to each batch of sample sets, and the total number of batches. Specifically, it can be calculated by substituting the batch number corresponding to the current batch of sample sets, the filtering rate corresponding to the current batch of sample sets, and the total number of batches into the formula (4) of the noise screening rate function to obtain the noise sample screening rate corresponding to each batch of sample sets. The formula (4) of the noise screening rate function is as follows:

[0183]

[0184] Where e represents which round of training the current model is in (e starts from 0), r is the filtering rate, which represents the ratio of the maximum number of noisy samples to be filtered in the current training step to the total number of samples, and g represents that at the g-th epoch, the filtering rate reaches the maximum value. R(e) represents the ratio of the number of samples to be retained in each batch in the e-th round of training to the total number of samples, that is, the noise sample screening rate.

[0185] It can be understood that based on the representation of formula (4), in the initial training rounds, both the first text classification model and the second text classification model can retain more samples, that is, the value of the number of retained samples at the beginning is larger. Then, as the number of training rounds increases, the value of the number of retained samples will become smaller, that is, the gradually increasing number of filtered noisy samples.

[0186] In step S402, according to the noise sample screening rate and the order of the N first loss values from small to large, determine the clean sample sentences corresponding to each batch of sample sets to obtain the first clean sample set;

[0187] In this embodiment, after obtaining the noise sample screening rate corresponding to each batch of sample sets and the N first loss values corresponding to the sample sentences in each batch of sample sets, the noisy samples in each batch of sample sets can be filtered according to the noise sample screening rate and the order of the N first loss values from small to large. Then, the set of the remaining sample sentences after filtering can be used as the first clean sample set corresponding to the batch of sample sets to obtain the first clean sample set, so that the second text classification model can be optimized using the first clean sample set in the subsequent steps, and the collaborative learning of the first text classification model based on the second clean sample set can be realized.

[0188] Specifically, in a training step, after obtaining the noise sample screening rate corresponding to the first batch of sample sets and the N first loss values corresponding to the sample sentences in the first batch of sample sets, the N first loss values corresponding to the sample sentences in the first batch of sample sets can be sorted according to the loss magnitude.

[0189] Furthermore, for a sample sentence with a relatively large loss value, it can be considered that the sample sentence is a relatively difficult-to-learn sample, that is, a noise sample. On the contrary, for a sample sentence with a relatively small loss value, it can be considered that the sample sentence is a simple sample, that is, a clean sample. Therefore, after obtaining the loss ranking for the first text classification model such as a CNN model, the number of clean samples to be retained can be calculated according to formula (5): R(e)*batch_size. Then, the first R(e)*batch_size samples with lower losses can be screened out from the first batch of sample sets after the loss ranking. Furthermore, the set of retained samples can be used as the first clean sample set corresponding to the first batch of sample sets.

[0190] Furthermore, in each subsequent training step, in the same way as obtaining the first clean sample set corresponding to the first batch of sample sets, the first clean sample sets corresponding to the second batch of sample sets, the third batch of sample sets, and other sample sets can be obtained in sequence, which will not be elaborated here, in order to obtain the first clean sample set.

[0191] In step S403, according to the noise sample screening rate and the ascending order of the N second loss values, the clean sample sentences corresponding to each batch of sample sets are determined to obtain the second clean sample set.

[0192] In this embodiment, after obtaining the noise sample screening rate corresponding to each batch of sample sets and the N second loss values corresponding to the sample sentences in each batch of sample sets, the noise samples in each batch of sample sets can be filtered according to the noise sample screening rate and the ascending order of the N second loss values. Then, the set of the remaining sample sentences after filtering can be used as the second clean sample set corresponding to the batch of sample sets to obtain the second clean sample set, so that the second clean sample set can be used to optimize the first text classification model in the subsequent steps, realizing the collaborative learning of the first text classification model based on the second clean sample set.

[0193] Specifically, in a training step, after obtaining the noise sample screening rate corresponding to the first batch of sample sets and the N second loss values corresponding to the sample sentences in the first batch of sample sets, the N second loss values corresponding to the sample sentences in the first batch of sample sets can be sorted according to the loss magnitude.

[0194] Further, after obtaining the loss ranking for the second text classification model such as the LSTM model, the number of clean samples to be retained can also be calculated according to formula (5): R(e)*batch_size. Then, the first R(e)*batch_size samples with lower losses can be screened out from the first batch of samples after loss ranking. Furthermore, the set of retained samples can be used as the second clean sample set corresponding to the first batch of samples.

[0195] Further, in each subsequent training step, in the same way as obtaining the second clean sample set corresponding to the first batch of samples, the second clean sample sets corresponding to the second batch of samples, the third batch of samples, etc. can be obtained in sequence, which will not be elaborated here to obtain the second clean sample set.

[0196] It should be noted that there is no necessary sequence between step S402 and step S403. Step S402 can be executed first, step S403 can be executed first, or step S402 and step S403 can be executed simultaneously, as long as it is executed after step S401, and specific details are not limited here.

[0197] Optionally, based on the above Figure 2 corresponding embodiment, in another optional embodiment of the training method of the text classification model provided by the embodiments of the present application, as Figure 5 shown, before step S102 inputs each sample sentence of the N sample sentences into the first text classification model and outputs the M first text category probabilities corresponding to the sample sentence through the first text classification model, the method further includes: step S501; step S102 includes: step S502 and step S503; and step S103 includes: step S504 and step S505;

[0198] In step S501, each batch of sample sets is subjected to word-level encoding processing to obtain the word vector encoding corresponding to each word;

[0199] In this embodiment, before inputting the sample sentences in each batch of sample sets into the first text classification model and outputting the M first text category probabilities corresponding to the sample sentence through the first text classification model, the sample sentences in each batch of sample sets can be subjected to word-level encoding processing to obtain the vector encoding at the word level corresponding to each sample sentence, that is, the word vector encoding corresponding to each word, so as to facilitate the first classification model to recognize or read the sample sentence, thereby improving the training efficiency of the text classification model to a certain extent.

[0200] Specifically, as Figure 10As shown, before inputting the sample sentences in each batch of sample sets into the first text classification model and outputting the M first text category probabilities corresponding to the sample sentences through the first text classification model, the sample sentences in each batch of sample sets can be subjected to character-level encoding processing. Specifically, the texts in each batch of sample sets can be first segmented into sentences to obtain N sample sentences. Then, the word embedding layer can be used to perform character-level vector encoding on each sample sentence to obtain the character-level vector encoding corresponding to each sample sentence, that is, the word vector encoding corresponding to each character.

[0201] Among them, the word embedding layer can be specifically manifested as Embedding Layer, Word2Vec (Word to Vector), or Doc2Vec (Document to Vector), etc., and can also be manifested as other word embedding algorithms, which are not specifically limited here. Among them, Embedding Layer performs one-hot encoding (hot encoding) on the characters or words in the cleaned sample sentences. Then, a single word or character is represented as a real number vector in a predefined vector space, and each word or character can be mapped to a vector. Among them, the size or dimension of the vector space is specified as a part of the model, such as 50, 100, or 300 dimensions, and the vectors are initialized with small random numbers. Embedding Layer can be used at the front end of the neural network and is supervised using the backpropagation algorithm.

[0202] In step S502, the word vector encoding is input into the first text classification model, and the first text classification model converts the word vector encoding into a sentence vector to obtain the first sample sentence vector corresponding to each sample sentence;

[0203] In step S503, category probability prediction is performed on the first sample sentence vector corresponding to each sample sentence to obtain the M first text category probabilities corresponding to the sample sentence;

[0204] In this embodiment, after performing character-level vector encoding on each sample sentence through the word embedding layer to obtain the word vector encoding, the word vector encoding can be input into the first text classification model, and the first text classification model converts the word vector encoding into a sentence vector to obtain the first sample sentence vector corresponding to each sample sentence, and the first text classification model performs category probability prediction on the first sample sentence vector corresponding to each sample sentence to obtain the M first text category probabilities corresponding to the sample sentence, so that the loss value corresponding to each sample sentence can be calculated based on the obtained M first text category probabilities corresponding to each sample sentence, and thus noise sample filtering can be better performed based on the loss value.

[0205] Specifically, as Figure 10As shown, in a training step, the word vector encodings corresponding to each sample sentence in the first batch of sample sets can be input into a first text classification model such as a convolutional neural network (CNN). The CNN can perform sentence encoding on these word vector encodings. Specifically, through the sentence encoding module in the CNN model, a sentence vector corresponding to each sample sentence can be generated, that is, the first sample sentence vector h corresponding to each sample sentence. cnn .

[0206] Further, after obtaining the first sample sentence vector h corresponding to each sample sentence cnn , a class probability prediction is performed on the first sample sentence vector corresponding to each sample sentence. Specifically, the first sample sentence vector h corresponding to each sample sentence can be cnn transmitted to the fully connected layer and softmax layer of the CNN model to calculate the prediction probability of the CNN model for the sample sentence, that is, the M first text class probabilities corresponding to each sample sentence. Similarly, in each subsequent training step, in the same way as obtaining the M first text class probabilities corresponding to the sample sentences in the first batch of sample sets, the second batch of sample sets, the third batch of sample sets, etc. can be input into the first text classification model in turn to output the M first text class probabilities corresponding to each sample sentence through the first text classification model, which will not be elaborated here.

[0207] In step S504, the word vector encoding is input into the second text classification model, and the second text classification model performs sentence vector transformation on the word vector encoding to obtain the second sample sentence vector corresponding to each sample sentence;

[0208] In step S505, a class probability prediction is performed on the second sample sentence vector corresponding to each sample sentence to obtain the M second text class probabilities corresponding to the sample sentence.

[0209] In this embodiment, after performing word-level vector encoding on each sample sentence through the word embedding layer and obtaining the word vector encoding, the word vector encoding can be input into the second text classification model, and the second text classification model performs sentence vector transformation on the word vector encoding to obtain the second sample sentence vector corresponding to each sample sentence, and the second text classification model performs class probability prediction on the second sample sentence vector corresponding to each sample sentence to obtain the M second text class probabilities corresponding to the sample sentence, so that the loss value corresponding to each sample sentence can be calculated based on the obtained M second text class probabilities corresponding to each sample sentence, and thus noise sample filtering can be better performed based on the loss value.

[0210] Specifically, such as Figure 10As shown, in a training step, the word vector encodings corresponding to each sample sentence in the first batch of sample sets can be input into a second text classification model such as a long short-term memory network (LSTM). The LSTM can perform sentence encoding on these word vector encodings. Specifically, it can generate a sentence vector corresponding to each sample sentence through the sentence encoding module in the LSTM model, that is, the second sample sentence vector h corresponding to each sample sentence. lstm .

[0211] Furthermore, after obtaining the second sample sentence vector h corresponding to each sample sentence lstm , a class probability prediction is performed on the second sample sentence vector corresponding to each sample sentence. Specifically, it can be by passing the second sample sentence vector h corresponding to each sample sentence lstm to the fully connected layer and softmax layer of the LSTM model to calculate the prediction probability of the LSTM model for the sample sentence, that is, the M second text class probabilities corresponding to each sample sentence. Similarly, in each subsequent training step, in the same way of obtaining the M second text class probabilities corresponding to the sample sentences in the first batch of sample sets, the second batch of sample sets, the third batch of sample sets, etc. can be input into the second text classification model in turn to output the M second text class probabilities corresponding to each sample sentence through the second text classification model, which will not be elaborated here.

[0212] It should be noted that there is no necessary order between steps S502 to S503 and steps S504 to S505. Steps S502 to S503 can be executed first, steps S504 to S505 can be executed first, or steps S502 to S503 and steps S504 to S505 can be executed simultaneously, as long as it is executed after step S501, and specific details are not limited here.

[0213] Optionally, based on the corresponding embodiment above Figure 3 , in another optional embodiment of the training method of the text classification model provided by the embodiments of the present application, as Figure 6 shown, after step S101 obtains multiple batches of sample sets corresponding to the target scenario from the original sample dataset, the method further includes:

[0214] In step S601, data augmentation processing is performed on the original sample dataset to obtain a strongly data-augmented sample set and a weakly data-augmented sample set corresponding to each batch of sample sets;

[0215] In this embodiment, after obtaining the original sample dataset, data augmentation processing can be performed on the obtained original sample dataset to construct a strong data augmentation sample set and a weak data augmentation sample set for each batch of sample sets in the original sample dataset. This can not only retain the semantic meaning of each sample sentence in the original sample dataset but also generate other different text representations, enabling the subsequent introduction of the strong data augmentation sample set and the weak data augmentation sample set into the training process of the text classification model. By enriching the sample data, the robustness of the text classification model can be enhanced, and the accuracy of classifying samples on the clean sample set by the text classification model can be improved, thereby improving the accuracy of determining the text category.

[0216] Among them, the data augmentation processing can specifically be manifested as one or more of reverse translation, vocabulary substitution, random noise injection, or literal surface transformation on the obtained original sample dataset. The strong data augmentation sample set refers to a sample set in which the samples obtained by the data augmentation processing have a large difference from the sample sentences in the original batch of sample sets. The weak data augmentation sample set refers to a sample set in which the samples obtained by the data augmentation processing have a small difference from the sample sentences in the original batch of sample sets.

[0217] Specifically, as Figure 9 shown, after obtaining the original sample dataset, data augmentation processing operations of reverse translation can be performed on each batch of sample sets in the obtained original sample dataset respectively to obtain the corresponding strong data augmentation sample set and weak data augmentation sample set for each batch of sample sets.

[0218] In step S602, the strong data augmentation sample set and the weak data augmentation sample set are respectively subjected to character-level encoding processing to obtain the strong data character vector encoding and the weak data character vector encoding corresponding to each character.

[0219] In this embodiment, during the collaborative learning process of the first text classification model and the second text classification model, a certain number of correctly labeled sample sentences may be filtered out, which may cause the text information in the filtered sample sentences to not be able to be re-added to the model learning process of the first text classification model and the second text classification model. Therefore, after obtaining the corresponding strong data augmentation sample set and weak data augmentation sample set for each batch of sample sets, the strong data augmentation sample set and the weak data augmentation sample set can be respectively subjected to character-level encoding processing to obtain the strong data character vector encoding and the weak data character vector encoding corresponding to each character in the same form of expression, so as to facilitate the recognition or reading by the text classification model. Thus, the strong data augmentation sample set and the weak data augmentation sample set can be introduced into the model learning process of the text classification model, enhancing the robustness of the text classification model and improving the accuracy of classifying samples on the clean sample set by the text classification model, thereby improving the accuracy of determining the text category.

[0220] Specifically, as Figure 9 shown, after obtaining the strongly data-augmented sample set and the weakly data-augmented sample set corresponding to each batch of sample sets, in this embodiment, consistent learning can be adopted to perform character-level encoding processing on the strongly data-augmented sample set and the weakly data-augmented sample set respectively. Specifically, the same method as in step S501 of performing character-level encoding processing on each batch of sample sets to obtain the character vector encoding corresponding to each character can be used, so as to obtain the strongly data character vector encoding and the weakly data character vector encoding corresponding to each character expressed in the same form, so that subsequently, based on the strongly data character vector encoding and the weakly data character vector encoding corresponding to each character expressed in the same form, the filtered noise samples can be better reintroduced into the learning process of the text classification model.

[0221] In step S603, the strongly data character vector encoding and the weakly data character vector encoding are respectively input into the first text classification model, and through the first text classification model, sentence vector conversion is performed to obtain the first strongly data-augmented sample sentence vector corresponding to each strongly data-augmented sample sentence, and the first weakly data-augmented sample sentence vector corresponding to each weakly data-augmented sample sentence;

[0222] In this embodiment, after obtaining the strongly data character vector encoding corresponding to the sample sentences in the strongly data-augmented sample set and the weakly data character vector encoding corresponding to the sample sentences in the weakly data-augmented sample set corresponding to each batch of sample sets, the strongly data character vector encoding and the weakly data character vector encoding can be respectively input into the first text classification model, and through the first text classification model, sentence vector conversion is performed to obtain the first strongly data-augmented sample sentence vector corresponding to each strongly data-augmented sample sentence, and the first weakly data-augmented sample sentence vector corresponding to each weakly data-augmented sample sentence, so that subsequently, the first text classification model and the second text classification model can be collaboratively learned based on the strongly data-augmented sample set and the weakly data-augmented sample set corresponding to each batch of sample sets.

[0223] Specifically, as Figure 12 shown, in the training step, two augmented strongly data-augmented samples X i " and weakly data-augmented samples X i'After that, then, a first text classification model such as a CNN model can be used to perform sentence vector transformation through the first text classification model to obtain the first strongly data-augmented sample sentence vectors corresponding to each strongly data-augmented sample sentence and the first weakly data-augmented sample sentence vectors corresponding to each weakly data-augmented sample sentence. Specifically, it can be the same as the method in step S502 of inputting the word vector encoding into the first text classification model and performing sentence vector transformation on the word vector encoding through the first text classification model to obtain the first sample sentence vectors corresponding to each sample sentence, which will not be elaborated here. The strongly data-augmented sample X i " and the weakly data-augmented sample X i ' can be used to generate corresponding prediction vectors, that is, the first strongly data-augmented sample sentence vector Z i ” and the first weakly data-augmented sample sentence vector Z i '.

[0224] In step S604, the strongly data word vector encoding and the weakly data word vector encoding are respectively input into the second text classification model, and sentence vector transformation is performed through the second text classification model to obtain the second strongly data-augmented sample sentence vectors corresponding to each strongly data-augmented sample sentence and the second weakly data-augmented sample sentence vectors corresponding to each weakly data-augmented sample sentence.

[0225] In this embodiment, after obtaining the strongly data word vector encodings corresponding to the sample sentences in the strongly data-augmented sample set corresponding to each batch of sample sets and the weakly data word vector encodings corresponding to the sample sentences in the weakly data-augmented sample set, the strongly data word vector encoding and the weakly data word vector encoding can be respectively input into the second text classification model, and sentence vector transformation is performed through the second text classification model to obtain the second strongly data-augmented sample sentence vectors corresponding to each strongly data-augmented sample sentence and the second weakly data-augmented sample sentence vectors corresponding to each weakly data-augmented sample sentence, so as to enable subsequent collaborative learning of the first text classification model and the second text classification model based on the strongly data-augmented sample set and the weakly data-augmented sample set corresponding to each batch of sample sets.

[0226] Specifically, as Figure 12 shown, in the training step, two augmented strongly data-augmented samples X″ i and weakly data-augmented samples X i'After that, then, a second text classification model such as an LSTM model can be used to perform sentence vector conversion through the second text classification model to obtain the second strongly data-augmented sample sentence vectors corresponding to each strongly data-augmented sample sentence and the second weakly data-augmented sample sentence vectors corresponding to each weakly data-augmented sample sentence. Specifically, it can be the same as the method in step S504 of encoding word vectors and inputting them into the second text classification model, and performing sentence vector conversion on the encoded word vectors through the second text classification model to obtain the second sample sentence vectors corresponding to each sample sentence. This will not be elaborated here. The strongly data-augmented sample X i " and the weakly data-augmented sample X i ' can generate corresponding prediction vectors, that is, the second strongly data-augmented sample sentence vectors Z i ” and the second weakly data-augmented sample sentence vectors Z i '.

[0227] It should be noted that there is no necessary order between step S603 and step S604. Step S603 can be executed first, step S604 can be executed first, or step S603 and step S604 can be executed simultaneously, as long as it is executed after step S602. Specifically, it is not limited here.

[0228] Optionally, based on the above Figure 6 corresponding embodiment, in another optional embodiment of the training method of the text classification model provided by the embodiments of the present application, as Figure 7 shown, after the strongly data-augmented word vector encoding and the weakly data-augmented word vector encoding are respectively input into the second text classification model in step S604, and sentence vector conversion is performed through the second text classification model to obtain the second strongly data-augmented sample sentence vectors corresponding to each strongly data-augmented sample sentence and the second weakly data-augmented sample sentence vectors corresponding to each weakly data-augmented sample sentence, the method further includes: step S701 and step S702; step S106 includes: step S703; step S107 includes: step S'704;

[0229] In step S701, loss calculation is performed according to the first strongly data-augmented sample sentence vectors, the first weakly data-augmented sample sentence vectors, and the number of sample sentences to obtain N third loss values corresponding to each batch of sample sets;

[0230] In this embodiment, after obtaining the first strongly data-augmented sample sentence vectors Z″ i and the first weakly data-augmented sample sentence vectors Z iAfter that, loss calculation can be performed based on the first strong data augmentation sample sentence vectors, the first weak data augmentation sample sentence vectors, and the number of sample sentences to obtain N third loss values corresponding to each batch of sample sets, so that the first text classification model can be optimized based on the N third loss values corresponding to each batch of sample sets obtained subsequently.

[0231] Specifically, through the collaborative learning method, each batch of sample sets can be divided into two parts: clean label data (i.e., the first clean sample set) and noisy label data (noisy samples). For the noisy samples with noisy labels, during the training rounds, their label information can be ignored and they can be used as unlabeled samples. As Figure 12 shown, the final loss function consists of two parts. The first part is the cross-entropy loss, that is, loss calculation is performed in formula (2) of the updated loss function corresponding to the first text classification model to obtain the first updated loss value corresponding to each sample sentence in the second clean sample set corresponding to the first batch of sample sets. The second part of the loss is the consistency learning loss, that is, loss calculation is performed based on the first strong data augmentation sample sentence vectors, the first weak data augmentation sample sentence vectors, and the number of sample sentences to obtain N third loss values corresponding to each batch of sample sets. Specifically, the first strong data augmentation sample sentence vector Z″ can be calculated through formula (6) of the loss function i and the first weak data augmentation sample sentence vector Z′ i The mean square error between them is used as the third loss value. Among them, formula (6) of the loss function is as follows:

[0232]

[0233] Among them, H represents the entire data set used to train the first text classification model at the current training step, that is, the first batch of sample sets, the weak data augmentation sample set, and the strong data augmentation sample set. N represents the total number of sample sentences in the entire data set used to train the first text classification model at the current training step. The meaning of this loss function is to punish the inconsistency between the strong data augmentation sample sentence vector and the weak data augmentation sample corresponding to the same sample sentence.

[0234] In step S702, loss calculation is performed based on the second strong data augmentation sample sentence vectors, the second weak data augmentation sample sentence vectors, and the number of sample sentences to obtain N fourth loss values corresponding to each batch of sample sets;

[0235] In this embodiment, after obtaining the second strong data augmentation sample sentence vector Z″ i and the second weak data augmentation sample sentence vector Z iAfter that, the loss can be calculated based on the second-strongest data-augmented sample sentence vectors, the second-weakest data-augmented sample sentence vectors, and the number of sample sentences, so as to obtain N fourth loss values corresponding to each batch of sample sets, enabling subsequent optimization of the second text classification model based on the N fourth loss values obtained for each batch of sample sets.

[0236] Specifically, it can be understood that in each step of training, the first text classification model and the second text classification model will use exactly the same strongly data-augmented data and weakly data-augmented data for training and updating the model to maintain the learning ability of the first text classification model and the second text classification model under the same data, enabling the first text classification model and the second text classification model to better perform collaborative learning. Therefore, the second-strongest data-augmented sample sentence vector Z i ” and the second-weakest data-augmented sample sentence vector Z i ' can be the same data as the first-strongest data-augmented sample sentence vector Z i ” and the first-weakest data-augmented sample sentence vector Z i '. Therefore, after obtaining the second-strongest data-augmented sample sentence vector Z″ i and the second-weakest data-augmented sample sentence vector Z′ i the loss is calculated according to the second-strongest data-augmented sample sentence vector, the second-weakest data-augmented sample sentence vector, and the number of sample sentences. Specifically, it can be the same as the method in step S701 of calculating the loss based on the first-strongest data-augmented sample sentence vector, the first-weakest data-augmented sample sentence vector, and the number of sample sentences to obtain N third loss values corresponding to each batch of sample sets, which will not be elaborated here, so as to obtain N fourth loss values corresponding to each batch of sample sets.

[0237] In step S703, the parameters of the first text classification model are adjusted according to the second clean sample set and the N third loss values to obtain the first target text classification model;

[0238] In this embodiment, after obtaining the N third loss values, the parameters of the first text classification model can be adjusted according to the second clean sample set, the M first text category probabilities corresponding to the sample sentences in the second clean sample set, the number of sample sentences in the second clean sample set, and the N third loss values, so as to obtain the first target text classification model, which can better introduce the filtered samples into the first text classification model for collaborative learning, enabling the first text classification model to learn the ability to classify samples on the clean sample set faster and more comprehensively, so as to save the training time of the first text classification model, thereby reducing the time cost and training cost to a certain extent.

[0239] Specifically, the parameters of the first text classification model are adjusted according to the second clean sample set, the M first text category probabilities corresponding to the sample sentences in the second clean sample set, the number of sample sentences in the second clean sample set, and the N third loss values. Specifically, it can be to first substitute the second clean sample set corresponding to the first batch of sample sets, the M first text category probabilities corresponding to the sample sentences in the second clean sample set, the number of sample sentences in the second clean sample set, and the N third loss values into the formula (7) of the first target loss function obtained by combining the formula (1) of the updated loss function corresponding to the first text classification model and the formula (6) of the loss function for loss calculation, so as to obtain the first target loss value corresponding to the first batch of sample sets. Among them, the formula (7) of the first target loss function is as follows:

[0240]

[0241] Furthermore, after obtaining the first target loss value corresponding to the first batch of sample sets, a parameter adjustment operation can be performed on the first text classification model. Specifically, the backpropagation gradient descent algorithm can be used to update the model parameters in a text classification model (such as a CNN model) until convergence, and a first intermediate text classification model can be obtained.

[0242] Furthermore, in each subsequent training step, in the same way as using the second clean sample set corresponding to the first batch of sample sets to adjust the parameters of the first text classification model, the first target loss values corresponding to sample sets such as the second batch of sample sets and the third batch of sample sets can be obtained in sequence to update the model parameters of the first intermediate text classification model updated in the previous training step. Details are not described here again until convergence to obtain the first target text classification model.

[0243] In step S704, the parameters of the second text classification model are adjusted according to the first clean sample set and the N fourth loss values to obtain the second target text classification model.

[0244] The parameters of the second text classification model are adjusted according to the first clean sample set, the M second text category probabilities corresponding to the sample sentences in the first clean sample set, the number of sample sentences in the first clean sample set, and the N fourth loss values to obtain the second target text classification model.

[0245] In this embodiment, after obtaining N fourth loss values, the second text classification model can be parameter-adjusted according to the first clean sample set, the M second text category probabilities corresponding to the sample sentences in the first clean sample set, the number of sample sentences in the first clean sample set, and the N fourth loss values, so as to obtain a second target text classification model, which can better introduce the filtered samples into the second text classification model for collaborative learning, enabling the second text classification model to learn the ability to classify samples on the clean sample set faster and more comprehensively, saving the training time of the second text classification model, and thus reducing the time cost and training cost to a certain extent.

[0246] Specifically, parameter-adjusting the second text classification model according to the first clean sample set, the M second text category probabilities corresponding to the sample sentences in the first clean sample set, the number of sample sentences in the first clean sample set, and the N fourth loss values can be specifically to first substitute the first clean sample set corresponding to the first batch of sample sets, the M second text category probabilities corresponding to the sample sentences in the first clean sample set, the number of sample sentences in the first clean sample set, and the N fourth loss values into the formula (8) of the second target loss function obtained by combining formula (2) of the updated loss function corresponding to the second text classification model and formula (6) of the loss function for loss calculation, so as to obtain the second target loss value corresponding to the first batch of sample sets. The formula (8) of the second target loss function is as follows:

[0247]

[0248] Further, after obtaining the second target loss value corresponding to the first batch of sample sets, a parameter adjustment operation can be performed on the second text classification model. Specifically, the backpropagation gradient descent algorithm can be used to update the model parameters in the second text classification model (such as the LSTM model) until convergence, and a second intermediate text classification model can be obtained.

[0249] Further, in each subsequent training step, in the same way as parameter-adjusting the second text classification model using the first clean sample set corresponding to the first batch of sample sets, the second target loss values corresponding to sample sets such as the second batch of sample sets and the third batch of sample sets can be sequentially obtained to update the model parameters of the second intermediate text classification model updated in the previous training step. Details are not elaborated here until convergence to obtain the second target text classification model.

[0250] It should be noted that there is no necessary sequence between step S701 and step S702. Step S701 can be executed first, step S702 can be executed first, or step S701 and step S702 can be executed simultaneously, as long as it is executed after step S604. Specific details are not limited here.

[0251] It can be understood that in this embodiment, the original annotation labels are destroyed on the TREC dataset to obtain mislabeled samples, and experimental comparisons are made. The specific experimental results are shown in Table 1 below:

[0252] Table 1

[0253]

[0254] Among them, as can be seen from Table 1, this embodiment has achieved the best results when the noise ratio exceeds 20%.

[0255] Optionally, on the basis of the above Figure 7 corresponding embodiment, in another optional embodiment of the training method of the text classification model provided by the embodiment of the present application, as Figure 8 shown, before calculating the loss according to the first strongly data-augmented sample sentence vectors, the first weakly data-augmented sample sentence vectors, and the number of sample sentences in step S701 to obtain N third loss values corresponding to each batch of sample sets, the method further includes: step S801 and step S802; step S703 includes: step S803; step S704 includes: step S804;

[0256] In step S801, according to the number of batches corresponding to each batch of sample sets, calculate the loss weight corresponding to each batch of sample sets;

[0257] In this embodiment, before obtaining the N third loss values corresponding to each batch of sample sets, the loss weight corresponding to each batch of sample sets can be calculated first according to the number of batches corresponding to each batch of sample sets, so that the third loss value or the fourth loss value can be obtained based on the loss weight corresponding to each batch of sample sets later, which can effectively avoid the negative impact brought by the inability of the text classification model to produce correct results for unlabeled data when it is unstable in the initial stage of training, so as to enhance the robustness of the text classification model.

[0258] Specifically, since the unsupervised weighting function usually gradually rises from zero along the Gaussian curve in the first 10 training epochs, therefore, according to the number of batches corresponding to each batch of sample sets, calculate the loss weight corresponding to each batch of sample sets. Specifically, the number of batches corresponding to each batch of sample sets can be substituted into formula (9) of the weighting function for calculation, where formula (9) of the weighting function is as follows:

[0259]

[0260] where t represents the number of batches corresponding to the batch of sample sets used to train the first text classification model or the second text classification model at the current training step.

[0261] In step S802, a loss calculation is performed based on the first loss value, the second loss value, the third loss value, the fourth loss value, and the loss weight to obtain a total loss value;

[0262] In this embodiment, after obtaining the first loss value, the second loss value, the third loss value, the fourth loss value, and the loss weight, a loss calculation can be performed based on the first loss value, the second loss value, the third loss value, the fourth loss value, and the loss weight to obtain a total loss value. A weighted sum of the loss values can be performed through a weight function related to one training epoch, more effectively combining the loss values of the supervised task and the unsupervised task, and effectively avoiding the negative impact brought by the inability of the text classification model to produce correct results for unlabeled data during the unstable initial stage of training, so as to enhance the robustness of the text classification model.

[0263] Specifically, performing a loss calculation based on the first loss value, the second loss value, the third loss value, the fourth loss value, and the loss weight can specifically be substituting into the total loss value formula (10) obtained from formula (7) of the first objective loss function, formula (8) of the second objective loss function, and formula (9) of the weighting function for loss calculation to obtain a total loss value. The total loss value formula (10) is as follows:

[0264]

[0265] Among them, w(t) represents the loss weight corresponding to each batch of sample sets, and L is used to represent the set of all clean samples filtered through collaborative learning. It can be understood that for the first text classification model (such as a CNN model), L refers to the set of clean samples filtered through the second text classification model (such as an LSTM model). Similarly, for the second text classification model (such as an LSTM model), L refers to the set of clean samples filtered through the first text classification model (such as a CNN model).

[0266] In step S803, the parameters of the first text classification model are adjusted according to the second clean sample set and the total loss value to obtain a first target text classification model;

[0267] In this embodiment, after obtaining the total loss value, the parameters of the first text classification model can be adjusted according to the second clean sample set and the total loss value to obtain a first target text classification model, which can better introduce the filtered samples into the first text classification model for collaborative learning, enabling the first text classification model to learn the ability to classify samples on the clean sample set faster and more comprehensively, so as to save the training time of the first text classification model, thereby reducing the time cost and training cost to a certain extent.

[0268] Specifically, after obtaining the total loss value corresponding to the first batch of sample sets, parameter adjustment operations can be performed on the first text classification model. Specifically, the backpropagation gradient descent algorithm can be used to update the model parameters in the first text classification model (such as a CNN model) until convergence, and a first intermediate text classification model can be obtained.

[0269] Furthermore, in each subsequent training step, in the same way as using the second clean sample set corresponding to the first batch of sample sets to perform parameter adjustment on the first text classification model, the total loss values corresponding to sample sets such as the second batch of sample sets and the third batch of sample sets can be sequentially obtained to update the model parameters of the first intermediate text classification model updated in the previous training step. Details are not described here again until convergence to obtain a first target text classification model.

[0270] In step S804, parameter adjustment is performed on the second text classification model according to the first clean sample set and the total loss value to obtain a second target text classification model.

[0271] In this embodiment, after obtaining the total loss value, parameter adjustment can be performed on the second text classification model according to the first clean sample set and the total loss value to obtain a second target text classification model, which can better introduce the filtered samples into the second text classification model for collaborative learning, enabling the second text classification model to learn the ability to classify samples on the clean sample set faster and more comprehensively, so as to save the training time of the second text classification model, thereby reducing the time cost and training cost to a certain extent.

[0272] Specifically, after obtaining the total loss value corresponding to the first batch of sample sets, parameter adjustment operations can be performed on the second text classification model. Specifically, the backpropagation gradient descent algorithm can be used to update the model parameters in the second text classification model (such as an LSTM model) until convergence, and a second intermediate text classification model can be obtained.

[0273] Furthermore, in each subsequent training step, in the same way as using the second clean sample set corresponding to the first batch of sample sets to perform parameter adjustment on the second text classification model, the total loss values corresponding to sample sets such as the second batch of sample sets and the third batch of sample sets can be sequentially obtained to update the model parameters of the second intermediate text classification model updated in the previous training step. Details are not described here again until convergence to obtain a second target text classification model.

[0274] Optionally, based on the corresponding embodiment above Figure 6 In another optional embodiment of the training method of the text classification model provided by the embodiments of the present application, data augmentation processing includes one or more of the following:

[0275] Perform back translation on the original sample dataset;

[0276] Perform lexical substitution on the original sample dataset;

[0277] Perform random noise injection on the original sample dataset;

[0278] Perform literal surface transformation on the original sample dataset.

[0279] In this embodiment, in order to be able to retain the semantic meaning of each sample sentence in the original sample dataset while also being able to generate other different text representations to be introduced into the training process of the text classification model, and to enhance the robustness of the text classification model by enriching the sample data. Therefore, this embodiment can perform one or more data augmentation processes such as back translation, lexical substitution, random noise injection, or literal surface transformation on the obtained original sample dataset to construct a strong data augmentation sample set and a weak data augmentation sample set for each batch of sample sets in the original sample dataset.

[0280] It can be understood that the difference between the strong data augmentation sample set obtained by the strong data augmentation method and the original sample dataset will be greater than that of the weak data augmentation sample set obtained by the weak data augmentation method.

[0281] Specifically, as Figure 11 shown, perform back translation on the original sample dataset to construct a strong data augmentation sample set and a weak data augmentation sample set for each batch of sample sets in the original sample dataset. Specifically, the Googletrans method can be used. Among them, the weak data augmentation method refers to performing only one intermediate translation, that is, translating the first language representation corresponding to a sample sentence into the second language representation, and then translating the second language representation back into the first language representation. For example, translating a sample sentence from Chinese to English and then back to Chinese; while the strong data augmentation method requires multiple intermediate translations, that is, after the first language representation corresponding to a sample sentence passes through multiple language translations, it is then translated back into the first language representation. For example, translating a sample sentence from Chinese to English, then back to Chinese, then translated into German, and finally translated back to Chinese. Among them, the first language representation and the second language representation correspond to different languages. The first language representation or the second language representation can specifically be Chinese, English, German, etc., and no specific limitation is made here.

[0282] Furthermore, perform lexical substitution on the original sample dataset to construct a strong data augmentation sample set and a weak data augmentation sample set for each batch of sample sets in the original sample dataset. Specifically, it can be based on the replacement of a thesaurus or word embedding replacement, and other replacement methods can also be used, such as word replacement based on TF-IDF, etc., and no specific limitation is made here.

[0283] Among them, the replacement of the thesaurus refers to extracting a random word from a sample sentence and then using the thesaurus to replace it with its synonym. For example, English synonyms can be found using the WordNet database and then the replacement is performed. Among them, the WordNet database is a database used to describe the relationships between words.

[0284] Among them, word embedding replacement refers to for pre-trained word embeddings, such as those pre-trained by models like Word2Vec, GloVe, FastText, or Sent2Vec, using the embedding space to obtain the nearest neighboring words of the pre-trained word embeddings, and using the obtained nearest neighboring words as replacements for certain words in the sample sentence.

[0285] Among them, TF-IDF-based word replacement means that since words with lower TF-IDF scores are meaningless, they can be replaced without affecting the true label of the sample sentence. Specifically, it can be to calculate the TF-IDF scores of words in the entire text and select the word with the lowest score to replace the original word in the sample sentence.

[0286] Furthermore, random noise injection is performed on the original sample dataset to construct strong data augmentation sample sets and weak data augmentation sample sets for each batch of sample sets in the original sample dataset. Specifically, it can be to adopt spelling error injection or blank noise injection, and other injection methods such as random insertion, random swapping, or random deletion can also be used, which are not specifically restricted here.

[0287] Among them, spelling error injection refers to adding spelling errors to some random words in the sample sentence. Among them, these spelling errors can be added programmatically or using a mapping of common spelling errors (such as an English list), which is not specifically restricted here.

[0288] Among them, blank noise injection refers to using a placeholder token to replace some random words in the sample sentence. Specifically, it can be to use "_" as the placeholder token. Blank noise injection can be used as a method to avoid overfitting on a specific context, or can be used as a smoothing mechanism for text classification models or language models.

[0289] Among them, random insertion refers to first selecting a random word from the sample sentence that is not a stop word, and then finding its synonym and inserting it at a random position in the sample sentence.

[0290] Among them, random swapping refers to randomly swapping the positions of any two words in a sample sentence.

[0291] Among them, random deletion means randomly deleting each word in the sample sentence with a certain probability p, where the probability p is set according to actual application requirements and is not specifically limited here.

[0292] Furthermore, perform literal surface transformation on the original sample data set to construct a strong data augmentation sample set and a weak data augmentation sample set for each batch of sample sets in the original sample data set. Specifically, it can be a simple pattern matching transformation method using regular expression applications. It can be to change the speech form of a sample sentence from contraction to expansion to obtain strong data augmentation samples. For example, "she’s" can be changed to "she is" or "she has", etc.; conversely, change the speech form of a sample sentence from expansion to contraction to obtain weak data augmentation samples. For example, "she has" is changed to "she’s".

[0293] Next, the method for text processing in this application will be introduced. Please refer to Figure 13 , an embodiment of the method for text processing in the embodiment of this application includes:

[0294] In step S1301, perform sentence segmentation on the text to be processed to obtain sentences to be processed;

[0295] In this embodiment, in actual scenarios, such as in various scenarios such as public opinion discovery or domain classification, the target object often generates some comment data on some target scenarios or target products. For example, object S has some comment data on virtual game Q, such as "The map F in the game is too complex, not friendly to novice players, and the experience is very poor". In order for the text classification model to better identify or analyze the text to be processed, the text to be processed can be first subjected to sentence segmentation processing to obtain one or more sentences to be processed.

[0296] Specifically, after obtaining the text to be processed, the text to be processed can be segmented into sentences according to the sentence segmentation delimiter to obtain one or more statements, and then the punctuation marks in each statement are filtered out to obtain clean statements without punctuation marks, that is, the sentences to be processed. Assuming that the comment data of object S on virtual game Q, such as "The map F in the game is too complex, not friendly to novice players, and the experience is very poor", is segmented, at least three sentences to be processed can be obtained, such as the first sentence to be processed "The map F in the game is too complex", the second sentence to be processed "Not friendly to novice players", and the third sentence to be processed "The experience is very poor". Among them, the punctuation marks in the statement can be filtered by means of regular filtering.

[0297] In step S1302, convert the sentences to be processed into vectors to obtain sentence vectors to be processed;

[0298] In this embodiment, after obtaining one or more sentences to be processed, each sentence to be processed can be vectorized to obtain a sentence vector to be processed corresponding to each sentence to be processed, so that the text classification model can better identify or analyze the sentence vector to be processed.

[0299] Specifically, after obtaining one or more sentences to be processed, each sentence to be processed can be vectorized. Specifically, the Word2Vec (Word to Vector) model or the Doc2Vec (Document to Vector) model can be used for vectorization. Other models such as the Glove model can also be used, and no specific limitation is made here. For example, each sentence to be processed can be vectorized based on the Word2Vec model. Specifically, each sentence to be processed can be segmented using a conventional word segmentation algorithm to obtain at least two words. Then, based on the Word2Vec model, vectors can be used to represent each word in the sentence to be processed first, and then a prediction objective function can be used to learn the parameters of these vectors to obtain a sentence vector to be processed corresponding to each sentence to be processed.

[0300] In step S1303, the sentence vector to be processed is input into the text classification model, and the text classification model outputs M category probabilities corresponding to the sentence vector to be processed, where the text classification model is the first target text classification model or the second target text classification model, and M is an integer greater than or equal to 1.

[0301] In this embodiment, after obtaining the sentence vector to be processed, each sentence vector to be processed can be input into the text classification model, and the text classification model outputs M category probabilities corresponding to each sentence vector to be processed, so that the target text category corresponding to each sentence to be processed can be quickly and accurately screened based on the M category probabilities in the subsequent process.

[0302] Among them, the text classification model is the above-mentioned trained first target text classification model or second target text classification model. One text category corresponds to one category probability. For example, in the scenario of some comment data of object S on virtual game Q, the text category can specifically be expressed as a positive emotion category, a negative emotion category, a neutral emotion category, etc.

[0303] Specifically, after obtaining the sentence vector to be processed, the sentence vector to be processed can be input into the text classification model. Specifically, each sentence vector to be processed can pass through the fully connected layer and the softmax layer of the text classification model to calculate the prediction probability of the text classification model for each sentence to be processed, that is, the M category probabilities corresponding to each sentence to be processed.

[0304] In step S1304, the target text category of the sentence to be processed is determined according to the M category probabilities.

[0305] In this embodiment, after obtaining the M category probabilities corresponding to each sentence to be processed, the target text category corresponding to each sentence to be processed can be quickly and accurately screened based on the M category probabilities, so that subsequent data annotation, domain classification, sentiment discovery, etc. of the text to be processed can be performed more accurately based on the target text category corresponding to each sentence to be processed.

[0306] Specifically, after obtaining the M category probabilities corresponding to each sentence to be processed, according to the M category probabilities, the target text category of each sentence to be processed is determined. Specifically, it can be by comparing the M category probabilities pairwise or sorting them by size, or other methods can also be used, which are not specifically limited here, to obtain the category probability with the largest value, and the text category corresponding to the category probability with the largest value can be determined as the target text category of the sentence to be processed.

[0307] Furthermore, if there is only one sentence to be processed in the text to be processed, the target text category of the sentence to be processed can be used as the text category of the text to be processed. Then, the text to be processed can be assigned or sent to the corresponding category database or category management department for storage or processing based on the text category corresponding to the text to be processed.

[0308] Furthermore, if there are multiple sentences to be processed in the text to be processed, the target text category corresponding to each sentence to be processed among the multiple sentences to be processed can be obtained through the above steps S1301 to S1304. For example, in the scenario of some comment data of object S on virtual game Q, the text category can specifically be manifested as a positive sentiment category, a negative sentiment category, a neutral sentiment category, etc. Therefore, after obtaining the target text category corresponding to each sentence to be processed, more accurate text classification of the text to be processed can be performed based on the target text category corresponding to each sentence to be processed. In this embodiment, the weighted sum can also be calculated according to the number of sentences belonging to the same target text category among each sentence to be processed and the category weight corresponding to the target text category to obtain the text score corresponding to the text to be processed. For example, assume that the text category of the first sentence to be processed "The map F in the game is too complex" is the neutral sentiment category, where the category weight corresponding to the neutral sentiment category is 0.3, the text category of the second sentence to be processed "Not friendly to novice players" is the negative sentiment category, where the category weight corresponding to the neutral sentiment category is 0.4, and the text category of the third sentence to be processed "The experience is very poor" is the negative sentiment category. Then, the weighted calculation is performed, such as the text score corresponding to the text to be processed is 1*0.3 + 2*0.4 = 1.1.

[0309] Further, after obtaining the text score corresponding to the text to be processed, the text category corresponding to the text to be processed can be determined according to the mapping relationship between the text score and the text category, or the threshold range into which the text score falls can be matched, and the text category corresponding to the threshold range can be determined as the text category corresponding to the text to be processed. For example, the text score 1.1 corresponding to the text to be processed falls into the threshold range of (1, 3), and the text category corresponding to this threshold range is the negative emotion category. The text category of the text to be processed can also be determined by other means, which are not specifically limited here. Then, based on the text category corresponding to the text to be processed, the text to be processed can be assigned or sent to the corresponding category database or category management department for storage or processing.

[0310] The training device of the text classification model in the present application will be described in detail below. Please refer to Figure 14 , Figure 14 FIG. is a schematic diagram of an embodiment of the training device of the text classification model in the embodiment of the present application. The training device 20 of the text classification model includes:

[0311] An obtaining unit 201, configured to obtain multiple batch sample sets corresponding to a target scenario from an original sample dataset, where each batch sample set includes N sample sentences, and N is an integer greater than or equal to 1;

[0312] A processing unit 202, configured to input each of the N sample sentences into a first text classification model for each of the N sample sentences, and output M first text category probabilities corresponding to the sample sentence through the first text classification model, where M is an integer greater than or equal to 1;

[0313] The processing unit 202 is further configured to calculate a loss according to the M first text category probabilities corresponding to the sample sentence and the number of sample sentences, and obtain N first loss values corresponding to each batch sample set;

[0314] A determining unit 203, configured to filter out noise samples for each batch sample set according to the N first loss values to obtain a first clean sample set, where K is an integer greater than or equal to 1;

[0315] The processing unit 202 is further configured to input each of the N sample sentences into a second text classification model for each of the N sample sentences, and output M second text category probabilities corresponding to the sample sentence through the second text classification model;

[0316] The processing unit 202 is further configured to calculate a loss according to the M second text category probabilities corresponding to the sample sentence and the number of sample sentences, and obtain N second loss values corresponding to each batch sample set;

[0317] The determination unit 203 is further configured to perform noise sample filtering on each batch of sample sets according to N second loss values to obtain a second clean sample set;

[0318] The processing unit 202 is further configured to perform parameter adjustment on the first text classification model according to the second clean sample set, the M first text category probabilities corresponding to the sample sentences in the second clean sample set, and the number of sample sentences in the second clean sample set to obtain a first target text classification model;

[0319] The processing unit 202 is further configured to perform parameter adjustment on the second text classification model according to the first clean sample set, the M second text category probabilities corresponding to the sample sentences in the first clean sample set, and the number of sample sentences in the first clean sample set to obtain a second target text classification model.

[0320] Optionally, based on the above Figure 14 corresponding embodiment, in another embodiment of the training device for a text classification model provided by the embodiments of the present application,

[0321] The processing unit 202 is further configured to calculate the noise sample screening rate corresponding to each batch of sample sets according to the number of batches corresponding to each batch of sample sets, the filtering rate corresponding to each batch of sample sets, and the total number of batches;

[0322] The determination unit 203 is specifically configured to: determine the clean sample sentences corresponding to each batch of sample sets according to the noise sample screening rate and the ascending order of the N first loss values to obtain a first clean sample set;

[0323] The determination unit 203 is specifically configured to: determine the clean sample sentences corresponding to each batch of sample sets according to the noise sample screening rate and the ascending order of the N second loss values to obtain a second clean sample set.

[0324] Optionally, based on the above Figure 14 corresponding embodiment, in another embodiment of the training device for a text classification model provided by the embodiments of the present application,

[0325] The processing unit 202 is further configured to perform character-level encoding processing on each batch of sample sets to obtain a character vector encoding corresponding to each character;

[0326] The processing unit 202 is specifically configured to:

[0327] Input the character vector encoding into the first text classification model, and perform sentence vector conversion on the character vector encoding through the first text classification model to obtain a first sample sentence vector corresponding to each sample sentence;

[0328] Perform class probability prediction on the first sample sentence vector corresponding to each sample sentence to obtain M first text class probabilities corresponding to the sample sentence;

[0329] The processing unit 202 can specifically be used for:

[0330] Input the word vector encoding into the second text classification model, and through the second text classification model, perform sentence vector transformation on the word vector encoding to obtain the second sample sentence vector corresponding to each sample sentence;

[0331] Perform class probability prediction on the second sample sentence vector corresponding to each sample sentence to obtain M second text class probabilities corresponding to the sample sentence.

[0332] Optionally, based on the above Figure 14 corresponding embodiment, in another embodiment of the training device for the text classification model provided by the embodiments of the present application,

[0333] The processing unit 202 is further configured to perform data augmentation processing on the original sample data set to obtain a strongly data-augmented sample set and a weakly data-augmented sample set respectively corresponding to each batch of sample sets;

[0334] The processing unit 202 is further configured to perform word-level encoding processing on the strongly data-augmented sample set and the weakly data-augmented sample set respectively to obtain strongly data word vector encodings and weakly data word vector encodings corresponding to each word;

[0335] The processing unit 202 is further configured to input the strongly data word vector encoding and the weakly data word vector encoding into the first text classification model respectively, and through the first text classification model, perform sentence vector transformation to obtain the first strongly data-augmented sample sentence vector corresponding to each strongly data-augmented sample sentence, and the first weakly data-augmented sample sentence vector corresponding to each weakly data-augmented sample sentence;

[0336] The processing unit 202 is further configured to input the strongly data word vector encoding and the weakly data word vector encoding into the second text classification model respectively, and through the second text classification model, perform sentence vector transformation to obtain the second strongly data-augmented sample sentence vector corresponding to each strongly data-augmented sample sentence, and the second weakly data-augmented sample sentence vector corresponding to each weakly data-augmented sample sentence.

[0337] Optionally, based on the above Figure 14 corresponding embodiment, in another embodiment of the training device for the text classification model provided by the embodiments of the present application,

[0338] The processing unit 202 is further configured to calculate losses according to the first strongly data-augmented sample sentence vector, the first weakly data-augmented sample sentence vector, and the number of sample sentences to obtain N third loss values corresponding to each batch of sample sets;

[0339] The processing unit 202 can specifically be used to: adjust the parameters of the first text classification model according to the second clean sample set, the M first text category probabilities corresponding to the sample sentences in the second clean sample set, the number of sample sentences in the second clean sample set, and the N third loss values, so as to obtain the first target text classification model;

[0340] The processing unit 202 is further used to calculate losses according to the second strongly data-augmented sample sentence vectors, the second weakly data-augmented sample sentence vectors, and the number of sample sentences, so as to obtain the N fourth loss values corresponding to each batch of sample sets;

[0341] The processing unit 202 can specifically be used to: adjust the parameters of the second text classification model according to the first clean sample set, the M second text category probabilities corresponding to the sample sentences in the first clean sample set, the number of sample sentences in the first clean sample set, and the N fourth loss values, so as to obtain the second target text classification model.

[0342] Optionally, on the basis of the above Figure 14 corresponding embodiment, in another embodiment of the training device for the text classification model provided by the embodiments of the present application,

[0343] The processing unit 202 is further used to calculate the loss weight corresponding to each batch of sample sets according to the batch number corresponding to each batch of sample sets;

[0344] The processing unit 202 can specifically be used to: calculate losses according to the first strongly data-augmented sample sentence vectors, the first weakly data-augmented sample sentence vectors, the number of sample sentences, and the loss weight, so as to obtain the N third loss values corresponding to each batch of sample sets;

[0345] The processing unit 202 can specifically be used to: calculate losses according to the second strongly data-augmented sample sentence vectors, the second weakly data-augmented sample sentence vectors, the number of sample sentences, and the loss weight, so as to obtain the N fourth loss values corresponding to each batch of sample sets.

[0346] Optionally, on the basis of the above Figure 14 corresponding embodiment, in another embodiment of the training device for the text classification model provided by the embodiments of the present application, the processing unit 202 can specifically be used to:

[0347] perform back translation on the original sample data set;

[0348] perform vocabulary substitution on the original sample data set;

[0349] perform random noise injection on the original sample data set;

[0350] perform literal surface conversion on the original sample data set.

[0351] The apparatus for text processing in the present application will be described in detail below. Please refer to Figure 15 , Figure 15 which is a schematic diagram of an embodiment of the apparatus for text processing in an embodiment of the present application. The apparatus 30 for text processing includes:

[0352] A processing unit 301, configured to perform sentence segmentation on the text to be processed to obtain sentences to be processed;

[0353] The processing unit 301 is further configured to perform vector transformation on the sentences to be processed to obtain sentence vectors to be processed;

[0354] The processing unit 301 is further configured to input the sentence vectors to be processed into the above-mentioned text classification model, and output M category probabilities corresponding to the sentence vectors to be processed through the text classification model, where the text classification model is a first target text classification model or a second target text classification model, and M is an integer greater than or equal to 1;

[0355] A determination unit 302, configured to determine the target text category of the sentence to be processed according to the M category probabilities.

[0356] On the other hand, the present application provides another schematic diagram of a computer device, as Figure 16 shown, Figure 16 which is a schematic diagram of the structure of a computer device provided in an embodiment of the present application. The computer device 300 may vary greatly due to different configurations or performances, and may include one or more central processing units (CPUs) 310 (for example, one or more processors) and a memory 320, and one or more storage media 330 for storing application programs 331 or data 332 (for example, one or more mass storage devices). Among them, the memory 320 and the storage media 330 may be transient storage or persistent storage. The program stored in the storage media 330 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the computer device 300. Further, the central processor 310 may be configured to communicate with the storage media 330 and execute a series of instruction operations in the storage media 330 on the computer device 300.

[0357] The computer device 300 may further include one or more power supplies 340, one or more wired or wireless network interfaces 350, one or more input / output interfaces 360, and / or one or more operating systems 333, such as Windows Server TM , Mac OS X TM , Unix TM , Linux TM, FreeBSD TM and so on.

[0358] The above computer device 300 is also used to execute the steps in the corresponding embodiments as Figures 2 to 8 shown, or execute the steps in the corresponding embodiments as Figure 12 shown.

[0359] On the other hand, the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps in the method described in the embodiments as Figures 2 to 8 shown, or executes the steps in the corresponding embodiments as Figure 12 shown.

[0360] On the other hand, the present application provides a computer program product including a computer program. When the computer program is executed by a processor, it implements the steps in the method described in the embodiments as Figures 2 to 8 shown, or executes the steps in the corresponding embodiments as Figure 12 shown.

[0361] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the above-described systems, devices, and units can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.

[0362] In the several embodiments provided by the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection can be through some interfaces, and the indirect coupling or communication connection of devices or units can be in electrical, mechanical, or other forms.

[0363] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0364] In addition, in each embodiment of the present application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.

[0365] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of this application. The foregoing storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.

Claims

1. A training method for a text classification model, characterized in that, Including: Obtain multiple batches of sample sets corresponding to the target scenario from the original sample dataset, where each batch of sample sets includes N sample sentences, and N is an integer greater than or equal to 1; For each of the N sample sentences, input the sample sentence into the first text classification model, and output M first text category probabilities corresponding to the sample sentence through the first text classification model, where M is an integer greater than or equal to 1; For each of the N sample sentences, input the sample sentence into the second text classification model, and output M second text category probabilities corresponding to the sample sentence through the second text classification model, where the second text classification model has a different model architecture from the first text classification model; Based on the M first text category probabilities, perform noise sample filtering on each batch of sample sets to obtain a first clean sample set; Based on the M second text category probabilities, perform noise sample filtering on each batch of sample sets to obtain a second clean sample set; Based on the second clean sample set, adjust the parameters of the first text classification model to obtain a first target text classification model; Based on the first clean sample set, adjust the parameters of the second text classification model to obtain a second target text classification model.

2. The method according to claim 1, wherein The performing noise sample filtering on each batch of sample sets based on the M first text category probabilities to obtain a first clean sample set includes: Calculate loss based on the M first text category probabilities corresponding to the sample sentence and the number of sample sentences to obtain N first loss values corresponding to each batch of sample sets; Perform noise sample filtering on each batch of sample sets according to the N first loss values to obtain the first clean sample set; The performing noise sample filtering on each batch of sample sets based on the M second text category probabilities to obtain a second clean sample set includes: Calculate loss based on the M second text category probabilities corresponding to the sample sentence and the number of sample sentences to obtain N second loss values corresponding to each batch of sample sets; Perform noise sample filtering on each batch of sample sets according to the N second loss values to obtain the second clean sample set.

3. The method according to claim 2, wherein Before performing noise sample filtering on each batch of sample sets according to the N first loss values to obtain the first clean sample set, the method further includes: Calculate the noise sample screening rate corresponding to each batch of sample sets according to the batch number corresponding to each batch of sample sets, the filtering rate corresponding to each batch of sample sets, and the total number of batches; The performing noise sample filtering on each batch of sample sets according to the N first loss values to obtain a first clean sample set includes: Determine the clean sample sentences corresponding to each batch of sample sets according to the noise sample screening rate and the order of the N first loss values from small to large to obtain the first clean sample set; The performing noise sample filtering on each batch of sample sets according to the N second loss values to obtain a second clean sample set includes: Determine the clean sample sentences corresponding to each batch of sample sets according to the noise sample screening rate and the ascending order of the N second loss values, so as to obtain the second clean sample set.

4. The method according to claim 1, wherein Before inputting each of the N sample sentences into the first text classification model and outputting the M first text category probabilities corresponding to the sample sentence through the first text classification model, the method further includes: Perform character-level encoding processing on each batch of sample sets to obtain the character vector encoding corresponding to each character; Inputting each of the N sample sentences into the first text classification model and outputting the M first text category probabilities corresponding to the sample sentence through the first text classification model includes: Input the character vector encoding into the first text classification model, and the first text classification model performs sentence vector conversion on the character vector encoding to obtain the first sample sentence vector corresponding to each sample sentence; Perform category probability prediction on the first sample sentence vector corresponding to each sample sentence to obtain the M first text category probabilities corresponding to the sample sentence; Inputting each of the N sample sentences into the second text classification model and outputting the M second text category probabilities corresponding to the sample sentence through the second text classification model includes: Input the character vector encoding into the second text classification model, and the second text classification model performs sentence vector conversion on the character vector encoding to obtain the second sample sentence vector corresponding to each sample sentence; Perform category probability prediction on the second sample sentence vector corresponding to each sample sentence to obtain the M second text category probabilities corresponding to the sample sentence.

5. The method according to claim 2, characterized in that After obtaining multiple batches of sample sets corresponding to the target scenario from the original sample dataset, the method further includes: Perform data augmentation processing on the original sample dataset to obtain a strongly data-augmented sample set and a weakly data-augmented sample set respectively corresponding to each batch of sample sets; Perform character-level encoding processing on the strongly data-augmented sample set and the weakly data-augmented sample set respectively to obtain the strongly data character vector encoding and the weakly data character vector encoding corresponding to each character; Input the strongly data character vector encoding and the weakly data character vector encoding into the first text classification model respectively, and the first text classification model performs sentence vector conversion to obtain the first strongly data-augmented sample sentence vector corresponding to each strongly data-augmented sample sentence and the first weakly data-augmented sample sentence vector corresponding to each weakly data-augmented sample sentence; Input the strongly data character vector encoding and the weakly data character vector encoding into the second text classification model respectively, and the second text classification model performs the sentence vector conversion to obtain the second strongly data-augmented sample sentence vector corresponding to each strongly data-augmented sample sentence and the second weakly data-augmented sample sentence vector corresponding to each weakly data-augmented sample sentence.

6. The method according to claim 5, characterized in that After inputting the encoded strong data word vectors and the encoded weak data word vectors into the second text classification model respectively, and performing the sentence vector transformation through the second text classification model to obtain the second strong data augmentation sample sentence vectors corresponding to each strong data augmentation sample sentence and the second weak data augmentation sample sentence vectors corresponding to each weak data augmentation sample sentence, the method further includes: Calculating losses based on the first strong data augmentation sample sentence vectors, the first weak data augmentation sample sentence vectors, and the number of sample sentences to obtain N third loss values corresponding to each batch of sample sets; Calculating losses based on the second strong data augmentation sample sentence vectors, the second weak data augmentation sample sentence vectors, and the number of sample sentences to obtain N fourth loss values corresponding to each batch of sample sets; The parameter adjustment of the first text classification model based on the second clean sample set to obtain the first target text classification model includes: Adjusting the parameters of the first text classification model according to the second clean sample set and the N third loss values to obtain the first target text classification model; The parameter adjustment of the second text classification model based on the first clean sample set to obtain the second target text classification model includes: Adjusting the parameters of the second text classification model according to the first clean sample set and the N fourth loss values to obtain the second target text classification model.

7. The method according to claim 6, wherein Before calculating losses based on the first strong data augmentation sample sentence vectors, the first weak data augmentation sample sentence vectors, and the number of sample sentences to obtain N third loss values corresponding to each batch of sample sets, the method further includes: Calculating the loss weights corresponding to each batch of sample sets according to the batch numbers corresponding to each batch of sample sets; Calculating the total loss value based on the first loss value, the second loss value, the third loss value, the fourth loss value, and the loss weights; The parameter adjustment of the first text classification model according to the second clean sample set and the N third loss values to obtain the first target text classification model includes: Adjusting the parameters of the first text classification model according to the second clean sample set and the total loss value to obtain the first target text classification model; The parameter adjustment of the second text classification model according to the first clean sample set and the N fourth loss values to obtain the second target text classification model includes: Adjusting the parameters of the second text classification model according to the first clean sample set and the total loss value to obtain the second target text classification model.

8. The method according to claim 5, wherein The data augmentation processing includes one or more of the following: Performing back translation on the original sample data set; Performing word substitution on the original sample data set; Performing random noise injection on the original sample data set; Performing literal surface transformation on the original sample data set.

9. A method for text processing, characterized in that, Including: Performing sentence segmentation on the text to be processed to obtain the sentences to be processed; Convert the sentence to be processed into a vector to obtain a sentence vector to be processed; Input the sentence vector to be processed into the text classification model according to any one of claims 1 to 8, and output M category probabilities corresponding to the sentence vector to be processed through the text classification model, where the text classification model is the first target text classification model or the second target text classification model, and M is an integer greater than or equal to 1; Determine the target text category of the sentence to be processed according to the M category probabilities.

10. A training device for a text classification model, characterized in that, It includes: An acquisition unit for acquiring multiple batch sample sets corresponding to a target scenario from an original sample dataset, where each batch sample set includes N sample sentences, and N is an integer greater than or equal to 1; A processing unit for inputting each of the N sample sentences into a first text classification model and outputting M first text category probabilities corresponding to the sample sentence through the first text classification model, where M is an integer greater than or equal to 1; The processing unit is further configured to input each of the N sample sentences into a second text classification model and output M second text category probabilities corresponding to the sample sentence through the second text classification model, where the second text classification model has a different model architecture from the first text classification model; A determination unit for filtering noise samples from each batch sample set based on the M first text category probabilities to obtain a first clean sample set; The determination unit is further configured to filter noise samples from each batch sample set based on the M second text category probabilities to obtain a second clean sample set; The processing unit is further configured to adjust the parameters of the first text classification model based on the second clean sample set to obtain a first target text classification model; The processing unit is further configured to adjust the parameters of the second text classification model based on the first clean sample set to obtain a second target text classification model.

11. The device according to claim 10, characterized in that, The processing unit is specifically configured to: Calculate loss according to the M first text category probabilities corresponding to the sample sentence and the number of sample sentences to obtain N first loss values corresponding to each batch sample set; Calculate loss according to the M second text category probabilities corresponding to the sample sentence and the number of sample sentences to obtain N second loss values corresponding to each batch sample set; The determination unit is specifically configured to: filter noise samples from each batch sample set according to the N first loss values to obtain the first clean sample set; Filter noise samples from each batch sample set according to the N second loss values to obtain the second clean sample set.

12. The device according to claim 11, characterized in that, The processing unit is further configured to: Calculate the noise sample screening rate corresponding to each batch sample set according to the number of batches corresponding to each batch sample set, the filtering rate corresponding to each batch sample set, and the total number of batches; The determining unit is specifically configured to: determine the clean sample sentences corresponding to each batch of sample sets according to the noise sample screening rate and the order of the N first loss values from small to large, so as to obtain the first clean sample set; Determine the clean sample sentences corresponding to each batch of sample sets according to the noise sample screening rate and the order of the N second loss values from small to large, so as to obtain the second clean sample set.

13. The device according to claim 10, characterized in that, The processing unit is further configured to: Perform character-level encoding processing on each batch of sample sets to obtain a character vector encoding corresponding to each character; The processing unit is specifically configured to: input the character vector encoding into the first text classification model, and perform sentence vector conversion on the character vector encoding through the first text classification model to obtain a first sample sentence vector corresponding to each sample sentence; Perform category probability prediction on the first sample sentence vector corresponding to each sample sentence to obtain M first text category probabilities corresponding to the sample sentence; Input the character vector encoding into the second text classification model, and perform sentence vector conversion on the character vector encoding through the second text classification model to obtain a second sample sentence vector corresponding to each sample sentence; Perform category probability prediction on the second sample sentence vector corresponding to each sample sentence to obtain M second text category probabilities corresponding to the sample sentence.

14. The device according to claim 11, wherein The processing unit is further configured to: Perform data augmentation processing on the original sample data set to obtain a strongly data-augmented sample set and a weakly data-augmented sample set respectively corresponding to each batch of sample sets; Perform character-level encoding processing on the strongly data-augmented sample set and the weakly data-augmented sample set respectively to obtain a strongly data character vector encoding and a weakly data character vector encoding corresponding to each character; Input the strongly data character vector encoding and the weakly data character vector encoding into the first text classification model respectively, and perform sentence vector conversion through the first text classification model to obtain a first strongly data-augmented sample sentence vector corresponding to each strongly data-augmented sample sentence and a first weakly data-augmented sample sentence vector corresponding to each weakly data-augmented sample sentence; Input the strongly data character vector encoding and the weakly data character vector encoding into the second text classification model respectively, and perform the sentence vector conversion through the second text classification model to obtain a second strongly data-augmented sample sentence vector corresponding to each strongly data-augmented sample sentence and a second weakly data-augmented sample sentence vector corresponding to each weakly data-augmented sample sentence.

15. The device according to claim 14, characterized in that, The processing unit is further configured to: Calculate losses according to the first strongly data-augmented sample sentence vector, the first weakly data-augmented sample sentence vector and the number of sample sentences to obtain N third loss values corresponding to each batch of sample sets; Calculate losses according to the second strongly data-augmented sample sentence vector, the second weakly data-augmented sample sentence vector and the number of sample sentences to obtain N fourth loss values corresponding to each batch of sample sets; The processing unit is specifically configured to: Adjust the parameters of the first text classification model according to the second clean sample set and the N third loss values to obtain the first target text classification model; Adjust the parameters of the second text classification model according to the first clean sample set and the N fourth loss values to obtain the second target text classification model.

16. The device according to claim 15, characterized in that, The processing unit is further configured to: Calculate the loss weight corresponding to each batch sample set according to the batch number corresponding to each batch sample set; Perform loss calculation based on the first loss value, the second loss value, the third loss value, the fourth loss value, and the loss weight to obtain the total loss value; Specifically, the processing unit is configured to: Adjust the parameters of the first text classification model according to the second clean sample set and the total loss value to obtain the first target text classification model; Adjust the parameters of the second text classification model according to the first clean sample set and the total loss value to obtain the second target text classification model.

17. The device according to claim 14, characterized in that, Specifically, the processing unit is configured to: Perform back translation on the original sample data set; Perform vocabulary substitution on the original sample data set; Perform random noise injection on the original sample data set; Perform literal surface conversion on the original sample data set.

18. An apparatus for text processing, characterized in that, Comprising: A processing unit for performing sentence segmentation on the text to be processed to obtain sentences to be processed; The processing unit is further configured to convert the sentences to be processed into vectors to obtain sentence vectors to be processed; The processing unit is further configured to input the sentence vectors to be processed into the text classification model according to any one of claims 1 to 8, and output M category probabilities corresponding to the sentence vectors to be processed through the text classification model, where the text classification model is the first target text classification model or the second target text classification model, and M is an integer greater than or equal to 1; A determination unit for determining the target text category of the sentence to be processed according to the M category probabilities.

19. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 8 or the method according to claim 9 are implemented.

20. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 8 or the method according to claim 9 are implemented.

21. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 8 or the method according to claim 9 are implemented.