Method and apparatus for training text classification model, and medium and electronic device
By using multiple teacher models to guide the method of training lightweight student models, the problem of long training and reasoning time caused by the large amount of parameters of the existing text classification model is solved, and more efficient text representation and classification accuracy are achieved.
Patent Information
- Application Number
- PCT/CN2024/127362
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-18
- Filing Date
- 2024-10-25
- Publication Date
- 2025-06-26
AI Technical Summary
The existing text classification model has a large amount of parameters, which leads to a long training, deployment and inference time, making it difficult to meet real-time needs.
Using the knowledge distillation method, lightweight student models are trained through pre-trained multiple teacher models, and the pseudo-standard results of the teacher model and the classification results of the student model are trained.
The teacher model guides students' training of the model, which improves the text representation ability and classification accuracy of the text classification model, and reduces the number of parameters and training time of the model.
Smart Images

Figure CN2024127362_26062025_PF_FP_ABST
Abstract
Description
A training method, device, medium and electronic device for a text classification model Technical Field
[0001] This specification relates to the field of computer technology, and in particular to a training method, device, medium, and electronic device for a text classification model. Background Art
[0002] With the development of information technology, text classification models are being used more and more widely. At the same time, privacy data has also attracted public attention.
[0003] Currently, text classification models are typically trained based on text data and its corresponding categories. However, due to the large number of parameters in text classification models, model training, deployment, and inference are time-consuming. Therefore, knowledge distillation can be used to train lightweight text classification models based on existing large-scale text classification models. Therefore, how to use knowledge distillation to train text classification models is a very important issue.
[0004] Based on this, this specification provides a training method for a text classification model.
[0005] Summary of the Invention
[0006] This specification provides a text classification model training method, device, medium and electronic device to partially solve the above-mentioned problems existing in the related art.
[0007] This manual adopts the following technical solutions.
[0008] This specification provides a training method for a text classification model, comprising: determining a text sample, and determining several pre-trained teacher models; wherein the parameter amounts of each teacher model are different; in order of the parameter amounts of the teacher models from small to large, for each teacher model, executing: inputting the text sample into the teacher model to determine a pseudo-labeling result, and inputting the text sample into a student model to be trained to determine a classification result, training the student model to be trained based on at least the pseudo-labeling result obtained based on the teacher model and the classification result; using the trained student model as a text classification model; wherein the text classification model is used to determine the classification result of the text to be classified based on the text to be classified.
[0009] Optionally, the student model to be trained includes a feature extraction layer and a classification layer; inputting the text sample into the student model to be trained to determine the classification result, specifically including: inputting the text sample into the feature extraction layer of the student model to be trained to determine the feature sequence corresponding to the text sample; using the feature of the position corresponding to the teacher model in the feature sequence as the output feature; inputting the output feature into the classification layer of the student model to be trained to determine the classification result.
[0010] Optionally, the student model to be trained is trained at least based on the pseudo-labeling result obtained based on the teacher model and the classification result, specifically including: taking the feature corresponding to the specified position in the feature sequence as the first feature; inputting the first feature into the classification layer of the student model to be trained to determine the first result; determining the annotation corresponding to the text sample; and training the student model to be trained based on the pseudo-labeling result obtained based on the teacher model, the classification result, the first result and the annotation.
[0011] Optionally, the student model to be trained is trained according to the pseudo-labeling result obtained based on the teacher model, the classification result, the first result and the annotation, specifically including: determining the first task loss according to the first result and the annotation; determining the second task loss according to the pseudo-labeling result and the classification result obtained based on the teacher model; and training the student model to be trained according to the first task loss and the second task loss.
[0012] Optionally, the student model to be trained is trained at least based on the pseudo-labeling result obtained based on the teacher model and the classification result, specifically including: determining other teacher models based on the parameter amount of the teacher model; wherein the parameter amount of the other teacher models is smaller than the parameter amount of the teacher model; taking the feature of the position corresponding to the other teacher model in the feature sequence as the second feature; inputting the second feature into the classification layer of the student model to be trained to determine the second result; determining the pseudo-labeling result corresponding to the other teacher model and using it as the other result; training the student model to be trained at least based on the pseudo-labeling result obtained based on the teacher model, the classification result, the second result and the other results.
[0013] Optionally, the student model to be trained is trained at least based on the pseudo-labeling result, the classification result, the second result and the other results obtained based on the teacher model, specifically including: taking the feature corresponding to the specified position in the feature sequence as the first feature; inputting the first feature into the classification layer of the student model to be trained to determine the first result; determining the annotation corresponding to the text sample; training the student model to be trained based on the pseudo-labeling result, the classification result, the second result, the other results, the first result and the annotation obtained based on the teacher model.
[0014] Optionally, the student model to be trained is trained according to the pseudo-labeling result, the classification result, the second result, the other results, the first result and the annotation obtained based on the teacher model, specifically including: determining the first task loss according to the first result and the annotation; determining the second task loss according to the pseudo-labeling result and the classification result obtained based on the teacher model; determining the third task loss according to the second result and the other results; and training the student model to be trained according to the first task loss, the second task loss and the third task loss.
[0015] Optionally, the student model to be trained is trained according to the first task loss, the second task loss and the third task loss, specifically including: weighting the first task loss, the second task loss and the third task loss respectively according to specified weights; and training the student model to be trained according to the weighted first task loss, the weighted second task loss and the weighted third task loss.
[0016] Optionally, pre-training several teacher models specifically includes: determining the annotations corresponding to the text samples; and training each teacher model to be trained based on the text samples and the annotations.
[0017] Optionally, after determining the text sample, the method further includes: determining an initial student model, and determining a label corresponding to the text sample; training the initial student model based on the text sample and the label to obtain a student model to be trained.
[0018] Optionally, the method further includes: determining the user's input text in response to the user's input operation; determining a pre-stored standard text; using the standard text and the input text as text to be classified; inputting the text to be classified into the text classification model to determine the classification result of the text to be classified; when the classification result is similar, determining the reply text corresponding to the standard text and displaying it to the user.
[0019] This specification provides a training device for a text classification model, comprising: a first determination module, used to determine a text sample and several pre-trained teacher models; wherein the parameter amounts of each teacher model are different; a first training module, used to execute, for each teacher model in ascending order of the parameter amounts of the teacher models: inputting the text sample into the teacher model to determine a pseudo-labeling result, and inputting the text sample into a student model to be trained to determine a classification result, and training the student model to be trained based on at least the pseudo-labeling result and the classification result obtained based on the teacher model; a second determination module, used to use the trained student model as a text classification model; wherein the text classification model is used to determine the classification result of the text to be classified based on the text to be classified.
[0020] Optionally, the student model to be trained includes a feature extraction layer and a classification layer; the first training module is specifically used to input the text sample into the feature extraction layer of the student model to be trained, determine the feature sequence corresponding to the text sample; use the features of the position corresponding to the teacher model in the feature sequence as output features; input the output features into the classification layer of the student model to be trained, and determine the classification results.
[0021] Optionally, the first training module is specifically used to take the feature corresponding to the specified position in the feature sequence as the first feature; input the first feature into the classification layer of the student model to be trained to determine a first result; determine the annotation corresponding to the text sample; and train the student model to be trained based on the pseudo-label result obtained based on the teacher model, the classification result, the first result and the annotation.
[0022] Optionally, the first training module is specifically used to determine a first task loss based on the first result and the labeling; determine a second task loss based on the pseudo-labeling result obtained based on the teacher model and the classification result; and train the student model to be trained based on the first task loss and the second task loss.
[0023] Optionally, the first training module is specifically used to determine other teacher models based on the parameter quantity of the teacher model; wherein the parameter quantity of the other teacher models is smaller than the parameter quantity of the teacher model; the feature of the position corresponding to the other teacher model in the feature sequence is used as the second feature; the second feature is input into the classification layer of the student model to be trained to determine the second result; the pseudo-label result corresponding to the other teacher model is determined and used as the other result; the student model to be trained is trained at least based on the pseudo-label result obtained based on the teacher model, the classification result, the second result and the other results.
[0024] Optionally, the first training module is specifically used to take the feature corresponding to the specified position in the feature sequence as the first feature; input the first feature into the classification layer of the student model to be trained to determine the first result; determine the annotation corresponding to the text sample; and train the student model to be trained based on the pseudo-label result obtained based on the teacher model, the classification result, the second result, the other results, the first result and the annotation.
[0025] Optionally, the first training module is specifically used to determine a first task loss based on the first result and the labeling; determine a second task loss based on the pseudo-labeling result obtained based on the teacher model and the classification result; determine a third task loss based on the second result and the other results; and train the student model to be trained based on the first task loss, the second task loss and the third task loss.
[0026] Optionally, the first training module is specifically used to weight the first task loss, the second task loss and the third task loss respectively according to specified weights; and train the student model to be trained based on the weighted first task loss, the weighted second task loss and the weighted third task loss.
[0027] Optionally, the device further includes: a second training module, configured to determine the annotation corresponding to the text sample; and for each teacher model to be trained, training the teacher model to be trained based on the text sample and the annotation.
[0028] Optionally, after determining the text sample, the first determination module is further used to determine an initial student model and determine the annotation corresponding to the text sample; based on the text sample and the annotation, the initial student model is trained to obtain a student model to be trained.
[0029] Optionally, the device also includes: an application module, used to determine the user's input text in response to the user's input operation; determine a pre-stored standard text; use the standard text and the input text as text to be classified; input the text to be classified into the text classification model to determine the classification result of the text to be classified; when the classification result is similar, determine the reply text corresponding to the standard text and display it to the user.
[0030] This specification provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the training method of the above-mentioned text classification model.
[0031] This specification provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the training method for the above-mentioned text classification model is implemented.
[0032] At least one of the above-mentioned technical solutions adopted in this specification can achieve the following beneficial effects: In the text classification model training method provided in this specification, a text sample and several pre-trained teacher models are determined. Then, in order of the number of parameters of each teacher model from small to large, the text sample is input into the teacher model to determine a pseudo-labeling result, and the text sample is input into the student model to be trained to determine a classification result. The student model to be trained is trained based on at least the pseudo-labeling result and the classification result obtained based on the teacher model. Thereafter, the trained student model is used as the text classification model.
[0033] As can be seen from the above method, when training a text classification model, this method determines text samples and several pre-trained teacher models. Then, in ascending order of the number of parameters of each teacher model, the text sample is input into each teacher model to determine a pseudo-labeling result. The text sample is also input into the student model to be trained to determine a classification result. The student model to be trained is then trained based on at least the pseudo-labeling and classification results obtained based on the teacher model. The trained student model is then used as the text classification model. By having each teacher guide the training of the student model, the text representation capability and classification accuracy of the text classification model are improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] The drawings described herein are used to provide a further understanding of this specification and constitute a part of this specification. The exemplary embodiments and descriptions of this specification are used to explain this specification and do not constitute an improper limitation of this specification. In the drawings:
[0035] FIG1 is a flow chart of a training method for a text classification model provided in this specification;
[0036] FIG2 is a schematic diagram of the structure of a student model provided in this specification;
[0037] FIG3 is a schematic diagram of a characteristic sequence provided in this specification;
[0038] FIG4 is a schematic diagram of a training device for a text classification model provided in this specification;
[0039] FIG5 is a schematic diagram of an electronic device corresponding to FIG1 provided in this specification. DETAILED DESCRIPTION
[0040] To make the objectives, technical solutions, and advantages of this specification more clear, the following will clearly and completely describe the technical solutions of this specification in conjunction with the specific embodiments of this specification and the corresponding drawings. Obviously, the embodiments described are only part of the embodiments of this specification, not all of the embodiments. Based on the embodiments in this specification, all other embodiments obtained by ordinary technicians in this field without making any creative efforts are within the scope of protection of this specification.
[0041] The embodiments of this specification provide a training method, device, medium, and electronic device for a text classification model. The technical solutions provided by the embodiments of this specification are described in detail below in conjunction with the accompanying drawings.
[0042] FIG1 is a flowchart of a text classification model training method provided in this specification, which specifically includes the following steps.
[0043] S100: Determine a text sample and determine several pre-trained teacher models; wherein the number of parameters of each teacher model is different.
[0044] In this specification, a device for training a text classification model can determine text samples and several pre-trained teacher models. The device for training a text classification model can be a server or an electronic device such as a desktop computer or laptop computer. For ease of description, the following description of the text classification model training method provided in this specification uses the server as the execution entity.
[0045] The above-mentioned text samples can be pre-collected text data or sample data from any existing text dataset. The annotation of text samples is related to the text classification model. When the text classification model is used to determine whether text samples are similar, in the intelligent customer service response scenario, the text sample can be pre-collected user-entered text. This user-entered text can be questions about transaction items, such as the size, dimensions, color, thickness, and usage instructions of the transaction items. The user-entered questions can serve as text samples. For example, the user-entered text can be "What is the size?". The user-entered text can also be questions about the transaction process, such as the entry point, initiator, overall process, transaction tools, and end point. For example, the user-entered text can be "How to conduct a transaction." The text sample includes at least two sentences, at least one of which corresponds to a reply text. This reply text is pre-labeled text from the user-entered text. The corresponding annotations for the text samples are either similar or dissimilar. A similar annotation indicates that the sentences in the text sample are similar, while a dissimilar annotation indicates that the sentences in the text sample are dissimilar. In addition, when the text classification model is used to determine the topic of a text sample, the text sample is labeled with various topic types, which are not specifically limited in this specification. For ease of explanation, the following example uses the determination of whether text samples are similar as an example, and the following text samples are labeled with either similar or dissimilar.
[0046] In this specification, the teacher model is a pre-trained model, and can also be any existing text classification model, which is not specifically limited in this specification. Different teacher models have different parameter amounts, that is, the parameter amounts of each teacher model are different, but the purposes of each teacher model are the same. For example, the teacher model includes a first model and a second model, the first model includes 12 network layers, and the second model includes 24 network layers, and the parameter amounts of the first model and the second model are different. In addition, each teacher model can be a model of BERT (Bidirectional Encoder Representation from Transformers) structure, and of course it can also be a model of other structures, which is not specifically limited in this specification.
[0047] S102: In order of the parameter amounts of the teacher models from small to large, for each teacher model, execute: inputting the text sample into the teacher model to determine the pseudo-labeling result, and inputting the text sample into the student model to be trained to determine the classification result, and training the student model to be trained at least based on the pseudo-labeling result obtained based on the teacher model and the classification result.
[0048] The server can execute the following process for each teacher model in order of the parameter amount of each teacher model from small to large: input the text sample into the teacher model to determine the pseudo-labeling result, and input the text sample into the student model to be trained to determine the classification result, and train the student model to be trained at least based on the pseudo-labeling result and the classification result obtained based on the teacher model.
[0049] The student model can be a BERT (Bidirectional Encoder Representation from Transformers) model. The student model has fewer parameters than the teacher model. For example, the student model includes four network layers. Both the classification result and the pseudo-labeling result are used to indicate whether sentences in the text sample are similar. The classification result can be either similar or dissimilar, and the pseudo-labeling result can be either similar or dissimilar.
[0050] When the student model to be trained is trained at least based on the pseudo-label results and classification results obtained based on the teacher model, the server can train the student model to be trained with the goal of at least minimizing the difference between the pseudo-label results and classification results obtained based on the teacher model.
[0051] In this specification, the server inputs text into each teacher model in ascending order of the number of parameters of each teacher model, determines a pseudo-labeling result, and inputs a text sample into a student model to be trained to determine a classification result. The student model to be trained is then trained based at least on the pseudo-labeling result and classification result obtained based on the teacher model. For example, the teacher model includes a first model and a second model, and the number of parameters of the first model is smaller than that of the second model. Therefore, the server first inputs text into the first model to determine a pseudo-labeling result, and inputs a text sample into the student model to be trained to determine a classification result. The student model to be trained is then trained based at least on the pseudo-labeling result and classification result obtained based on the first model. Subsequently, the server inputs text into the second model to determine a pseudo-labeling result, and inputs a text sample into the student model to be trained after being guided by the first model to determine a classification result. The student model to be trained after being guided by the first model is then trained based at least on the pseudo-labeling result and classification result obtained based on the second model.
[0052] In addition, in order to better train the student model and obtain a text classification model, the server can select a reference model from each teacher model, and initialize the model parameters of the student model to be trained according to the model parameters of the reference model. Then train the initialized student model to be trained. Among them, the server can randomly select a reference model from each teacher model, and the server can also select a model with the smallest number of parameters from each teacher model as a reference model. This specification does not make specific restrictions. For example, when the reference model selected by the server includes 12 network layers, and the student model to be trained includes 4 network layers, the server can initialize the model parameters of the student model to be trained according to the parameters of the first 4 network layers of the reference model.
[0053] S104: Using the trained student model as a text classification model; wherein the text classification model is used to determine a classification result of the text to be classified based on the text to be classified.
[0054] The server can use the trained student model as a text classification model. The trained student model is the model obtained in step S102 after being trained with the teacher models in ascending order of parameter size. That is, the model obtained after all teacher models have trained the student model to be trained. The text classification model is used to determine the classification result of the text to be classified.
[0055] As can be seen from the above method, when training a text classification model, the server of this application can determine a text sample and several pre-trained teacher models. Then, according to the parameter amount of each teacher model from small to large, for each teacher model, the text sample is input into the teacher model in turn to determine the pseudo-labeling result, and the text sample is input into the student model to be trained to determine the classification result. At least according to the pseudo-labeling result and the classification result obtained based on the teacher model, the student model to be trained is trained. Afterwards, the trained student model is used as a text classification model. The student model is trained under the guidance of several teachers, and the guidance is gradually provided in the order of the parameter amount of the teacher model from small to large, so that the student model can learn more text representations of the teacher model, and the student model gradually learns the text representations of the teacher model, thereby improving the student model's ability to represent the text and preventing the student model from forgetting the text representation of the teacher model. The trained student model is used as a text classification model to improve the classification accuracy of the text classification model.
[0056] In this specification, the student model to be trained includes a feature extraction layer and a classification layer. Therefore, in step S102 above, when a text sample is input into the student model to be trained and the classification result is determined, as shown in FIG2 , which is a schematic diagram of the structure of a student model provided in this specification, the server can input the text sample into the feature extraction layer of the student model to be trained to determine the features corresponding to the text sample. The features are then input into the classification layer of the student model to be trained to determine the classification result.
[0057] In addition, the server can flexibly align the student model with different teacher models through independent representation, without considering the differences in representation of multiple teacher models. Therefore, in the above step S102, when the text sample is input into the student model to be trained and the classification result is determined, the server can input the text sample into the feature extraction layer of the student model to be trained to determine the feature sequence corresponding to the text sample. The feature of the position corresponding to the teacher model in the feature sequence is used as the output feature. The output feature is input into the classification layer of the student model to be trained to determine the classification result.
[0058] Wherein, as shown in Figure 3, Figure 3 is a schematic diagram of the feature sequence provided in this specification. The feature sequence in Figure 3 includes features corresponding to the CLS position, the teacher position and the text position respectively. The text position is the position corresponding to each word or word in the text. There is a corresponding relationship between the teacher position and the teacher model. There are multiple teacher models to guide the student model for training, and there are as many teacher positions as there are in the feature sequence. Usually, the CLS position is located at the first position of the feature sequence. The teacher position can be between the CLS position and the text position (i.e., the T1 to Tn positions in Figure 3). The teacher position can also be the last position of the feature sequence, i.e., after the text position. This specification does not make specific restrictions. Figure 3 only takes the position corresponding to the two teacher models (i.e., the first teacher position and the second teacher position in Figure 3) as an example, which can be between the CLS position and other text positions. In addition, the features corresponding to the CLS position and the teacher position are all features that characterize the entire text. The above classification result is determined based on the features of the position corresponding to the teacher model in the feature sequence. Therefore, the student model to be trained is subsequently trained at least based on the classification result and the pseudo-label result obtained based on the teacher model.
[0059] In addition, when inputting the text sample into the teacher model in step S102 and determining the pseudo-labeling result, the server can input the text sample into the teacher model, determine the feature sequence corresponding to the text sample, and use the feature corresponding to the specified position in the feature sequence as the text feature. The pseudo-labeling result is then determined based on the text feature. Since the feature sequence determined based on the teacher model only includes the CLS position and the features corresponding to the position, the specified position is the CLS position, and the text feature is the feature corresponding to the CLS position.
[0060] Based on this, in the above step S102, when the student model to be trained is trained based on at least the pseudo-labeling results and classification results obtained based on the teacher model, in addition to training the student model to be trained based on the pseudo-labeling results and classification results obtained based on the teacher model, the server can also determine a first result based on the feature corresponding to the specified position in the feature sequence, and then train the student model to be trained based on the first result and the annotation corresponding to the text sample. Specifically, the server can use the feature corresponding to the specified position in the feature sequence as the first feature. The first feature is input into the classification layer of the student model to be trained to determine the first result. The annotation corresponding to the text sample is determined. Then, the student model to be trained is trained based on the pseudo-labeling results, classification results, first result and annotation obtained based on the teacher model. Among them, the specified position is the CLS position, and the first result is used to characterize whether the sentences in the text sample are similar, and the first result is one of similarity and dissimilarity. The annotation corresponding to the text sample is one of similarity and dissimilarity.
[0061] When training the student model based on the pseudo-labeled results, classification results, first results, and annotations obtained based on the teacher model, the server can train the student model with the goal of minimizing the difference between the pseudo-labeled results and classification results obtained based on the teacher model, and with the goal of minimizing the difference between the first results and annotations. The server can also determine a first task loss based on the first results and annotations. Determine a second task loss based on the pseudo-labeled results and classification results obtained based on the teacher model. The student model is trained based on the first task loss and the second task loss.
[0062] Furthermore, in order to better balance the first task loss and the second task loss so that the student model can progressively learn the text representation of the teacher model, when training the student model to be trained based on the first task loss and the second task loss, the server can weight the first task loss and the second task loss according to specified weights. The student model to be trained is then trained based on the weighted first task loss and the weighted second task loss. The specified weights can be pre-set weights corresponding to each task loss, for example, the weights corresponding to the first task loss and the second task loss can both be 1.
[0063] When determining the first task loss based on the first result and the annotation, the server can calculate the cross-entropy loss between the first result and the annotation and use it as the first task loss. When determining the second task loss based on the pseudo-labeled result and the classification result obtained based on the teacher model, the server can calculate the relative entropy, i.e., the KL divergence, between the pseudo-labeled result and the classification result obtained based on the teacher model and use it as the second task loss.
[0064] In this specification, if the teacher model is the teacher model with the smallest number of parameters among all the teacher models, then when using the teacher model to guide the student model, the server can train the student model to be trained based on the pseudo-label results and classification results obtained based on the teacher model. The server can also train the student model to be trained based on the pseudo-label results, classification results, first results and annotations obtained based on the teacher model. The specific process is consistent with the process in the above step S102 and will not be repeated here.
[0065] However, if the teacher model is not the teacher model with the smallest number of parameters among all the teacher models, that is, there is a teacher model with a smaller number of parameters than the teacher model, and the teacher model with a smaller number of parameters than the teacher model has been trained on the student model to be trained before the teacher model. Therefore, in the above step S102, when the student model to be trained is trained at least based on the pseudo-labeling results and classification results obtained based on the teacher model, the server can determine other teacher models based on the parameter amount of the teacher model. The features of the positions corresponding to other teacher models in the feature sequence are used as the second features. The second features are input into the classification layer of the student model to be trained to determine the second result. The pseudo-labeling results corresponding to other teacher models are determined and used as other results. At least based on the pseudo-labeling results, classification results, second results and other results obtained based on the teacher model, the student model to be trained is trained to prevent the student model from forgetting the text representation learned from other teacher models. The parameters of the other teacher models are smaller than those of the teacher model. Therefore, the other teacher models train the student model to be trained before the teacher model. The server can directly determine the pseudo-labeling results of the other teacher models and use them as other results. The pseudo-labeling results of the other teacher models are disguised results obtained when the other teacher models guide the training of the student model to be trained. The second result is used to indicate whether the sentences in the text sample are similar. The second result can be either similar or dissimilar.
[0066] When the student model to be trained is trained at least based on the pseudo-label results, classification results, second results and other results obtained based on the teacher model, the server can train the student model to be trained with the goal of at least minimizing the difference between the pseudo-label results and classification results obtained based on the teacher model and minimizing the difference between the second results and other results. The server can also determine the second task loss based on the pseudo-label results and classification results obtained based on the teacher model. Determine the third task loss based on the second result and other results. Then, train the student model to be trained based on at least the second task loss and the third task loss. The process of determining the third task loss based on the second result and other results is similar to the process of determining the second task loss based on the pseudo-label results and classification results obtained based on the teacher model, and will not be repeated here.
[0067] Furthermore, to better balance the second task loss and the third task loss so that the student model can progressively learn the text representation of the teacher model, when training the student model based on at least the second task loss and the third task loss, the server can weight the second task loss and the third task loss according to specified weights. The student model is then trained based on at least the weighted second task loss and the weighted third task loss.
[0068] Furthermore, when training the student model to be trained based on at least the pseudo-labeling results, classification results, second results, and other results obtained based on the teacher model, the server may use the feature corresponding to a specified position in the feature sequence as a first feature. The first feature is input into the classification layer of the student model to be trained to determine a first result. A label corresponding to the text sample is determined. The student model to be trained is trained based on the pseudo-labeling results, classification results, second results, other results, first results, and labels obtained based on the teacher model.
[0069] When training the student model to be trained based on the pseudo-labeled results, classification results, second results, other results, first results, and annotations obtained based on the teacher model, the server can train the student model to be trained with the goal of minimizing the difference between the pseudo-labeled results and the classification results obtained by the teacher model, minimizing the difference between the second results and other results, and minimizing the difference between the first result and the annotation. The server can also determine the first task loss based on the first result and the annotation. Determine the second task loss based on the pseudo-labeled results and classification results obtained based on the teacher model. Determine the third task loss based on the second result and other results. Then, train the student model to be trained based on the first task loss, the second task loss, and the third task loss.
[0070] Furthermore, to better balance the first task loss, the second task loss, and the third task loss so that the student model can progressively learn the text representations of each teacher model, when training the student model to be trained based on the first task loss, the second task loss, and the third task loss, the server can weight the first task loss, the second task loss, and the third task loss according to specified weights. The student model to be trained is trained based on the weighted first task loss, the weighted second task loss, and the weighted third task loss.
[0071] In this specification, each of the aforementioned teacher models is trained using text samples and their corresponding annotations. Therefore, when pre-training several teacher models, the server determines the annotations corresponding to the text samples and then trains each teacher model based on the text samples and annotations.
[0072] Specifically, when training the teacher model based on the text sample and the annotation, the server can input the text sample into the teacher model to determine the output result. The teacher model is trained with the goal of minimizing the difference between the output result and the annotation.
[0073] In this specification, after determining the text sample in the above step S100, the server can determine the initial student model and determine the annotation corresponding to the text sample. Then, based on the text sample and the annotation, the initial student model is trained to obtain the student model to be trained. Among them, the model parameters of the initial student model can be initialized according to the model parameters of the teacher model. The specific process is as in the above step S102 and will not be repeated here. The student model to be trained can be a model trained with text samples and annotations. However, when using text samples and annotations to train the initial student model, the server uses the initial student model that has not fully converged as the student model to be trained, that is, the initial student model that has not been fully trained is used as the student model to be trained. Therefore, the server can train the initial student model according to the output results and annotations until the specified number of times is reached, and the initial student model after the last training is used as the student model to be trained.
[0074] In addition, each time a teacher model is used to guide the training of the student model to be trained in step S102, that is, each time a teacher model is used in step S102, in addition to the teacher model with the largest number of parameters, for example, when the teacher model with the smallest number of parameters is used to guide the training of the student model to be trained, the server may not train the student model to be trained to full convergence. Specifically, the server may only use the teacher model with the smallest number of parameters to guide the training of the student model to be trained for a preset number of times. When the preset number of times is reached, the server may use the next teacher model after the teacher model with the smallest number of parameters to guide the training of the student model to be trained. The preset number of times is the number of training times pre-set by the server. However, when the teacher model with the largest number of parameters is used to guide the training of the student model to be trained, the server is required to train the student model to be trained to full convergence. Specifically, to determine when the student model to be trained has fully converged, the server may set an end condition. When the student model to be trained meets the end condition, it is determined that the student model to be trained has fully converged. Subsequently, in step S104, the server may use the trained student model (i.e., the fully converged student model to be trained) as the text classification model. The termination condition may be that the number of training times of the student model to be trained reaches a preset threshold, which may be a value preset by the server. The termination condition may also be that the output results of two consecutive student models to be trained are similar. Of course, the termination condition may also be any other existing termination condition for determining that the model has fully converged, and this specification does not specifically limit this.
[0075] In this specification, after obtaining a text classification model, the server may determine a text to be classified. The text to be classified is input into the text classification model to determine a classification result for the text to be classified. Based on the classification result, the text to be classified is classified. The text to be classified may include at least two sentences, and the classification result may be either similar or dissimilar.
[0076] In this specification, in the intelligent customer service response scenario, after obtaining the text classification model, the server can determine the user's input text in response to the user's input operation. A pre-stored standard text is determined, and the standard text and the input text are used as the text to be classified. The text to be classified is then input into the text classification model to determine the classification result of the text to be classified. When the classification result is similar, the reply text corresponding to the standard text is determined and displayed to the user. Among them, there is a corresponding reply text for the standard text. For example, if the standard text is "What is the size", the reply text corresponding to the standard text can be "The size is 37". The standard text can be the text previously input by the user collected by the server, and the reply text corresponding to the standard text can be pre-annotated by the operator or pre-collected by the server. This specification does not make specific restrictions. When the standard text and the input text are used as the text to be classified, the server can splice the standard text and the input text, and use the spliced text as the text to be classified. The classification result indicates whether the input text and the standard text are similar. The classification result is one of similarity and dissimilarity.
[0077] When inputting the text to be classified into the text classification model and determining the classification result for the text to be classified, the server can input the text to be classified into the feature extraction layer of the text classification model to determine a feature sequence. Features corresponding to designated positions in the feature sequence are determined and used as designated features. Features corresponding to positions in each teacher model in the feature sequence are also determined and used as teacher features. The teacher features are fused with the designated features to obtain classification features. The classification features are then input into the classification layer of the text classification model to determine the classification result.
[0078] When the classification result is dissimilar, the server can redetermine the standard text, and then splice the redetermined standard text with the input text to obtain a new text to be classified, and continue to determine the classification result until the classification result is similar.
[0079] In this manual, the model structures of the initial student model, the student model to be trained, and the text classification model are the same, but they are models obtained in different training stages. These student models, in addition to including a feature extraction layer and a classification layer, may also include a global average pooling layer, i.e., a Pooler layer, which is used for global average pooling of features. Therefore, specifically taking the process of inputting the output feature into the classification layer of the student model to be trained in the above-mentioned step S102 and determining the classification result as an example, the server can input the output feature into the Pooler layer of the student model to be trained to obtain the features after pooling. The features after pooling are then input into the classification layer of the student model to be trained to determine the classification result.
[0080] The above is a training method for a text classification model provided in one or more embodiments of this specification. Based on the same idea, this specification also provides a corresponding training device for a text classification model, as shown in FIG4 .
[0081] FIG4 is a schematic diagram of a training device for a text classification model provided in this specification, which specifically includes:
[0082] A first determination module 200 is used to determine a text sample and a plurality of pre-trained teacher models, wherein each teacher model has a different number of parameters;
[0083] The first training module 202 is configured to, for each teacher model, sequentially execute the following steps, in ascending order of the number of parameters of the teacher models: inputting the text sample into the teacher model to determine a pseudo-labeling result; inputting the text sample into a student model to be trained to determine a classification result; and training the student model to be trained based on at least the pseudo-labeling result and the classification result obtained based on the teacher model;
[0084] The second determination module 204 is configured to use the trained student model as a text classification model; wherein the text classification model is configured to determine a classification result of the text to be classified based on the text to be classified.
[0085] Optionally, the student model to be trained includes a feature extraction layer and a classification layer; the first training module 202 is specifically used to input the text sample into the feature extraction layer of the student model to be trained, determine the feature sequence corresponding to the text sample; use the features of the position corresponding to the teacher model in the feature sequence as output features; input the output features into the classification layer of the student model to be trained, and determine the classification results.
[0086] Optionally, the first training module 202 is specifically used to take the feature corresponding to the specified position in the feature sequence as the first feature; input the first feature into the classification layer of the student model to be trained to determine the first result; determine the annotation corresponding to the text sample; and train the student model to be trained based on the pseudo-label result obtained based on the teacher model, the classification result, the first result and the annotation.
[0087] Optionally, the first training module 202 is specifically used to determine a first task loss based on the first result and the labeling; determine a second task loss based on the pseudo-labeling result and the classification result obtained based on the teacher model; and train the student model to be trained based on the first task loss and the second task loss.
[0088] Optionally, the first training module 202 is specifically used to determine other teacher models based on the parameter quantity of the teacher model; wherein the parameter quantity of the other teacher models is smaller than the parameter quantity of the teacher model; the feature of the position corresponding to the other teacher model in the feature sequence is used as the second feature; the second feature is input into the classification layer of the student model to be trained to determine the second result; the pseudo-label result corresponding to the other teacher model is determined and used as the other result; the student model to be trained is trained at least based on the pseudo-label result obtained based on the teacher model, the classification result, the second result and the other results.
[0089] Optionally, the first training module 202 is specifically used to take the feature corresponding to the specified position in the feature sequence as the first feature; input the first feature into the classification layer of the student model to be trained to determine the first result; determine the annotation corresponding to the text sample; and train the student model to be trained based on the pseudo-label result obtained based on the teacher model, the classification result, the second result, the other results, the first result and the annotation.
[0090] Optionally, the first training module 202 is specifically used to determine a first task loss based on the first result and the labeling; determine a second task loss based on the pseudo-labeling result obtained based on the teacher model and the classification result; determine a third task loss based on the second result and the other results; and train the student model to be trained based on the first task loss, the second task loss and the third task loss.
[0091] Optionally, the first training module 202 is specifically used to weight the first task loss, the second task loss and the third task loss respectively according to specified weights; and train the student model to be trained based on the weighted first task loss, the weighted second task loss and the weighted third task loss.
[0092] Optionally, the apparatus further includes: a second training module 206 for determining the annotation corresponding to the text sample; and for each teacher model to be trained, training the teacher model to be trained based on the text sample and the annotation.
[0093] Optionally, after determining the text sample, the first determination module 200 is further used to determine an initial student model and determine the annotation corresponding to the text sample; based on the text sample and the annotation, the initial student model is trained to obtain a student model to be trained.
[0094] Optionally, the device also includes: an application module 208, which is used to determine the user's input text in response to the user's input operation; determine a pre-stored standard text; use the standard text and the input text as text to be classified; input the text to be classified into the text classification model to determine the classification result of the text to be classified; when the classification result is similar, determine the reply text corresponding to the standard text and display it to the user.
[0095] This specification also provides a computer-readable storage medium, which stores a computer program. The computer program can be used to execute the training method of the text classification model shown in Figure 1 above.
[0096] This specification also provides a schematic diagram of an electronic device shown in Figure 5. As shown in Figure 5, Figure 5 is a schematic diagram of the electronic device provided in this specification corresponding to Figure 1. At the hardware level, the electronic device includes a processor, an internal bus, a network interface, a memory, and a non-volatile memory. Of course, it may also include hardware required for other services. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to implement the training method of the text classification model shown in Figure 1 above.
[0097] Of course, in addition to software implementation, this specification does not exclude other implementation methods, such as logic devices or a combination of software and hardware, etc. That is to say, the execution subject of the following processing flow is not limited to each logic unit, but can also be hardware or logic devices.
[0098] In the 1990s, technological improvements could be clearly distinguished as either hardware improvements (for example, improvements to circuit structures like diodes, transistors, and switches) or software improvements (improvements to process flows). However, with the advancement of technology, many process flow improvements today can now be considered direct improvements to hardware circuit structures. Designers almost always create the corresponding hardware circuit structure by programming the improved process flow into the hardware circuit. Therefore, it cannot be said that a process flow improvement cannot be implemented using hardware modules. For example, a programmable logic device (PLD), such as a field programmable gate array (FPGA), is an integrated circuit whose logical function is determined by user programming. Designers can "integrate" a digital system on a PLD by programming it themselves, without having to hire a chip manufacturer to design and manufacture a dedicated integrated circuit chip. Moreover, nowadays, instead of manually fabricating integrated circuit chips, this programming is mostly done using "logic compiler" software. This is similar to the software compiler used when developing programs. Before compilation, the original code must also be written in a specific programming language, called a hardware description language (HDL). There is not just one HDL, but many, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. The most commonly used are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art will also understand that by simply programming the method flow in one of these hardware description languages and then programming it into an integrated circuit, a hardware circuit that implements the logic method flow can be easily obtained.
[0099] The controller can be implemented in any suitable manner. For example, the controller can take the form of a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., software or firmware) executable by the (micro)processor, logic gates, switches, application-specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art will also know that in addition to implementing the controller in a purely computer-readable program code format, the controller can be implemented in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers by logically programming the method steps. Therefore, such a controller can be considered a hardware component, and the devices included therein for implementing various functions can also be considered as structures within the hardware component. Or even, the devices for implementing various functions can be considered as both software modules that implement the method and structures within the hardware component.
[0100] The systems, devices, modules, or units described in the above embodiments may be implemented by computer chips or entities, or by products having certain functions. A typical implementation device is a computer. Specifically, the computer may be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smartphone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.
[0101] For the convenience of description, the above devices are described as being divided into various units according to their functions. Of course, when implementing this specification, the functions of each unit can be implemented in the same or multiple software and / or hardware.
[0102] It will be understood by those skilled in the art that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0103] The present invention is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device produce a device for implementing the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.
[0104] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce a product including an instruction device that implements the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.
[0105] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.
[0106] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.
[0107] Memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.
[0108] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media (transitory media), such as modulated data signals and carrier waves.
[0109] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.
[0110] Those skilled in the art will appreciate that the embodiments of this specification may be provided as methods, systems, or computer program products. Thus, this specification may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0111] This specification may be described in the general context of computer-executable instructions, such as program modules, executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, and the like that perform specific tasks or implement specific abstract data types. This specification may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communications network. In a distributed computing environment, program modules may be located in both local and remote computer storage media, including storage devices.
[0112] The various embodiments in this specification are described in a progressive manner. Similar parts between the various embodiments can be referred to in conjunction with each other. Each embodiment focuses on the differences between the other embodiments. In particular, the system embodiments are generally similar to the method embodiments, so the description is relatively simple. For relevant parts, refer to the description of the method embodiments.
[0113] The foregoing is merely an example of the present invention and is not intended to limit the present invention. Various modifications and variations are possible for those skilled in the art. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be included within the scope of the claims of the present invention.
Claims
1. A training method for a text classification model, comprising: Determine a text sample and determine a plurality of pre-trained teacher models; wherein the number of parameters of each teacher model is different; According to the order of the parameter amount of each teacher model from small to large, for each teacher model, executing: inputting the text sample into the teacher model to determine the pseudo-label result, and inputting the text sample into the student model to be trained to determine the classification result, and training the student model to be trained at least according to the pseudo-label result obtained based on the teacher model and the classification result; The trained student model is used as a text classification model; wherein the text classification model is used to determine the classification result of the text to be classified based on the text to be classified.
2. The method according to claim 1, wherein the student model to be trained comprises a feature extraction layer and a classification layer; The text sample is input into the student model to be trained to determine the classification result, specifically including: Inputting the text sample into the feature extraction layer of the student model to be trained, and determining the feature sequence corresponding to the text sample; Taking the feature of the position corresponding to the teacher model in the feature sequence as the output feature; The output features are input into the classification layer of the student model to be trained to determine the classification result.
3. The method according to claim 2, training the student model to be trained at least according to the pseudo-label result obtained based on the teacher model and the classification result, specifically comprising: Taking the feature corresponding to the specified position in the feature sequence as the first feature; Inputting the first feature into the classification layer of the student model to be trained to determine a first result; Determining a label corresponding to the text sample; The student model to be trained is trained according to the pseudo-label result obtained based on the teacher model, the classification result, the first result and the annotation.
4. The method according to claim 3, training the student model to be trained according to the pseudo-label result obtained based on the teacher model, the classification result, the first result and the annotation, specifically comprising: Determining a first task loss according to the first result and the annotation; Determining a second task loss according to the pseudo-label result obtained based on the teacher model and the classification result; The student model to be trained is trained according to the first task loss and the second task loss.
5. The method according to claim 2, training the student model to be trained at least according to the pseudo-labeling result obtained based on the teacher model and the classification result, specifically comprising: Determine other teacher models according to the parameter amount of the teacher model; wherein the parameter amount of the other teacher models is smaller than the parameter amount of the teacher model; Taking the feature of the position corresponding to the other teacher model in the feature sequence as the second feature; Inputting the second feature into the classification layer of the student model to be trained to determine a second result; Determine the pseudo-label results corresponding to the other teacher models and use them as other results; The student model to be trained is trained at least based on the pseudo-label result, the classification result, the second result and the other results obtained based on the teacher model.
6. The method according to claim 5, training the student model to be trained based on at least the pseudo-label result, the classification result, the second result and the other results obtained based on the teacher model, specifically comprising: Taking the feature corresponding to the specified position in the feature sequence as the first feature; Inputting the first feature into the classification layer of the student model to be trained to determine a first result; Determining a label corresponding to the text sample; The student model to be trained is trained according to the pseudo-label result obtained based on the teacher model, the classification result, the second result, the other results, the first result and the annotation.
7. The method according to claim 6, training the student model to be trained according to the pseudo-labeled result obtained based on the teacher model, the classification result, the second result, the other results, the first result and the annotation, specifically comprising: Determining a first task loss according to the first result and the annotation; Determining a second task loss according to the pseudo-label result obtained based on the teacher model and the classification result; Determine a third task loss according to the second result and the other results; The student model to be trained is trained according to the first task loss, the second task loss and the third task loss.
8. The method according to claim 7, training the student model to be trained according to the first task loss, the second task loss and the third task loss, specifically comprising: According to the specified weights, weighting the first task loss, the second task loss and the third task loss respectively; The student model to be trained is trained according to the weighted first task loss, the weighted second task loss and the weighted third task loss.
9. The method as claimed in claim 1, wherein a plurality of teacher models are pre-trained, specifically comprising: Determining a label corresponding to the text sample; For each teacher model to be trained, the teacher model to be trained is trained based on the text sample and the annotation.
10. The method according to claim 1, after determining the text sample, the method further comprises: Determining an initial student model and determining a label corresponding to the text sample; Based on the text samples and the annotations, the initial student model is trained to obtain a student model to be trained.
11. The method of claim 1, further comprising: In response to an input operation of a user, determining an input text of the user; Determining pre-stored standard text; Using the standard text and the input text as text to be classified; Inputting the text to be classified into the text classification model to determine the classification result of the text to be classified; When the classification result is similar, the reply text corresponding to the standard text is determined and displayed to the user.
12. A training device for a text classification model, comprising: A first determination module is used to determine a text sample and a plurality of pre-trained teacher models, wherein the number of parameters of each teacher model is different; The first training module is used to execute, for each teacher model in order of the parameter amount of each teacher model from small to large, the following steps: inputting the text sample into the teacher model to determine a pseudo-labeling result, and inputting the text sample into a student model to be trained to determine a classification result, and training the student model to be trained at least according to the pseudo-labeling result obtained based on the teacher model and the classification result; The second determination module is used to use the trained student model as a text classification model; wherein the text classification model is used to determine the classification result of the text to be classified based on the text to be classified.
13. The apparatus of claim 12, wherein the student model to be trained comprises a feature extraction layer and a classification layer; The first training module is specifically used to input the text sample into the feature extraction layer of the student model to be trained, determine the feature sequence corresponding to the text sample; use the feature of the position corresponding to the teacher model in the feature sequence as the output feature; input the output feature into the classification layer of the student model to be trained, and determine the classification result.
14. In the device as described in claim 13, the first training module is specifically used to use the feature corresponding to the specified position in the feature sequence as the second feature; input the second feature into the classification layer of the student model to be trained to determine the second result; determine the annotation corresponding to the text sample; and train the student model to be trained according to the pseudo-label result obtained based on the teacher model, the classification result, the second result and the annotation.
15. In the device as described in claim 14, the first training module is specifically used to determine the first task loss based on the second result and the annotation; determine the second task loss based on the pseudo-label result obtained based on the teacher model and the classification result; and train the student model to be trained based on the first task loss and the second task loss.
16. The apparatus according to claim 13, wherein the first training module is specifically used to determine other teacher models according to the parameter quantity of the teacher model; wherein The parameter amount of the other teacher model is smaller than the parameter amount of the teacher model; the feature of the position corresponding to the other teacher model in the feature sequence is used as the first feature; Inputting the first feature into the classification layer of the student model to be trained to determine a first result; Determine the pseudo-label results corresponding to the other teacher models and use them as other results; The student model to be trained is trained at least according to the pseudo-label result, the classification result, the first result and the other results obtained based on the teacher model.
17. In the device as described in claim 16, the first training module is specifically used to use the feature corresponding to the specified position in the feature sequence as the second feature; input the second feature into the classification layer of the student model to be trained to determine the second result; determine the annotation corresponding to the text sample; and train the student model to be trained according to the pseudo-label result obtained based on the teacher model, the classification result, the first result, the other results, the second result and the annotation.
18. In the device as described in claim 17, the first training module is specifically used to determine the first task loss based on the second result and the annotation; determine the second task loss based on the first result and the other results; determine the third task loss based on the pseudo-label result obtained based on the teacher model and the classification result; train the student model to be trained based on the first task loss, the second task loss and the third task loss.
19. In the device as described in claim 18, the first training module is specifically used to weight the first task loss, the second task loss and the third task loss respectively according to specified weights; and train the student model to be trained according to the weighted first task loss, the weighted second task loss and the weighted third task loss.
20. The apparatus of claim 12, further comprising: A second training module is used to determine the annotation corresponding to the text sample; For each teacher model to be trained, the teacher model to be trained is trained based on the text sample and the annotation.
21. In the device as described in claim 12, the first determination module, after determining the text sample, is also used to determine the initial student model and determine the annotation corresponding to the text sample; based on the text sample and the annotation, the initial student model is trained to obtain the student model to be trained.
22. The apparatus of claim 12, further comprising: An application module, configured to determine the user's input text in response to the user's input operation; Determining pre-stored standard text; The standard text and the input text are used as texts to be classified; the texts to be classified are input into the text classification model to determine the classification results of the texts to be classified; when the classification results are similar, the reply text corresponding to the standard text is determined and displayed to the user.
23. A computer-readable storage medium, wherein the storage medium stores a computer program, wherein the computer program implements the method according to any one of claims 1 to 11 when executed by a processor.
24. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method according to any one of claims 1 to 11 when executing the program.
Citation Information
Patent Citations
Knowledge distillation using deep clustering
CN114626518A
Method and device for training text auditing model
CN114970540A
Semi-supervised video classification method and system based on neighbor consistency and comparative learning
CN115311605A
Pedestrian re-identification model training method and device based on semi-supervised knowledge distillation
CN115546840A
Text classification model training method and device, medium and electronic equipment
CN117786107A
Cited By
Literature classification method and device, electronic equipment and storage medium
CN121833962A
A document classification method and device, electronic equipment and storage medium
CN121833962B