A training method, device and storage medium for an emotion recognition model
By using the first text training sample of pre-labeled emotion labels to train the emotion recognition model, and setting the emotion label for the second text training sample without labels, the problem of excessive manually labeled samples in the training of the emotion recognition model is solved, and efficient model training and resource saving is achieved.
Patent Information
- Application Number
- CN202211022781.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-25
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2042-08-25
AI Technical Summary
In the prior art, training of emotion recognition models requires a large number of manual labeling samples, resulting in high labor costs and low training efficiency.
The first text emotion recognition model is trained by using the first text training sample of the pre-labeled emotion label, and the model is used to set the emotion prediction probability value for the unlabeled second text training sample as a non-manually labeled emotion label, and then the two are combined to train to generate the second text emotion recognition model.
Reduce dependence on manual annotation samples, save human resources and improve model training efficiency.
Smart Images

Figure CN115391531B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of Internet technologies, and in particular, to a method for training an emotion recognition model, an emotion recognition method, a computer-readable storage medium, and a computer program product. Background Art
[0002] In order to accurately recognize the emotions contained in text, it is usually necessary to manually label a large number of training samples to train and generate a model, which will waste a large amount of human costs and reduce the training efficiency. Summary of the Invention
[0003] In view of the above-mentioned prior art, embodiments of the present invention disclose a method for training an emotion recognition model, which can overcome the defect of a large number of manually labeled samples, and achieve the purpose of saving human costs and improving the efficiency of training the model.
[0004] In view of this, a method for training an emotion recognition model proposed in an embodiment of this application includes:
[0005] Training and generating a first text emotion recognition model using a first text training sample, where the first text training sample has a pre-determined emotion label, and the first text emotion recognition model is a deep learning model for recognizing the emotion expressed by text;
[0006] Processing a second text training sample using the first text emotion recognition model to obtain an emotion prediction probability value, and setting a non-manually labeled emotion label for the second text training sample according to the emotion prediction probability value;
[0007] Using the first text training sample and the second text training sample as a text training set, training and generating a second text emotion recognition model using the samples in the text training set, and using the trained second text emotion recognition model as the final deep learning model for recognizing the emotion expressed by text.
[0008] Further,
[0009] The step of training and generating a first text emotion recognition model using the first text training sample includes:
[0010] Fusing corpora with context association relationships to generate an unlabeled first text training sample;
[0011] Manually labeling the emotion label for the first text training sample, and using the first text training sample with the manually labeled emotion label as a third text training sample;
[0012] Reconstruct the third text training sample using a pre - set hint template, so that the reconstructed third text training sample includes the hint template, the hint template corresponds to the manually labeled emotion label, and the hint template is used to hint the emotion expressed by the training sample;
[0013] Fine - tune the first pre - trained model using the reconstructed third text training sample to obtain the first text emotion recognition model, where the first pre - trained model is a model in the underlying state that has been pre - trained in advance.
[0014] Further,
[0015] The hint template includes M hint templates, and the first pre - trained model includes M first pre - trained models. One hint template corresponds to one pre - trained model, and M is a natural number greater than or equal to 1;
[0016] The step of reconstructing the third text training sample using the pre - set hint template includes: reconstructing the third text training sample separately using the M pre - set hint templates to obtain M reconstructed third text training samples;
[0017] The step of fine - tuning the first pre - trained model using the reconstructed third text training sample to obtain the first text emotion recognition model includes: for the M reconstructed third text training samples and the corresponding M first pre - trained models, fine - tune each corresponding first pre - trained model using each reconstructed third text training sample to obtain the corresponding M first text emotion recognition models.
[0018] Further,
[0019] The step of fine - tuning each corresponding first pre - trained model using each reconstructed third text training sample includes:
[0020] Pre - process the reconstructed third text training sample to generate a fourth text training sample;
[0021] Extract features according to the fourth text training sample to determine the sample input features;
[0022] Input the sample input features into the corresponding first pre - trained model to obtain semantic features;
[0023] Determine the feature vector values at the hint template positions according to the semantic features, and determine the N emotion probability values corresponding to the N pre - set emotion labels from the feature vector values at the hint template positions according to the existing language model dictionary, where N is a natural number greater than or equal to 1;
[0024] Input the N emotion probability values and the manually labeled emotion tags as parameters into the cross-entropy loss function to determine the weight parameters of the first pre-trained model;
[0025] Return the weight parameters to the first pre-trained model for fine-tuning.
[0026] Furthermore,
[0027] The step of using the first text emotion recognition model to process the second text training sample to obtain an emotion prediction probability value and setting an unmanually labeled emotion tag for the second text training sample according to the emotion prediction probability value includes:
[0028] Use the pre-set M prompt templates to reconstruct the second text training sample respectively to obtain M reconstructed second text training samples. The second text training sample is an unlabeled sample, and each of the M reconstructed second text training samples includes the corresponding prompt template;
[0029] Preprocess the M reconstructed second text training samples to generate M fifth text training samples;
[0030] Input the M fifth text training samples into the corresponding M first text emotion recognition models respectively, and each of the first text emotion recognition models obtains N emotion probability values corresponding to the N emotion tags;
[0031] Determine one of the N emotion probability values respectively obtained by the M first text emotion recognition models as the emotion prediction probability value;
[0032] Set the emotion tag corresponding to the emotion prediction probability value as the unmanually labeled emotion tag of the second text training sample.
[0033] Furthermore,
[0034] The step of determining one of the N emotion probability values respectively obtained by the M first text emotion recognition models as the emotion prediction probability value includes:
[0035] Take the first emotion tag as the emotion tag to be processed, and take the emotion probability value corresponding to the emotion tag to be processed as the emotion probability value to be processed;
[0036] Average or weighted average the emotion probability values to be processed corresponding to the emotion tags to be processed of the M first text emotion recognition models, and take the calculation result as the average emotion probability value corresponding to the emotion tag to be processed;
[0037] Take the next emotion label as the emotion label to be processed, and return to the step of averaging or weighted-averaging the probability values of the emotions corresponding to the emotion labels to be processed of the M first text emotion recognition models until all the emotion labels are processed;
[0038] Take the largest of the average probability values of the emotions to be processed as the emotion prediction probability value.
[0039] Furthermore,
[0040] The first text training sample is the text obtained by performing speech recognition on the first speech sample;
[0041] and / or
[0042] The second text training sample is the text obtained by performing speech recognition on the second speech sample;
[0043] and / or
[0044] The emotion recognized by the trained second text emotion recognition model for the text is used as the emotion expressed by the corresponding speech.
[0045] The embodiment of the present application also discloses an emotion recognition method, which can overcome the defect of a large number of manually labeled samples, and achieve the purpose of saving labor costs and improving the efficiency of the training model.
[0046] An emotion recognition method proposed by the present application includes:
[0047] Obtain a third text to be recognized;
[0048] Use the trained second text emotion recognition model described in claim 1 to recognize the third text, and obtain an emotion recognition result for the third text.
[0049] Furthermore,
[0050] The third text is the text obtained by performing speech recognition on the third speech information;
[0051] The emotion recognition result for the third text is used as the emotion recognition result for the third speech information.
[0052] The embodiment of the present invention discloses a training device for a text emotion recognition model, which can overcome the defect of a large number of manually labeled samples, and achieve the purpose of saving labor costs and improving the efficiency of the training model.
[0053] A training device for a text emotion recognition model, the device includes: a first training module, a non-manual annotation module, and a second training module;
[0054] The first training module uses the first text training samples with manually labeled emotion tags to train and generate a first text emotion recognition model, which is a deep learning model for recognizing the emotions expressed in text;
[0055] The non-manual annotation module uses the first text emotion recognition model to process the second text training samples, obtains emotion prediction probability values, and sets non-manually labeled emotion tags for the second text training samples according to the emotion prediction probability values;
[0056] The second training module is used to use the first text training samples and the second text training samples as a text training set, and train and generate a second text emotion recognition model using the samples in the text training set. The trained second text emotion recognition model is used as the final deep learning model for recognizing the emotions expressed in text.
[0057] An embodiment of the present invention discloses an electronic device for training a text emotion recognition model. The electronic device includes:
[0058] A processor;
[0059] A memory for storing executable instructions of the processor;
[0060] The processor is configured to read the executable instructions from the memory and execute the instructions to implement the training method of the text emotion recognition model described in any one of the above.
[0061] An embodiment of the present invention also discloses a computer-readable storage medium, on which computer instructions are stored. When the instructions are executed by a processor, the training method of the text emotion recognition model described in any one of the above can be implemented.
[0062] The present invention also discloses a computer program product, including computer instructions, which implement the training method of the text emotion recognition model as described in any one of the above when executed by a processor.
[0063] In summary, the solution of the embodiment of the present application can use only a small number of first text training samples for manual annotation, thereby training a first text emotion recognition model, and then using the first text emotion recognition model to process a large number of unlabeled second text training samples, and setting non-manually labeled emotion tags for the second text training samples, so that a large number of training samples with emotion tags can be obtained with a small amount of manual intervention, and the final second text emotion recognition model can be trained. Therefore, it is not necessary to use the manual annotation method for all the large number of training samples required for training the model, thereby saving human resources and improving the efficiency of model training. Description of the Drawings
[0064] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0065] Figure 1 It is a flowchart of the first embodiment of the method for training a text emotion recognition model implemented by the present application.
[0066] Figure 2 It is a flowchart of the second embodiment of the method for training a text emotion recognition model implemented by the present application.
[0067] Figure 3 It is a flowchart of the method for fine-tuning the corresponding first pre-trained model using each reconstructed third text training sample.
[0068] Figure 4 It is a flowchart of the specific implementation method of step 205 in the second method embodiment.
[0069] Figure 5 It is a logical schematic diagram of the model training process in the actual application scenario.
[0070] Figure 6 It is a schematic structural diagram of the first embodiment of the device for training a text emotion recognition model implemented by the present application.
[0071] Figure 7 It is a schematic structural diagram of an electronic device provided by the present application. Specific implementation manners
[0072] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.
[0073] The terms "first", "second", "third", "fourth", etc. (if any) in the description, claims and above-mentioned drawings of the present invention are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0074] The technical solution of the present invention will be described in detail below with specific embodiments. The following several specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.
[0075] In the embodiment solution of the present application, first, a first text training sample with manually labeled emotion tags is used to train and generate a first text emotion recognition model. Then, the first text emotion recognition model is used to set non-manually labeled emotion tags for a second text training sample. At this time, both the first text training sample and the second text training sample have emotion tags. Then, the first text training sample and the second text training sample are used as a training set to train and generate a second text emotion recognition model, which is the final deep learning model used to recognize the emotion expressed by the text.
[0076] Figure 1 It is a flowchart of the first embodiment of the training method for implementing an emotion recognition model in the present application. As Figure 1 shown, the method includes:
[0077] Step 101: Use a first text training sample to train and generate a first text emotion recognition model. The first text training sample has a pre-determined emotion tag, and the first text emotion recognition model is a deep learning model for recognizing the emotion expressed by the text.
[0078] The first text training samples here have pre-determined emotion labels. For example, the emotion labels can be set in a manual annotation manner, and the purpose is to train the first text emotion recognition model. The first text emotion recognition model is a deep learning model for recognizing the emotions expressed in text, such as the Bidirectional Encoder Representations from Transformer (BERT) model, the Recurrent Neural Network (RNN) model, the Long Short-Term Memory (LSTM) neural network model, etc. To distinguish it from subsequent models, the model generated by training with the first text training samples is called the "first text emotion recognition model" here.
[0079] Step 102: Process the second text training samples using the first text emotion recognition model to obtain emotion prediction probability values, and set non-manually annotated emotion labels for the second text training samples according to the emotion prediction probability values.
[0080] The second text training samples here are unannotated training samples. In this step, the already trained first text emotion recognition model is used to process the second text training samples to obtain the emotion prediction probability values of the second text training samples. In practical applications, multiple emotion labels can be set in advance, such as emotion labels like "angry", "complaining", "dissatisfied", etc. After the first text emotion recognition model processes the second text training samples, the probability value of the second text training sample containing a certain emotion can be predicted. Therefore, the embodiment of this application is called the "emotion prediction probability value". Thereafter, non-manually annotated emotion labels can be set for the second text training samples according to the emotion prediction probability values. Example: Suppose the first text emotion recognition model predicts that the probability values of the second text training sample containing emotions such as "angry", "complaining", "dissatisfied" are 0.7, 0.9, 0.2 respectively. Then it can be considered that the probability of the second text training sample containing the emotion of "complaining" is the highest. Therefore, the emotion label of "complaining" can be set for the second text training sample.
[0081] In addition, it should be noted that the emotion labels of the first text training samples are manually annotated, while for the second text training samples, since the emotion prediction probability values are predicted according to the first text emotion recognition model and no manual intervention is required, the emotion labels of the second text training samples are non-manually annotated.
[0082] Step 103: Use the first text training sample and the second text training sample as a text training set, and train and generate a second text emotion recognition model using the samples in the text training set. The trained second text emotion recognition model is used as the final deep learning model for recognizing the emotion expressed by the text.
[0083] As described above, both the first text training sample and the second text training sample have been labeled with emotion labels, so they can be directly used to train the second text emotion recognition model. The second text emotion recognition model here is also a deep learning model for recognizing the emotion expressed by the text, such as a BERT model, an RNN model, an LSTM model, etc. Different from the purpose of the first text emotion recognition model, the first text emotion recognition model is to predict the emotion prediction probability value of the unlabeled second text training sample, so as to perform non-artificial labeling on the second text training sample, while the second text emotion recognition model is the model obtained after training the solution of this application and can be directly used for emotion prediction of actual texts.
[0084] Applying the solution of the embodiment of this application, since the first text emotion recognition model is trained first, and the first text emotion recognition model is used to process the unlabeled second text training sample to set non-artificial labeled emotion labels for the second text training sample, therefore, it is not necessary to use all artificial labeling methods for a large number of training samples required for training the model, thus saving human resources and improving the efficiency of model training.
[0085] To better illustrate the solution of the embodiment of this application, the following will be described in detail using other preferred embodiments.
[0086] In the second embodiment of the method of this application, it is assumed that both the first text emotion recognition model and the second text emotion recognition model are BERT models. Among them, the first text emotion recognition model is the first BERT model, and the second text emotion recognition model is the second BERT model. The BERT model is a pre-trained language model. During the training process in natural language processing (NLP), the training samples of downstream tasks are relatively few. Usually, the underlying and common partial models in these tasks are trained in advance, and then the respective surface models are trained with the sample data of their respective downstream tasks, which can greatly accelerate the convergence speed. In the embodiment of the method of this application, it is assumed that the first BERT model is trained based on the first pre-trained model, and the second BERT model is trained based on the second pre-trained model.
[0087] The embodiment of this application pre-sets 5 different emotion labels, such as "angry", "complaining", "dissatisfied", "general" and "satisfied", to express different emotions. Of course, in actual applications, other types and numbers of emotion labels can also be set, which are not limited here.
[0088] In addition, for the accuracy of training, the embodiments of the present application can also set a prompt template during training, and the prompt template corresponds to an unknown emotion label. Suppose the embodiments of the present application set six different prompt templates:
[0089] Template 1: True ___. Sample. That is, add "True __." before the sample, and it becomes a single sentence.
[0090] Template 2: Sample. In short, true ___. That is, add "In short, true __." after the sample, and it becomes a single sentence.
[0091] Template 3: Too ___. || Sample. That is, add "Too __." before the sample, and it becomes two single sentences.
[0092] Template 4: Sample || In short, it makes people feel ___. That is, add "In short, it makes people feel __." after the sample, and it becomes two single sentences.
[0093] Template 5: Sample || I'm ___. That is, add "| I'm __." after the sample, and it becomes two single sentences.
[0094] Template 6: Sample || ___. That is, add "__." after the sample, and it becomes two single sentences.
[0095] In practical applications, how to set the template and the number of templates can be set according to the situation, and there is no limitation here.
[0096] Figure 2 It is a flowchart of the second embodiment of the training method for the emotion recognition model implemented in the present application. As Figure 2 shown, the method includes:
[0097] Step 201: Integrate the corpora with context correlation to generate an unannotated first text training sample.
[0098] Considering that it may be difficult for a single-sentence text to express emotions in practical applications, while multiple sentences with context correlation can express emotions, the embodiments of the present application integrate the corpora with context correlation and use them as the unannotated first text training sample. The meaning of "unannotated" here means that the first text training sample has not been labeled with emotion labels.
[0099] In practical applications, there may be context correlation between the previous sentence of the current text or the text of the previous sentence. Therefore, several corpora with context correlation can be integrated. For example, the corpora with context correlation include:
[0100] "Painting like this is really too ugly."
[0101] "@Zhang San, Engineer Zhang, can we tile that area?";
[0102] "I contacted Li Si but couldn't reach him. I guess he was receiving a client and didn't see the message. I'll ask him to call you back later.";
[0103] After fusion, it can become:
[0104] "[CLS]Painting like this looks really ugly[SEP][CLS]@Zhang San, Engineer Zhang, can we tile that area?[SEP][CLS]I contacted Li Si but couldn't reach him. I guess he was receiving a client and didn't see the message. I'll ask him to call you back later[SEP]";
[0105] Among them, "[CLS]" and "[SEP]" are special marker bits. "[CLS]" indicates the start of a sentence, and "[SEP]" indicates the end of a sentence, which are used for separation.
[0106] Step 202: Manually annotate emotion labels for the first text training samples, and use the first text training samples with manually annotated emotion labels as the third text training samples.
[0107] As described above, the embodiments of the present application pre-set 5 different emotion labels. Assuming that the above example is manually annotated with the emotion label of "complaint" here. That is to say, the emotion label corresponding to the training sample of "[CLS]Painting like this looks really ugly[SEP][CLS]@Zhang San, Engineer Zhang, can we tile that area?[SEP][CLS]I contacted Li Si but couldn't reach him. I guess he was receiving a client and didn't see the message. I'll ask him to call you back later[SEP]" is "complaint".
[0108] Step 203: Reconstruct the third text training samples using the pre-set prompt templates, so that the reconstructed third text training samples include the prompt templates. The prompt templates correspond to the manually annotated emotion labels, and the prompt templates are used to prompt the emotions expressed by the training samples.
[0109] Assuming that the above template one is used to reconstruct the third text training samples here, after reconstruction, it is:
[0110] "[CLS]Really [MASK][MASK]. Painting like this looks really ugly[SEP][CLS]@Zhang San, Engineer Zhang, can we tile that area?[SEP][CLS]I contacted Li Si but couldn't reach him. I guess he was receiving a client and didn't see the message. I'll ask him to call you back later[SEP]".
[0111] Among them, "[MASK]" represents the blank marker bit in the template.
[0112] That is to say, the above-mentioned Template 1 to Template 6 can correspond to the manually labeled emotion tags, which are used to prompt the emotions expressed by the training samples.
[0113] Step 204: Fine-tune the first pre-trained model with the reconstructed third text training samples to obtain a first text emotion recognition model. The first pre-trained model is a model in the underlying state that has been pre-trained in advance.
[0114] The pre-trained model only reflects the model in the underlying state and cannot be directly used to complete the task of predicting text emotions. Based on the first pre-trained model, in this step, fine-tuning is performed using the reconstructed third text training samples to obtain a trained first text emotion recognition model. The above steps only list an example of a training sample. In actual applications, multiple third text training samples can be used for training. The so-called "fine-tuning" here refers to the process of training a model (the first text emotion recognition model) for a specific downstream task.
[0115] The above steps 201 to 204 are the steps of generating a first text emotion recognition model using the first text training samples, that is, the specific implementation method of step 101 in Method Embodiment 1.
[0116] In actual applications, it can be set to M hint templates and M first pre-trained models. One hint template corresponds to one pre-trained model, and M is a natural number greater than or equal to 1. That is to say, for the M reconstructed third text training samples and the corresponding M first pre-trained models, each reconstructed third text training sample is used to fine-tune the corresponding first pre-trained model to obtain the corresponding M first text emotion recognition models.
[0117] In the method embodiment of the present application, M = 6, that is, 6 hint templates and 6 corresponding first pre-trained models are set. Suppose the 6 corresponding first pre-trained models are First Pre-trained Model A, First Pre-trained Model B, First Pre-trained Model C, First Pre-trained Model D, First Pre-trained Model E, and First Pre-trained Model F respectively. A certain first training sample is reconstructed into different third text training samples using 6 hint templates, which are Third Text Training Sample a, Third Text Training Sample b, Third Text Training Sample a, Third Text Training Sample b, Third Text Training Sample c, Third Text Training Sample d, Third Text Training Sample e, and Third Text Training Sample f respectively. Among them, Third Text Training Sample a is input to First Pre-trained Model A, Third Text Training Sample b is input to First Pre-trained Model B, and so on. After the 6 pre-trained models are fine-tuned, 6 different first text emotion recognition models will be obtained.
[0118] In practical applications, the method of fine-tuning the corresponding first pre-training model using each reconstructed third text training sample is as follows: Figure 3 As shown, including:
[0119] Step 301: preprocess the reconstructed third text training sample to generate a fourth text training sample.
[0120] In practical applications, the training samples also need to be preprocessed, including: splitting the reconstructed third text training sample into the smallest character unit in the natural language model, and encoding each smallest character unit using the language model dictionary. For example, in the above example, "True" is a smallest character unit, and its corresponding code is "4696". Therefore, the above example will become: "101, 4696, 103, 103... 102" These codes constitute the fourth text training sample.
[0121] Step 302: extract features based on the fourth text training sample to determine sample input features.
[0122] In actual applications, the input features of the first pre-trained model are of fixed length, including input_ids and attention_mask. Among them, input_ids is the encoded training sample, and attention_mask is a sequence of 0 and 1. The position corresponding to 1 represents the valid part of the sample, and 0 represents a fixed-length supplementary bit.
[0123] Step 303: Input the sample input features into the corresponding first pre-training model to obtain semantic features.
[0124] The BERT model can be expressed as follows:
[0125] H = BERT(X) Formula 1
[0126] Among them, X represents the input features and H represents the semantic features.
[0127] Step 304: Determine the feature vector value at the prompt template position according to the semantic features, and determine N emotion probability values corresponding to N pre-set emotion tags from the feature vector value at the prompt template position according to the existing language model dictionary, where N is a natural number greater than 1 or equal to 1.
[0128] In practical applications, the model can calculate the semantic features of each minimum character unit of the input sample, which represents the probability value encoded at that position. The feature vector value obtained from the position of the prompt template [MASK] also represents the probability value logits encoded at the [MASK] position. [MASK] . Probability value logits [MASK]is a vector with the length of the language model dictionary length. Therefore, according to the positions of the N emotion labels set in advance in the language model dictionary, the N emotion probability values corresponding to the N emotion labels can be determined. In the method embodiment of the present application, 5 emotion labels are set in advance. Therefore, in this step 304, the emotion probability values corresponding to the 5 emotion labels of "angry", "complaining", "dissatisfied", "general", and "satisfied" can be determined respectively.
[0129] Step 305: Input the N emotion probability values and the manually labeled emotion labels as parameters into the cross-entropy loss function to determine the weight parameters of the first pre-trained model.
[0130] Although the N (N = 5) emotion probability values are determined here, in the above step 202, it has been manually labeled as "complaining". Therefore, there is a loss between the calculation result and the actual result. Then, the weight parameters of the first pre-trained model can be calculated through the cross-entropy loss function.
[0131] Step 306: Return the weight parameters to the first pre-trained model for fine-tuning.
[0132] In this step, the calculated weight parameters of the first text emotion recognition model are returned to the first pre-trained model, and the first pre-trained model will modify the weight parameters again. Therefore, the purpose of fine-tuning can be achieved. Of course, steps 301 to 306 are an example of a training sample. In practical applications, multiple samples can be used to fine-tune the first pre-trained model. Steps 301 to 306 can be repeatedly executed until the training is completed. At this time, the fine-tuned first pre-trained model is the first text emotion recognition model.
[0133] Figure 3 The method shown is a method of fine-tuning the corresponding first pre-trained model for each reconstructed third text training sample. In the method embodiment of the present application, there are M (M = 6) reconstructed third text training samples, and M (M = 6) corresponding first text emotion recognition models can be generated. Using the third text training samples reconstructed by one hint template to generate the first text emotion recognition model can predict the emotion prediction probability value of the second text training sample in the subsequent step 205. However, for more accurate prediction, different hint templates are used here. In practical applications, other types and quantities of hint templates can also be set, which will not be listed one by one here.
[0134] Step 205: Process the second text training sample using the first text emotion recognition model to obtain the emotion prediction probability value, and set the non-manually labeled emotion label for the second text training sample according to the emotion prediction probability value.
[0135] This step is the same as 102 in the first method embodiment.
[0136] Still assuming that there are M (M = 6) prompt templates, the first pre-trained model is M (M = 6) first pre-trained models, and M (M = 6) first text emotion recognition models are trained and generated according to M (M = 6) reconstructed third text training samples. In this case, the specific method of step 205 is as Figure 4 shown, including:
[0137] Step 401: Use the M pre-set prompt templates to reconstruct the second text training samples respectively, obtaining M reconstructed second text training samples. The second text training samples are unlabeled samples, and each of the M reconstructed second text training samples includes the corresponding prompt template.
[0138] The reconstruction method described in this step is the same as the reconstruction in step 203, both using the set prompt templates for reconstruction. In addition, different from the first text training samples, the second text training samples are unlabeled samples and do not require manual annotation, but are input to the first text emotion recognition model to automatically obtain emotion labels.
[0139] Step 402: Preprocess the M reconstructed second text training samples to generate M fifth text training samples.
[0140] The preprocessing described in this step is similar to the preprocessing described in step 301, including steps such as splitting the reconstructed second text training samples into the smallest character units in the natural language model and encoding each smallest character unit using the language model dictionary.
[0141] Step 403: Input the M fifth text training samples into the corresponding M first text emotion recognition models respectively, and each first text emotion recognition model obtains N emotion probability values corresponding to N emotion labels.
[0142] When inputting the fifth text training samples into the first text emotion recognition model in this step, it is usually also necessary to extract features to determine the input features, and then input the input features into the corresponding first text emotion recognition model. Since each first text emotion recognition model is a model trained through the manually annotated first text training samples and can be directly used for emotion prediction, N emotion probability values corresponding to N emotion labels can be obtained.
[0143] Thereafter, in the embodiments of the present application, the N emotion probability values respectively obtained by the M first text emotion recognition models can be subsequently determined, and one of them is determined as the emotion prediction probability value. Since each first text emotion recognition model can obtain N emotion probability values, then the M first text emotion recognition models can obtain M * N emotion probability values, and the method for determining the emotion prediction probability value from them is as shown in steps 404 to 407.
[0144] Step 404: Take the first emotion label as the emotion label to be processed, and take the emotion probability value corresponding to the emotion label to be processed as the emotion probability value to be processed.
[0145] Step 405: Average or weighted-average the emotion probability values to be processed corresponding to the emotion labels to be processed of the M first text emotion recognition models, and take the calculation result as the average emotion probability corresponding to the emotion label to be processed.
[0146] Suppose there are N (N = 5) emotion labels, the first emotion label is "angry", and M emotion probability values corresponding to "angry" can be obtained from M first text emotion recognition models respectively. Average or weighted-average these M emotion probability values. For example, if the M (M = 6) emotion probability values are 0.3, 0.4, 0.25, 0.35, 0.5, 0.35 respectively, then their average value is 0.36. According to this method, suppose the emotion probability values corresponding to M "complaint" emotion labels are 0.8, 0.9, 0.75, 0.85, 0.95, 0.82 respectively, and their average value is 0.85. And so on. By repeatedly executing steps 502 to 503, N (N = 5) average emotion probabilities can be obtained.
[0147] In practical applications, the method of averaging or weighted-averaging the emotion probability values to be processed can be calculated using the following formula 2:
[0148]
[0149] Among them, P represents the set of all models, p represents one specific model, w(p) represents the weight of each model, Z represents the sum of the weights of all models, and s p (l|x) represents the emotion probability value to be processed of each model for the input sample x. When w(p) = 1, it means that the weights of the emotion probability values to be processed corresponding to each model are the same. When w(p) ≠ 1, it means that the weights of the emotion probability values to be processed corresponding to each model are different, and calculations are performed according to their respective weights.
[0150] Step 406: Take the next emotion label as the label to be processed, and return to step 405 until all emotion labels are processed.
[0151] Step 407: Take the largest of the average emotion probabilities to be processed as the emotion prediction probability value.
[0152] In the above example, it is assumed that the average values of the N (N = 5) emotion probabilities calculated are as follows: the average emotion probability corresponding to the emotion label "anger" is 0.36, the average emotion probability corresponding to the emotion label "complaint" is 0.85, the average emotion probability corresponding to the emotion label "dissatisfaction" is 0.25, the average emotion probability corresponding to the emotion label "general" is 0.23, and the average emotion probability corresponding to the emotion label "satisfaction" is 0.2. Then, the average emotion probability corresponding to the emotion label "complaint" is the largest, and it is used as the emotion prediction probability value.
[0153] In the above steps 404 to 407, an emotion prediction probability value can be determined from the N emotion probability values respectively obtained by the M first text emotion recognition models.
[0154] Step 408: Set the emotion label corresponding to the emotion prediction probability value as the non-artificially labeled emotion label of the second text training sample.
[0155] Assume that the average emotion probability corresponding to the emotion label "complaint" is the largest and it is used as the emotion prediction probability value. Then, in this step, the emotion label "complaint" can be set as the non-artificially labeled emotion label of the second text training sample. So far, the above steps 401 to 408 have processed the second text training sample using the first text emotion recognition model, obtained the emotion prediction probability value, and set the non-artificially labeled emotion label for the second text training sample according to the emotion prediction probability value. In practical applications, the above steps 401 to 408 can be used to automatically set non-artificially labeled emotion labels for a large number of unlabeled second text training samples.
[0156] Step 206: Use the first text training sample and the second text training sample as a text training set, and train and generate a second text emotion recognition model using the samples in the text training set. The trained second text emotion recognition model is used as the final deep learning model for recognizing the emotion expressed by the text.
[0157] Here, the first text training sample has an artificially labeled emotion label, and the second text training sample also already has a non-artificially labeled emotion label. Therefore, they can be jointly used as the training set for training the second text emotion recognition model. In practical applications, the second text emotion recognition model can be trained and generated based on the second pre-trained model. During specific training, the samples in the training set can be preprocessed, then feature extraction can be performed to determine the input features, and then the input features are input into the second pre-trained model to obtain the output semantic features. Finally, the loss between the emotion prediction probability value and the emotion label corresponding to the sample is determined, and then the weight parameters of the second pre-trained model are adjusted until the training is completed to generate the second text emotion recognition model.
[0158] By applying the solution of the embodiment of the present application, the first text emotion recognition model can be trained with the first text training samples with a small number of manually labeled emotion tags, and then the first text emotion recognition model can be used to automatically set non-manually labeled emotion tags for a large number of unlabeled second text training samples. Finally, the second text emotion recognition model is trained and generated using the first text training samples and the second text training samples, thereby saving a large amount of human resources and improving the efficiency of model training.
[0159] In the first embodiment of the training method of the emotion recognition model and the second embodiment of the training method of the emotion recognition model, the first text training samples and / or the second text training samples can originally be texts, or can be texts obtained after speech recognition. For example: performing speech recognition on the first speech sample to obtain a text as the first text training sample, and / or performing speech recognition on the second speech sample to obtain a text as the second text training sample. In this case, the emotion expressed by the text recognized by the trained second text emotion recognition model can be used as the emotion expressed by the corresponding speech.
[0160] The above are the embodiments described in the training method of the emotion recognition model. In practical applications, the trained second text emotion recognition model can be further used to directly recognize emotions. Specifically, the third text to be recognized can be obtained first, and then the trained second text emotion recognition model is used to recognize the third text to obtain an emotion recognition result for the third text. The third text here can be a text directly obtained from, for example, a certain chat platform, or a text obtained by speech recognition in a chat platform. Then, when using the trained second text emotion recognition model to recognize the third text to obtain an emotion recognition result, this emotion recognition result can also be used as the emotion recognition result of the corresponding speech.
[0161] In an actual application scenario, assume that both the first text emotion recognition model and the second text emotion recognition model are BERT models. Among them, the first text emotion recognition model is the first BERT model, and the second text emotion recognition model is the second BERT model. Assume Figure 5 is a logical schematic diagram of the model training process in this scenario. As Figure 5As shown in the figure, the process includes: generating a first text training sample by fusing corpus context, manually annotating an emotion label for the first text training sample to obtain a third text training sample (Steps 201 to Step 202); reconstructing the third text training sample using six preset prompt templates (Step 203); respectively fine-tuning six first pre-trained models using the reconstructed third text training sample to obtain six first BERT models (Step 204); respectively processing the unannotated second text training sample using the six first BERT models to set an unmanually annotated emotion label for the second text training sample (Step 205); using the first text training sample and the second text training sample as a training set to fine-tune a second pre-trained model to obtain a second BERT model (Step 206).
[0162] After that, the generated second BERT model can be directly used in actual applications. For example, in a certain decoration project, customers, construction workers, and managers form a chat platform. If a customer mentions a text like "Why hasn't this wall been leveled yet?" in the chat, after the second BERT model captures this text, it can accurately analyze that this text contains a "dissatisfied" emotion and report it to the management platform. Managers will handle it properly according to the reported situation later, so as to provide high-quality services for customers.
[0163] The embodiment of the present application also provides a training device for an emotion recognition model. Figure 6 It is a schematic structural diagram of the first embodiment of the training device for the emotion recognition model. As Figure 6 shown, the device includes a first training module 601, an unmanually annotating module 602, and a second training module 603. Among them:
[0164] The first training module 601 uses the first text training sample to train and generate a first text emotion recognition model. The first text training sample has a manually annotated emotion label, and the first text emotion recognition model is a deep learning model for recognizing the emotion expressed by the text.
[0165] The unmanually annotating module 602 processes the second text training sample using the first text emotion recognition model to obtain an emotion prediction probability value, and sets an unmanually annotated emotion label for the second text training sample according to the emotion prediction probability value.
[0166] The second training module 603 is used to use the first text training sample and the second text training sample as a text training set, and train and generate a second text emotion recognition model using the samples in the text training set. The trained second text emotion recognition model is used as the final deep learning model for recognizing the emotion expressed by the text.
[0167] That is to say, the first training module 601 uses the first text training samples to train and generate the first text emotion recognition model; the non-artificial annotation module 602 processes the second text training samples by using the first text emotion recognition model to obtain emotion prediction probability values, and sets non-artificial annotation emotion labels for the second text training samples according to the emotion prediction probability values; the second training module 603 uses the first text training samples and the second text training samples as a text training set, and trains and generates the second text emotion recognition model by using the samples in the text training set.
[0168] Applying the solution of the embodiment of the present application, since the first text emotion recognition model is trained first, and the first text emotion recognition model is used to process the unlabeled second text training samples to set non-artificial annotation emotion labels for the second text training samples, therefore, a large number of training samples required for training the model do not have to be all manually annotated, thus saving human resources and improving the efficiency of model training.
[0169] In practical applications, the first training module 601 can further specifically refer to steps 201 to 204 in the second method embodiment above to complete the steps of generating the first text emotion recognition model by using the first text training samples; the non-artificial annotation module 602 can further specifically refer to steps 401 to 408 in the second method embodiment above to complete the steps of setting non-artificial annotation emotion labels for the second text training samples.
[0170] The embodiment of the present application further provides a computer-readable medium. The computer-readable storage medium stores instructions, and the instructions can execute the steps in the training method of the text emotion recognition model as described above when executed by a processor. In practical applications, the computer-readable medium may be included in the device / device / system described in the above embodiment, or may exist separately without being assembled into the device / device / system. The above computer-readable storage medium carries one or more programs, and when the above one or more programs are executed, the methods of data visualization display described in the above embodiments can be implemented. According to the embodiments disclosed in the present application, the computer-readable storage medium may be a non-volatile computer-readable storage medium, for example, it may include but is not limited to: portable computer disks, hard disks, random access memories (RAMs), read-only memories (ROMs), erasable programmable read-only memories (EPROMs or flash memories), portable compact disk read-only memories (CD-ROMs), optical storage devices, magnetic storage devices, or any suitable combination of the above, but is not used to limit the scope of protection of the present application. In the embodiments disclosed in the present application, the computer-readable storage medium may be any tangible medium that contains or stores a program, and the program can be used by or combined with an instruction execution system, device, or device.
[0171] Such asFigure 7 As shown, an embodiment of the present invention further provides an electronic device, which may integrate the device for implementing the method of the present application. As Figure 7 shown, it shows a schematic structural diagram of the electronic device involved in the embodiment of the present invention. Specifically:
[0172] The electronic device may include a processor 701 with one or more processing cores, a memory 702 of one or more computer-readable storage media, and a computer program stored in the memory and executable on the processor. When executing the program in the memory 702, the training method of the above text emotion recognition model can be implemented.
[0173] Specifically, in practical applications, the electronic device may further include components such as a power supply 703, an input unit 704, and an output unit 705. Those skilled in the art can understand that Figure 7 the structure of the electronic device shown in
[0174] does not limit the electronic device, and it may include more or fewer components than shown, or combine certain components, or have different component arrangements. Among them:
[0175] The processor 701 is the control center of the electronic device, connecting various parts of the entire electronic device through various interfaces and lines, and by running or executing software programs and / or modules stored in the memory 702, and calling data stored in the memory 702, it executes various functions of the server and processes data, thereby monitoring the entire electronic device.
[0176] The electronic device further includes a power supply 703 for supplying power to each component, which can be logically connected to the processor 701 through a power management system, so as to implement functions such as management of charging, discharging, and power consumption management through the power management system. The power supply 703 may further include any components such as one or more DC or AC power supplies, a recharge system, a power failure detection circuit, a power converter or inverter, and a power status indicator.
[0177] The electronic device may further include an input unit 704, which may be configured to receive input numerical or character information, and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function controls.
[0178] The electronic device may further include an output unit 705, which may be configured to display information input by the user or information provided to the user, as well as various graphical user interfaces, which may be composed of graphics, text, icons, videos, and any combination thereof.
[0179] The embodiment of the present application further provides a computer program product, which includes computer instructions that, when executed by a processor, implement the method described in any of the above embodiments.
[0180] The flowcharts and block diagrams in the accompanying drawings of the present application illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments disclosed in the present application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the above-mentioned module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in the order marked in different drawings. For example, two consecutively represented blocks may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, as well as combinations of blocks in the block diagram or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.
[0181] Those skilled in the art can understand that the features described in the various embodiments and / or claims of the present disclosure can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly described in the present application. In particular, without departing from the spirit and teachings of the present application, the features described in the various embodiments and / or claims of the present application can be combined and / or combined in various ways, and all such combinations and / or combinations fall within the scope of the disclosure of the present application.
[0182] In this text, specific embodiments are used to elaborate on the principles and implementation manners of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention, and is not used to limit this application. For those skilled in the art, they can make changes in the specific implementation manners and application scope according to the idea, spirit and principle of the present invention. Any modifications, equivalent replacements, improvements, etc. made by them shall be included within the scope protected by this application.
Claims
1. A training method for an emotion recognition model, characterized in that, The method includes: Training and generating a first text emotion recognition model using a first text training sample, which includes: fusing corpora with context association relationships to generate an unlabeled first text training sample; manually labeling emotion labels for the first text training sample, and using the first text training sample with manually labeled emotion labels as a third text training sample; reconstructing the third text training sample using a pre-set prompt template, such that the reconstructed third text training sample includes the prompt template, the prompt template corresponding to the manually labeled emotion label, the prompt template being used to prompt the emotion expressed by the training sample; and fine-tuning a first pre-trained model using the reconstructed third text training sample to obtain the first text emotion recognition model, the first pre-trained model being a pre-trained model in an underlying state, the first text training sample having a pre-determined emotion label, the first text emotion recognition model being a deep learning model for recognizing the emotion expressed by text; Processing a second text training sample using the first text emotion recognition model to obtain an emotion prediction probability value, and setting an unmanually labeled emotion label for the second text training sample according to the emotion prediction probability value; Using the first text training sample and the second text training sample as a text training set, training and generating a second text emotion recognition model using the samples in the text training set, and using the trained second text emotion recognition model as the final deep learning model for recognizing the emotion expressed by text, wherein, the prompt template includes M prompt templates, the first pre-trained model includes M first pre-trained models, one prompt template corresponds to one pre-trained model, and M is a natural number greater than or equal to 1; the step of reconstructing the third text training sample using the pre-set prompt template includes: respectively reconstructing the third text training sample using the pre-set M prompt templates to obtain M reconstructed third text training samples; the step of fine-tuning the first pre-trained model using the reconstructed third text training sample to obtain the first text emotion recognition model includes: for the M reconstructed third text training samples and the corresponding M first pre-trained models, fine-tuning each corresponding first pre-trained model using each reconstructed third text training sample to obtain M corresponding first text emotion recognition models.
2. The method according to claim 1, wherein The step of fine-tuning each corresponding first pre-trained model using each reconstructed third text training sample includes: Preprocessing the reconstructed third text training sample to generate a fourth text training sample; Extracting features according to the fourth text training sample to determine sample input features; Inputting the sample input features into the corresponding first pre-trained model to obtain semantic features; Determine the eigenvector value at the position of the prompt template according to the semantic feature, and determine the N emotion probability values corresponding to N preset emotion labels from the eigenvector value at the position of the prompt template according to the existing language model dictionary, where N is a natural number greater than 1 or equal to 1; Input the N emotion probability values and the manually labeled emotion label as parameters into the cross-entropy loss function to determine the weight parameters of the first pre-trained model; Return the weight parameters to the first pre-trained model for fine-tuning.
3. The method according to claim 2, wherein The step of using the first text emotion recognition model to process the second text training sample to obtain an emotion prediction probability value and setting an unmanually labeled emotion label for the second text training sample according to the emotion prediction probability value includes: Use the M preset prompt templates to reconstruct the second text training sample respectively to obtain M reconstructed second text training samples. The second text training sample is an unlabeled sample, and each of the M reconstructed second text training samples includes the corresponding prompt template; Preprocess the M reconstructed second text training samples to generate M fifth text training samples; Input the M fifth text training samples into the corresponding M first text emotion recognition models respectively, and each of the first text emotion recognition models obtains the N emotion probability values corresponding to the N emotion labels; Determine one of the N emotion probability values respectively obtained by the M first text emotion recognition models as the emotion prediction probability value; Set the emotion label corresponding to the emotion prediction probability value as the unmanually labeled emotion label of the second text training sample.
4. The method according to claim 3, characterized in that The step of determining one of the N emotion probability values respectively obtained by the M first text emotion recognition models as the emotion prediction probability value includes: Take the first emotion label as the emotion label to be processed, and take the emotion probability value corresponding to the emotion label to be processed as the emotion probability value to be processed; Average or weighted average the emotion probability values to be processed corresponding to the emotion label to be processed of the M first text emotion recognition models, and take the calculation result as the average emotion probability value corresponding to the emotion label to be processed; Take the next emotion label as the emotion label to be processed, and return to the step of averaging or weighted averaging the emotion probability values to be processed corresponding to the emotion label to be processed of the M first text emotion recognition models until all the emotion labels are processed; Take the largest of the average emotion probability values to be processed as the emotion prediction probability value.
5. The method according to any one of claims 1 to 4, wherein The first text training sample is the text obtained by performing speech recognition on the first speech sample; and / or The second text training sample is the text obtained by performing speech recognition on the second speech sample; and / or The trained second text emotion recognition model recognizes the emotion expressed by the text as the emotion expressed by the corresponding speech.
6. A method for emotion recognition, characterized in that, The method includes: Obtain a third text to be recognized; The trained second text emotion recognition model according to claim 1 is used to recognize the third text, and an emotion recognition result for the third text is obtained.
7. A computer-readable storage medium having computer instructions stored thereon, characterized in that, When the instruction is executed by a processor, the emotion recognition model training method according to any one of claims 1 to 5 can be implemented.
8. A computer program product, comprising computer instructions, where when the computer instructions are executed by a processor, the text emotion recognition model training method according to any one of the above claims 1 to 5 is implemented.
Citation Information
Patent Citations
Language model training method and device, electronic equipment and storage medium
CN114036300A
Text processing method and device, computer equipment and storage medium
CN114091444A