Training Method, Device, Equipment and Storage Medium for Emotion Recognition Model

By determining the loss adjustment weight and emotion recognition loss during the training process of the emotion recognition model, the data imbalance problem is solved, and the training effect and accuracy of the emotion recognition model are improved.

CN118245602BActive Publication Date: 2025-06-20TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202410339691.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-03-22
Publication Date
2025-06-20
Estimated Expiration
2044-03-22

AI Technical Summary

Technical Problem

The prior art is difficult to effectively deal with the problem of data imbalance in emotion recognition, resulting in a decrease in generalization of head labels and insufficient label learning benefits, thereby reducing the accuracy of emotion recognition.

Method used

During the training process of the emotion recognition model, the loss adjustment weight is determined based on the sample characteristics and emotion recognition results of the sample text, and the emotion recognition loss is determined based on the loss adjustment weight, emotion recognition results and emotion truth labels. Based on the loss training model, the training effect of the emotion recognition model is optimized.

Benefits of technology

The adjustment of emotion recognition loss in easy samples and difficult samples is achieved, the performance of the emotion recognition model on unbalanced label samples is improved, and the training effect and output accuracy of the emotion recognition model are optimized.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118245602B_ABST
    Figure CN118245602B_ABST
Patent Text Reader

Abstract

The embodiments of the present application disclose a training method, device, equipment and storage medium for an emotion recognition model, belonging to the technical field of emotion recognition. The method includes: performing emotion recognition on the sample texts in the training dataset through the emotion recognition model to obtain the sample emotion recognition results corresponding to each sample text, where the training dataset includes sample texts and the emotion true value labels corresponding to the sample texts; determining the loss adjustment weight corresponding to the sample text according to the sample features of the sample text and the sample emotion recognition result, where the sample features include at least one of the sample type and the number of label samples of the emotion true value label corresponding to the sample text; determining the emotion recognition loss of the sample text based on the loss adjustment weight, the sample emotion recognition result and the emotion true value label; training the emotion recognition model based on the emotion recognition loss; adopting the solution provided by the embodiments of the present application can optimize the training effect of the emotion recognition model and improve the accuracy of emotion recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the technical field of emotion recognition, and particularly to a training method, device, equipment and storage medium for an emotion recognition model. Background Art

[0002] The understanding of film and television drama scripts requires the emotional analysis of the scripts submitted by writers or to be filmed, so as to understand the emotional trends of the script characters, especially the leading male and female characters, so as to evaluate whether the script is full of ups and downs or grasp the emotional points during filming. Since human emotions are complex, a single sentence can express multiple emotions, such as extreme joy turning to sorrow, or crying with joy. And in multi-label prediction, it is often easy to miss predicting some labels.

[0003] In the related art, the method of balanced sampling is adopted. By counting the number of samples n of each label category, and then when sampling the samples, the samples of this label category are drawn with a probability of 1 / n, so that a larger sampling coverage can be obtained for the label categories with less data volume.

[0004] In the case of extremely unbalanced data, some samples of the head labels (label categories with more sample numbers) will always not be sampled and learned, resulting in a reduction in the generalization of the head labels, and samples with multiple label categories will also bring the problems of sampling redundancy and insufficient label learning benefits, thus reducing the accuracy of emotion recognition. Summary of the Invention

[0005] The embodiments of the present application provide a training method, device, equipment and storage medium for an emotion recognition model, which can optimize the training effect of the emotion recognition model and improve the accuracy of emotion recognition. The technical solutions are as follows:

[0006] On the one hand, the embodiments of the present application provide a training method for an emotion recognition model, and the method includes:

[0007] Performing emotion recognition on the sample texts in the training dataset through the emotion recognition model to obtain the sample emotion recognition results corresponding to each sample text, where the training dataset includes the sample texts and the emotion true value labels corresponding to the sample texts;

[0008] Determining the loss adjustment weight corresponding to the sample text according to the sample features of the sample text and the sample emotion recognition result, where the sample features include at least one of the sample type and the number of label samples of the emotion true value label corresponding to the sample text, the sample type includes easy samples and difficult samples, and the number of label samples refers to the number of sample texts with the emotion true value label in the training dataset;

[0009] Determine the emotion recognition loss of the sample text based on the loss adjustment weight, the sample emotion recognition result, and the emotion ground truth label;

[0010] Train the emotion recognition model based on the emotion recognition loss.

[0011] On the other hand, an embodiment of the present application provides a training device for an emotion recognition model, and the device includes:

[0012] A first emotion recognition module, configured to perform emotion recognition on the sample text in the training dataset through an emotion recognition model, and obtain the sample emotion recognition result corresponding to each sample text. The training dataset includes the sample text and the emotion ground truth label corresponding to the sample text;

[0013] A weight determination module, configured to determine the loss adjustment weight corresponding to the sample text according to the sample features of the sample text and the sample emotion recognition result. The sample features include at least one of the sample type and the number of label samples of the emotion ground truth label corresponding to the sample text. The sample type includes easy samples and difficult samples, and the number of label samples refers to the number of sample texts with the emotion ground truth label in the training dataset;

[0014] A loss determination module, configured to determine the emotion recognition loss of the sample text based on the loss adjustment weight, the sample emotion recognition result, and the emotion ground truth label;

[0015] A model training module, configured to train the emotion recognition model based on the emotion recognition loss.

[0016] On the other hand, an embodiment of the present application provides a computer device, which includes a processor and a memory. At least one instruction is stored in the memory, and the at least one instruction is loaded and executed by the processor to implement the training method of the emotion recognition model as described in the above aspect.

[0017] On the other hand, an embodiment of the present application provides a computer-readable storage medium, in which at least one instruction is stored, and the at least one instruction is loaded and executed by a processor to implement the training method of the emotion recognition model as described in the above aspect.

[0018] On the other hand, an embodiment of the present application provides a computer program product, which includes at least one instruction, and the at least one instruction is stored in a computer-readable storage medium. The processor of the computer device reads the at least one instruction from the computer-readable storage medium, and the processor executes the at least one instruction, so that the computer device executes the training method of the emotion recognition model as described in the above aspect.

[0019] In the embodiments of the present application, after the emotion recognition model performs emotion recognition on the sample texts in the training dataset to obtain the sample emotion recognition results of each sample text, the emotion recognition loss is not directly determined based on the sample emotion recognition results and the emotion true value labels. Instead, the sample features of each sample text in the training dataset are fully considered. First, the loss adjustment weight is determined according to the sample features of the sample text and the sample emotion recognition results, and then the emotion recognition loss is jointly determined by combining the loss adjustment weight, the sample emotion recognition results, and the emotion true value labels. Furthermore, the emotion recognition model is trained based on the emotion recognition loss. By adopting the solution provided by the embodiments of the present application, the emotion recognition loss of easy samples and difficult samples can be adjusted by using the loss adjustment weight, and the emotion recognition loss of unbalanced label samples can be adjusted when the number of sample texts with different emotion true value labels is unbalanced. Therefore, by training the emotion recognition model based on the adjusted emotion recognition loss, the training effect of the emotion recognition model can be optimized, and the accuracy of the output of the emotion recognition model can be improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0021] Figure 1 Shows a flowchart of a method for training an emotion recognition model provided by an exemplary embodiment of the present application;

[0022] Figure 2 Shows a schematic structural diagram of a BERT model provided by an exemplary embodiment of the present application;

[0023] Figure 3 Shows a schematic structural diagram of a Transformer encoder provided by an exemplary embodiment of the present application;

[0024] Figure 4 Shows a schematic diagram of the input data format of an emotion recognition model provided by an exemplary embodiment of the present application;

[0025] Figure 5 Shows a schematic diagram of the distribution of the number of sample texts corresponding to different emotion labels provided by an exemplary embodiment of the present application;

[0026] Figure 6 Shows a schematic diagram of the staged training of an emotion recognition model provided by an exemplary embodiment of the present application;

[0027] Figure 7 The flowchart of the training method of the emotion recognition model provided by another exemplary embodiment of the present application is shown;

[0028] Figure 8 The schematic diagram of determining the second loss adjustment weight in the multi-label sample text provided by an exemplary embodiment of the present application is shown;

[0029] Figure 9 The schematic structural diagram of the emotion recognition model including two networks provided by an exemplary embodiment of the present application is shown;

[0030] Figure 10 The flowchart of emotion recognition by applying the emotion recognition model provided by an exemplary embodiment of the present application is shown;

[0031] Figure 11 The schematic diagram of the emotion development curves corresponding to two key figures provided by an exemplary embodiment of the present application is shown;

[0032] Figure 12 The structural block diagram of the training device of the emotion recognition model provided by an exemplary embodiment of the present application is shown;

[0033] Figure 13 The schematic structural diagram of the computer device provided by an exemplary embodiment of the present application is shown. Detailed implementation manners

[0034] To make the objectives, technical solutions and advantages of the present application clearer, the embodiments of the present application will be further described in detail below with reference to the accompanying drawings.

[0035] Artificial Intelligence (AI) is to use a digital computer or a machine controlled by a digital computer to simulate, extend and expand human intelligence, a theory, method, technology and application system that perceives the environment, acquires knowledge and uses knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology of computer science, which attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable the machines to have the functions of perception, reasoning and decision-making.

[0036] Artificial intelligence technology is an interdisciplinary subject with a wide range of fields involved, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0037] Machine Learning (ML) is an interdisciplinary subject that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and learning from demonstration.

[0038] With the research and progress of artificial intelligence technology, artificial intelligence technology has been studied and applied in multiple fields. For example, common ones include smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, driverless, autonomous driving, drones, robots, smart healthcare, smart customer service, etc. It is believed that with the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role.

[0039] The solution provided in the embodiments of this application involves technologies such as machine learning in artificial intelligence, which will be specifically described through the following embodiments.

[0040] In some embodiments, the implementation environment in the embodiments of this application may include a terminal and a server. Among them, data communication is carried out between the terminal and the server through a communication network. Optionally, the communication network can be a wired network or a wireless network, and this communication network can be at least one of a local area network, a metropolitan area network, and a wide area network.

[0041] The terminal is an electronic device installed with an application program having the function of training an emotion recognition model. Among them, this function of training an emotion recognition model can be a function of a native application in the terminal, or a function of a third-party application; the terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart TV, a wearable device, a vehicle terminal, etc.

[0042] The server can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms. In the embodiments of the present application, the server can be the background server of an application program with the function of training an emotion recognition model.

[0043] In some embodiments, there is data interaction between the server and the terminal. Taking the example that the training method of the emotion recognition model provided by the embodiments of the present application is executed by the server, after the terminal obtains a large number of sample texts and determines the emotion truth value labels corresponding to each sample text, it generates a training data set and sends the training data set to the server. Then, the server performs emotion recognition on the sample texts in the training data set through the emotion recognition model to obtain the sample emotion recognition results corresponding to each sample text, and determines the loss adjustment weights for each sample text according to the sample features and sample emotion recognition results of the sample texts. Furthermore, the server determines the emotion recognition loss of the sample texts based on the loss adjustment weights, sample emotion recognition results, and emotion truth value labels, and trains the emotion recognition model with the emotion recognition loss. Finally, after completing the training of the emotion recognition model, the server returns the model parameters of the emotion recognition model to the terminal, and the terminal uses the trained emotion recognition model to perform emotion recognition.

[0044] Please refer to Figure 1 , which shows a flowchart of the training method of the emotion recognition model provided by an exemplary embodiment of the present application. In this embodiment, the method is described by taking it as an example for a computer device (including a terminal and / or a server). The method includes the following steps:

[0045] Step 101, perform emotion recognition on the sample texts in the training data set through the emotion recognition model to obtain the sample emotion recognition results corresponding to each sample text. The training data set includes sample texts and the emotion truth value labels corresponding to the sample texts.

[0046] Optionally, the training data set includes sample texts and the emotion truth value labels corresponding to the sample texts. Among them, the sample text can be a descriptive text or a dialogue text, and the embodiments of the present application do not make specific limitations on the text type of the sample text.

[0047] Optionally, the emotion truth value label is used to represent whether the sample text contains a certain emotion type. If the sample text contains the emotion type, the emotion truth value label can be marked as 1; if the sample text does not contain the emotion type, the emotion truth value label can be marked as 0.

[0048] Optionally, the emotion truth value label can further characterize the emotional concentration of the sample text containing a certain emotion type. For example, the value range of the emotion truth value label can be 0 to 1, and the higher the emotional concentration, the larger the value of the emotion truth value label.

[0049] Optionally, a sample text can have a single emotion truth value label, that is, it only contains a single emotion type, such as joy, sadness, anger, etc.; a sample text can also have multiple emotion truth value labels, that is, it contains at least two emotion types, such as crying with joy, grief and indignation, etc.

[0050] Optionally, the model structure of the emotion recognition model can be Chinese-BERT-WWM. The model input is the sample text, and the model output is the emotion recognition result. Among them, BERT (Bidirectional Encoder Representations from Transformers) is an open-source language model. Based on the BERT model, a text learning task paradigm can be formed by pre-training on a large-scale language data and then fine-tuning on the target task.

[0051] Schematically, as Figure 2 shown, the BERT model includes an embedding process from the input text (Token) to generate the input sequence E, a model core architecture composed of multiple Transformer encoders ( Figure 2 abbreviated as Trm in the figure), and an output for the target classification task with T layers. Among them, each Transformer encoder contains two main parts: a self-attention module and a feed-forward neural network. The structure of the Transformer encoder is as Figure 3As shown, the Multi-head attention is a self-attention module. The self-attention module includes three sequence input heads, which are respectively used to input Query, Key, and Value, so as to respectively use a Linear layer to transform the input sequence, and calculate and process it through a Scaled dot-product Attention (SDPA) module to obtain attention scores. Then, a Concat module is used for concatenation processing, and finally the result is output through a Linear layer; Feed-Forward (feed-forward network) is an intermediate layer module, including a fully connected layer and an activation layer (such as tanh activation); Add&Norm represents residual connection and layer normalization. The operation of the Add&Norm layer is to first sum the input of the previous layer (such as the Feed-Forward layer or the Multi-head attention layer) and the input of this layer, and then calculate the layer normalization process.

[0052] In some embodiments, in order to train the emotion recognition model, the computer device can first perform emotion recognition on the sample texts in the training dataset through the emotion recognition model, so as to obtain the sample emotion recognition results corresponding to each sample text.

[0053] Optionally, the sample texts in the training dataset can be non-related texts collected through different text collection methods. For example, the sample texts can be comment contents on social platforms, descriptive texts in books and magazines, or other texts containing emotions. The embodiments of the present application do not limit this.

[0054] Optionally, the sample texts in the training dataset can also be dialogue texts with relevance between contexts, such as interview dialogue texts, scenario dialogue texts, plot dialogue texts, etc. The embodiments of the present application do not limit this.

[0055] In a possible implementation manner, the computer device sequentially inputs each sample text into the emotion recognition model. First, through the Embedding processing layer in the emotion recognition model, each word in the sample text is mapped to the corresponding dictionary ID (Token ID), and the Embedding of this dictionary ID is used as the Embedding of this word. Then, through the encoding learning of multiple Transformer encoders, the emotion prediction probabilities corresponding to each emotion label task can be output, that is, the sample emotion recognition results.

[0056] Schematically, the input data format of the emotion recognition model can be as Figure 4As shown, the input data includes the CLS token and the tokenized results of the sample text (including token embeddings and position embeddings). In the embodiments of the present application, the emotion labels are divided into 8 emotion types, that is, the CLS is represented by eight emotion label tasks, namely cls1, cls2, cls3, cls4, cls5, cls6, cls7, and cls8. Each token is used to output the emotion recognition results corresponding to each emotion label respectively.

[0057] In a possible implementation manner, when the training data set contains a large number of sample texts, in order to improve the training efficiency of the emotion recognition model, the computer device can also process the training data set in batches. For example, when the training data set includes N sample texts, each m sample texts can be used as a batch, for a total of N / m batches. Thus, the computer device can perform emotion recognition on the sample texts of each batch in turn to obtain the sample emotion recognition results corresponding to the sample texts of each batch.

[0058] Step 102: Determine the loss adjustment weight corresponding to the sample text according to the sample features of the sample text and the sample emotion recognition results. The sample features include at least one of the sample type and the number of label samples of the emotion true value label corresponding to the sample text. The sample type includes easy samples and difficult samples. The number of label samples refers to the number of sample texts with emotion true value labels in the training data set.

[0059] Optionally, considering that the emotion recognition model has different prediction difficulties for different sample texts, some sample texts are easier to be predicted by the emotion recognition model, while some sample texts are difficult to be predicted by the emotion recognition model. Therefore, the sample type can be divided into easy samples and difficult samples. An easy sample is a sample text that is easy to be predicted by the emotion recognition model, and a difficult sample is a sample text that is difficult to be predicted by the emotion recognition model.

[0060] Optionally, due to the uncertainty in the collection of sample texts, that is, there may be a large difference in the number of sample texts corresponding to different emotion labels, which affects the training effect of the emotion recognition model. For example, some emotion labels correspond to a large number of sample texts, while some emotion labels only correspond to a small number of sample texts. Therefore, the emotion labels can be divided into head labels and tail labels. A head label is an emotion label corresponding to a relatively large number of sample texts, and a tail label is an emotion label corresponding to a relatively small number of sample texts.

[0061] Schematically, such as Figure 5As shown, it shows the distribution of the number of sample texts corresponding to eight emotion labels. Among them, the number of sample texts corresponding to the three emotion labels of "worry", "doubt", and "joy" is significantly larger, which are the head labels; the number of sample texts corresponding to the five emotion labels of "faith", "resentment", "fear", "expectation", and "love" is significantly smaller, which are the tail labels.

[0062] Different from the related art, after obtaining the sample emotion recognition results of each sample text, the emotion recognition loss is determined according to the sample emotion recognition results and the emotion true value labels, and then the emotion recognition model is trained using the emotion recognition loss. In the embodiments of the present application, in order to optimize the optimization effect of the emotion recognition model, after obtaining the sample emotion recognition results of each sample text, the emotion recognition loss is not directly determined. Instead, according to the sample features of each sample text and the sample emotion recognition results, the loss adjustment weights of each sample text are first determined.

[0063] Among them, the sample features of the sample text include at least one of the sample type and the number of label samples of the emotion true value label corresponding to the sample text. The sample type includes easy samples and difficult samples, and the number of label samples refers to the number of sample texts with emotion true value labels in the training dataset.

[0064] In a possible implementation manner, since it is easier for the emotion recognition model to predict easy samples, while it is more difficult for the emotion recognition model to predict difficult samples, the accuracy of the sample emotion recognition results corresponding to easy samples is higher than that of the sample emotion recognition results corresponding to difficult samples. Therefore, in order to optimize the learning effect of the emotion recognition model on difficult samples during the model training process, the computer device needs to first determine the loss adjustment weights of the sample texts according to the sample types of the sample texts, so that the loss adjustment weights of difficult samples are greater than those of easy samples, and the emotion recognition model can focus on learning difficult samples.

[0065] In another possible implementation, since the number of sample texts corresponding to different emotion labels is different, for the head labels with a larger number of sample texts, the emotion recognition model may continuously repeat learning the emotion features; for the tail labels with a smaller number of sample texts, the emotion recognition model may not be able to fully learn the emotion features. Therefore, during the model training process, it is necessary to balance the learning effects of the emotion recognition model on different emotion features. And for a tail label, there is a single-label sample text and a multi-label sample text respectively. Due to the co-occurrence influence of other head labels in the multi-label sample text, the loss contribution of the multi-label sample text to this tail label will be greater than that of the single-label sample text, thus increasing the learning difficulty of the emotion recognition model for the tail label. Therefore, during the model training process, it is also necessary to adjust the loss of the multi-label sample text containing the tail label. That is, the computer device can first determine the loss adjustment weight of the sample text according to the number of label samples of the emotion true value label corresponding to the sample text, so as to reduce the contribution of the head label to the model loss, avoid the repeated learning of the emotion features of the head label with a large number of sample texts, and improve the label weight of the tail label in the multi-label sample text.

[0066] Step 103: Determine the emotion recognition loss of the sample text based on the loss adjustment weight, the sample emotion recognition result, and the emotion true value label.

[0067] In some embodiments, after determining the loss adjustment weight of each sample text, the computer device can determine the emotion recognition loss of each sample text based on the loss adjustment weight, the sample emotion recognition result, and the emotion true value label.

[0068] In a possible implementation, the computer device can first determine the emotion recognition loss according to the emotion true value label and the sample emotion recognition result of the sample text based on the principle of cross-entropy loss determination, and then adjust the emotion recognition loss using the loss adjustment weight to obtain the adjusted emotion recognition loss.

[0069] Step 104: Train the emotion recognition model based on the emotion recognition loss.

[0070] In some embodiments, after determining the emotion recognition loss of each sample text, the computer device can execute the backpropagation algorithm to calculate the parameter gradients of each model parameter in the emotion recognition model based on the emotion recognition loss, and then re-determine the new model parameters according to the model parameters, the parameter gradients, and the learning rate to update the parameters of the emotion recognition model, that is, complete one training of the emotion recognition model.

[0071] In a possible implementation, after obtaining the emotion recognition losses of each sample text, the computer device can accumulate the emotion recognition losses of each sample text to obtain the total emotion recognition loss corresponding to the current training round, and then use the total emotion recognition loss to train the emotion recognition model.

[0072] In another possible implementation, after obtaining the total emotion recognition loss, the computer device can further obtain the average emotion recognition loss of the current training round according to the number of sample texts in the training dataset, and then use the average emotion recognition loss to train the emotion recognition model.

[0073] In a possible implementation, when the sample texts in the training dataset are processed in batches, after determining the emotion recognition losses of each batch of sample texts, the computer device can first train the emotion recognition model based on the emotion recognition loss of the current batch of sample texts, and then use the trained emotion recognition model to perform emotion recognition on the next batch of sample texts, so as to improve the model training efficiency.

[0074] In summary, in the embodiments of the present application, after the emotion recognition model performs emotion recognition on the sample texts in the training dataset to obtain the sample emotion recognition results of each sample text, the emotion recognition loss is not directly determined according to the sample emotion recognition results and the emotion true value labels. Instead, the sample characteristics of each sample text in the training dataset are fully considered. First, the loss adjustment weight is determined according to the sample characteristics and the sample emotion recognition results of the sample text, and then the emotion recognition loss is jointly determined by combining the loss adjustment weight, the sample emotion recognition result, and the emotion true value label. Furthermore, the emotion recognition model is trained based on the emotion recognition loss. By using the loss adjustment weight to determine the emotion recognition loss, the proposed solution in the embodiments of the present application can adjust the emotion recognition losses of easy samples and difficult samples, and can also adjust the emotion recognition losses of unbalanced label samples when the number of sample texts with different emotion true value labels is unbalanced. Therefore, by training the emotion recognition model based on the adjusted emotion recognition loss, the training effect of the emotion recognition model can be optimized, and the accuracy of the output of the emotion recognition model can be improved.

[0075] In some embodiments, considering that the prediction difficulty of the emotion recognition model is different for easy samples and difficult samples, and the learning effect of the emotion recognition model is also different for head labels and tail labels, in order to ensure refined learning of the sample text and improve the training efficiency of the emotion recognition model, the computer device may also divide the process of training the emotion recognition model into two stages. In the first training stage, the emotion recognition model is trained using sample texts with a single emotion ground truth label. In the second training stage, the emotion recognition model is trained using sample texts with at least one emotion ground truth label, that is, the emotion recognition model is obtained through two-stage training.

[0076] Among them, in the first training stage, the computer device trains the emotion recognition model based on the first emotion recognition loss, and the first emotion recognition loss is the emotion recognition loss corresponding to the sample text in the first training dataset, and the sample text in the first training dataset has a single emotion ground truth label.

[0077] In a schematic example, the emotion labels can be divided into eight types, respectively indicating eight emotions: trust, resentment, joy, fear, expectation, anxiety, doubt, and love. The computer device respectively obtains the sample texts corresponding to each emotion type, that is, the sample text with only the emotion of "trust", the sample text with only the emotion of "resentment", the sample text with only the emotion of "joy", the sample text with only the emotion of "fear", the sample text with only the emotion of "expectation", the sample text with only the emotion of "anxiety", the sample text with only the emotion of "doubt", and the sample text with only the emotion of "love", so as to form the first training dataset based on the sample texts of the above emotion types.

[0078] After training the emotion recognition model with single-label sample texts, the emotion recognition model has basic emotion recognition capabilities. Further, considering that the emotions contained in the text are relatively rich and there may be texts containing multiple emotion types, in order to optimize the emotion recognition model, the computer device also needs to train the emotion recognition model using multi-label sample texts.

[0079] Therefore, in the second training stage, based on the emotion recognition model trained in the first training stage, the computer device trains the emotion recognition model based on the second emotion recognition loss, and the second emotion recognition loss is the emotion recognition loss corresponding to the sample text in the second training dataset, and the sample text in the second training dataset has at least one emotion ground truth label.

[0080] Optionally, in addition to single-label sample texts, the second training dataset further includes multi-label sample texts. Among them, the multi-label sample text is a sample text containing at least two emotion types. For example, a sample text with two emotions of "faith" and "resentment", a sample text with two emotions of "joy" and "faith", a sample text with two emotions of "worry" and "fear", and so on.

[0081] Schematically, as Figure 6 shown, it shows a schematic diagram of the staged training of the emotion recognition model provided by an exemplary embodiment of the present application. First, in the first training stage, the computer device inputs the first training dataset 601 into the emotion recognition model 605, so as to obtain the sample emotion recognition results 606 corresponding to each sample text in the first training dataset 601 output by the emotion recognition model 605. Then, after determining the first emotion recognition loss 602, the computer device performs the first-stage training on the emotion recognition model 605 based on the first emotion recognition loss 602. On the basis of completing the first-stage training, the computer device inputs the second training dataset 603 into the emotion recognition model 605, so as to obtain the sample emotion recognition results 607 corresponding to each sample text in the second training dataset 603 output by the emotion recognition model 605. Then, after determining the second emotion recognition loss 604, the computer device performs the second-stage training on the emotion recognition model 605 based on the second emotion recognition loss 604.

[0082] Please refer to Figure 7 , which shows a flowchart of the training method of the emotion recognition model provided by another exemplary embodiment of the present application. In this embodiment, it is described by taking this method as being used in a computer device (including a terminal and / or a server) as an example. The method includes the following steps:

[0083] Step 701, perform emotion recognition on the sample texts in the first training dataset through the emotion recognition model to obtain the sample emotion recognition results corresponding to each sample text. The first training dataset includes sample texts and single emotion true value labels corresponding to the sample texts.

[0084] In some embodiments, in order to improve the model training efficiency, before using the emotion recognition model to perform emotion recognition on the sample texts, the computer device can first perform parameter initialization processing on the emotion recognition model. For example, in the case of adopting the basic model structure of Chinese-BERT-WWM, the computer device can use the pre-training parameters of the Chinese-BERT-WWM model to perform parameter initialization processing on the emotion recognition model and set the learning rate of the emotion recognition model. Then, after completing the parameter initialization and setting the learning parameters, the computer device performs pre-training on the emotion recognition model.

[0085] Optionally, the first training dataset contains sample texts and the corresponding single emotion true value label. In the embodiments of the present application, the emotion types are divided into eight categories, namely faith, resentment, joy, fear, expectation, anxiety, doubt, and love. Among them, the number of sample texts corresponding to each emotion type can be the same or different, and the embodiments of the present application do not limit this.

[0086] Optionally, when the sample text has a certain emotion type, its corresponding emotion true value label can be 1; when the sample text does not have a certain emotion type, its corresponding emotion true value label can be 0. Schematically, for the sample text "Yes, I just believe her!", its corresponding emotion true value label is faith-1.

[0087] In a possible implementation manner, the computer device inputs the sample text and the corresponding emotion true value label into the emotion recognition model. Thus, the Embedding processing layer in the emotion recognition model first uses the dictionary vocab.txt to map each word in the sample text to its corresponding dictionary ID (Token ID), and at the same time adds the marker CLS as the start of the text sequence. And considering that the text lengths of different sample texts are different, the model input can also be normalized, such as using token=0 for padding, so that the number of tokens of the sample text reaches a certain value. Thus, the Embedding of the dictionary ID can be used as the word Embedding, and then through the encoding learning of multiple Transformer encoders in the emotion recognition model, the model output corresponding to the sample text can be obtained.

[0088] Among them, the model output includes the sample emotion recognition result and the text learning content. The sample emotion recognition result corresponds to the marker CLS, and CLS is represented by eight emotion label tasks, namely cls1, cls2, cls3, cls4, cls5, cls6, cls7, cls8. The output probability corresponding to each marker represents the emotion recognition result of the emotion recognition model for the emotion indicated by the marker.

[0089] Step 702, determine the first loss adjustment weight corresponding to the sample text according to the sample type of the sample text and the sample emotion recognition result. The first loss adjustment weight is used to increase the emotion recognition loss of difficult samples.

[0090] In some embodiments, considering that the emotion recognition model has different prediction difficulties for different sample texts, after obtaining the sample emotion recognition result of the sample text, the computer device can also determine the first loss adjustment weight according to the sample type of the sample text and the sample emotion recognition result, so as to increase the emotion recognition loss of difficult samples during the model training process.

[0091] Optionally, the first loss adjustment weight may be negatively correlated with the prediction difficulty of the sample text, that is, for easy samples, it is necessary to reduce their contribution to the model training loss; for difficult samples, it is necessary to increase their contribution to the model training loss.

[0092] Optionally, when the emotion ground truth label is 1, the first loss adjustment weight can be expressed as When the emotion ground truth label is 0, the first loss adjustment weight can be expressed as Where is the sample emotion recognition result, that is, the prediction probability of sample text i on emotion label k, β is the adjustment index, and β is preset.

[0093] For easy samples, their output probability is usually large and close to 1, so is small. When the exponent β is greater than 1, the first loss adjustment weight corresponding to easy samples will be relatively small. For example, when the exponent β is 2, is 0.9, its first loss adjustment weight is equal to 0.01; when the exponent β is 2, is 0.95, its first loss adjustment weight is equal to 0.0025, thus reducing the contribution of easy samples to the model training loss.

[0094] For difficult samples, since the emotion recognition model is difficult to predict them, their output probability is usually small. At this time is large. When the exponent β is greater than 1, the first loss adjustment weight corresponding to difficult samples will be relatively large. For example, when the exponent β is 2, is 0.5, its first loss adjustment weight is equal to 0.25; when the exponent β is 2, is 0.2, its first loss adjustment weight is equal to 0.64. Compared with easy samples, the first loss adjustment weight corresponding to difficult samples increases exponentially, thus increasing the contribution of difficult samples to the model training loss.

[0095] Step 703: Determine the second loss adjustment weight corresponding to the sample text according to the number of label samples of the emotion ground truth label and the number of first label categories in the first training dataset. The number of first label categories refers to the total number of label categories included in the first training dataset. The second loss adjustment weight is negatively correlated with the number of label samples.

[0096] In some embodiments, considering the problem of unbalanced sample distribution caused by the head tags and tail tags, after obtaining the sample emotion recognition result, the computer device may further determine the second loss adjustment weight corresponding to the sample text according to the number of label samples of the emotion true value label and the number of first label categories in the first training dataset.

[0097] Wherein, the number of first label categories refers to the total number of label categories included in the first training dataset. The second loss adjustment weight has a negative correlation with the number of label samples. The larger the number of sample texts corresponding to an emotion label, the larger the second loss adjustment weight corresponding to the sample text with that emotion label.

[0098] Optionally, the second loss adjustment weight can be expressed as Where C represents the total number of label categories included in the first training dataset, C is equal to 8, and n k represents the number of label samples corresponding to the emotion label k in the first training dataset. For example, for an emotion label with only 200 sample texts, its corresponding second loss adjustment weight is 1 / 8 * 1 / 200; for an emotion label with 10,000 sample texts, its corresponding second loss adjustment weight is 1 / 8 * 1 / 10000. Thus, through the second loss adjustment weight, the contribution of the sample text with the head tag to the model training loss can be reduced, avoiding the repeated learning of the head tag information by the emotion recognition model, and increasing the contribution of the sample text with the tail tag to the model training loss.

[0099] Step 704, determine the first emotion recognition loss of the sample text based on the first loss adjustment weight, the second loss adjustment weight, the sample emotion recognition result, and the emotion true value label.

[0100] In some embodiments, after respectively determining the first loss adjustment weight and the second loss adjustment weight, the computer device can determine the first emotion recognition loss corresponding to the sample text according to the first loss adjustment weight, the second loss adjustment weight, the sample emotion recognition result, and the emotion true value label of each sample text, based on the cross-entropy loss principle.

[0101] Optionally, the first emotion recognition loss can be expressed as Where is the annotation record of the sample text i on the emotion label k, is the sample emotion recognition result, that is, the predicted probability of the sample text i on the emotion label k, β is the adjustment index, and R is the second loss adjustment weight.

[0102] Optionally, in the case of determining the emotion true value label by the dichotomy method, is 0 or 1, and the first emotion recognition loss can also be expressed as:

[0103]

[0104] Step 705: Train the emotion recognition model based on the first emotion recognition loss.

[0105] In some embodiments, after determining the first emotion recognition loss of the sample text, the computer device can use the first emotion recognition loss to pre-train the emotion recognition model. That is, based on the first emotion recognition loss, execute the backpropagation algorithm to calculate the parameter gradients of each model parameter in the emotion recognition model, and then re-determine the new model parameters according to the model parameters, parameter gradients, and learning rate, and update the parameters of the emotion recognition model, that is, complete one pre-training of the emotion recognition model.

[0106] In a possible implementation, when the first training data set contains a large number of sample texts, in order to improve the training efficiency of the emotion recognition model, the computer device can also process the first training data set in batches. For example, when the first training data set includes N sample texts, every m sample texts can be used as a batch, for a total of N / m batches. Completing N / m batches means completing one round (epoch) of iteration.

[0107] Optionally, in order to improve the model convergence accuracy and the model generalization ability, the computer can also perform a learning rate decay during the model training process. For example, the learning rate can be set to 0.005 in the model initialization stage, and during the model training process, it can be set to reduce the learning rate to 0.1 times the original value after every 10 training rounds.

[0108] In a possible implementation, during the process of training the emotion recognition model, the computer device can record the average emotion recognition loss corresponding to the first training data set in each training round. Thus, when the decrease in the average emotion recognition loss is significantly small or no longer decreases after multiple rounds of training, the first-stage training of the emotion recognition model is ended.

[0109] Step 706: Based on the emotion recognition model trained in the first training stage, perform emotion recognition on the sample texts in the second training data set through the emotion recognition model to obtain the sample emotion recognition results corresponding to each sample text. The second training data set includes sample texts and at least one emotion ground-truth label corresponding to the sample texts.

[0110] In some embodiments, after the emotion recognition model is obtained by completing the first-stage training, the emotion recognition model already has basic emotion recognition capabilities. To improve the recognition accuracy of the emotion recognition model and its ability to recognize emotions in multi-label texts, the computer device can also use the emotion recognition model to recognize the emotions of the sample texts in the second training dataset, and obtain the sample emotion recognition results corresponding to each sample text.

[0111] Among them, the second training dataset includes sample texts and at least one emotion ground-truth label corresponding to the sample texts. In the embodiments of the present application, the emotion types are divided into eight categories, namely, faith, resentment, joy, fear, expectation, anxiety, doubt, and love.

[0112] In a possible implementation, the computer device inputs the sample text and the emotion ground-truth label corresponding to the sample text into the emotion recognition model. Thus, the Embedding processing layer in the emotion recognition model first uses the dictionary vocab.txt to map each word in the sample text to its corresponding dictionary ID (Token ID), and at the same time adds the marker CLS as the start of the text sequence. Considering that the text lengths of different sample texts are different, the model input can also be normalized, for example, using token = 0 for padding, so that the number of tokens in the sample text reaches a certain value. Thus, the Embedding of the dictionary ID can be used as the word Embedding, and then through the encoding learning of multiple Transformer encoders in the emotion recognition model, the model output corresponding to the sample text can be obtained.

[0113] Among them, the model output includes the sample emotion recognition result and the text learning content. The sample emotion recognition result corresponds to the marker CLS, and CLS is represented by eight emotion label tasks, namely, cls1, cls2, cls3, cls4, cls5, cls6, cls7, and cls8. The output probability corresponding to each marker represents the emotion recognition result of the emotion recognition model for the emotion indicated by the marker.

[0114] Step 707: Determine the first loss adjustment weight corresponding to the sample text according to the sample type of the sample text and the sample emotion recognition result. The first loss adjustment weight is used to increase the emotion recognition loss of difficult samples.

[0115] In some embodiments, considering that the emotion recognition model has different prediction difficulties for different sample texts, after obtaining the sample emotion recognition results of the sample texts, the computer device can also determine the first loss adjustment weight according to the sample type of the sample text and the sample emotion recognition result, so as to increase the emotion recognition loss of difficult samples during the model training process.

[0116] Optionally, the first loss adjustment weight may be negatively correlated with the prediction difficulty of the sample text, that is, for easy samples, it is necessary to reduce their contribution to the model training loss; for difficult samples, it is necessary to increase their contribution to the model training loss.

[0117] Optionally, when the emotion ground truth label is 1, the first loss adjustment weight can be expressed as When the emotion ground truth label is 0, the first loss adjustment weight can be expressed as Where is the sample emotion recognition result, that is, the prediction probability of the sample text i on the emotion label k, β is the adjustment index, and β is preset.

[0118] For easy samples, their output probability is usually large and close to 1. Therefore is small. When the exponent β is greater than 1, the first loss adjustment weight corresponding to the easy sample will be relatively small. For example, when the exponent β is 2, is 0.9, its first loss adjustment weight is equal to 0.01; when the exponent β is 2, is 0.95, its first loss adjustment weight is equal to 0.0025, thus reducing the contribution of easy samples to the model training loss.

[0119] For difficult samples, since it is difficult for the emotion recognition model to predict them, their output probability is usually small. At this time is large. When the exponent β is greater than 1, the first loss adjustment weight corresponding to the difficult sample will be relatively large. For example, when the exponent β is 2, is 0.5, its first loss adjustment weight is equal to 0.25; when the exponent β is 2, is 0.2, its first loss adjustment weight is equal to 0.64. Compared with easy samples, the first loss adjustment weight corresponding to difficult samples increases exponentially, thus increasing the contribution of difficult samples to the model training loss.

[0120] Step 708: Determine the second loss adjustment weight corresponding to the sample text according to the number of sample label categories of the sample text, the number of label samples of the emotion ground truth label, and the number of second label categories in the second training data set.

[0121] In some embodiments, considering the problem of unbalanced sample distribution caused by head tags and tail tags, as well as the co-occurrence migration problem existing in multi-label sample texts, after obtaining the sample emotion recognition result, the computer device can also determine the second loss adjustment weight corresponding to the sample text according to the number of sample label categories of the sample text, the number of label samples of the emotion true value label, and the number of second label categories in the second training dataset.

[0122] In a possible implementation manner, when the number of sample label categories of the sample text is 1, that is, the sample text is a single-label sample, similar to determining the second loss adjustment weight in the first training stage, the computer device can determine the second loss adjustment weight corresponding to the sample text based on the number of label samples of the emotion true value label and the number of second label categories in the second training dataset.

[0123] Optionally, the second loss adjustment weight can be expressed as where C represents the total number of label categories included in the second training dataset, C is equal to 8, and n k represents the number of label samples corresponding to the emotion label k in the second training dataset.

[0124] In another possible implementation manner, when the number of sample label categories of the sample text is greater than 1, that is, the sample text is a multi-label sample, considering the co-occurrence migration problem existing in the multi-label sample text, the computer device first needs to determine the label weights of each emotion true value label according to the number of label samples of each emotion true value label, and determine the target emotion true value label as the emotion true value label with the least number of label samples (i.e., the tail tag in the multi-label sample text). Then, according to the number of second label categories, and the weight difference between the label weight of the target emotion true value label and the label weights of other emotion true value labels, the computer device determines the second loss adjustment weight corresponding to the sample text.

[0125] Optionally, the label weight can be m = 1 / n k The more the number of label samples of the emotion true value label, the lower its corresponding label weight, that is, the label weight of the tail tag in the multi-label sample text is higher than that of the head tag.

[0126] Optionally, the second loss adjustment weight can be expressed as where That is, by subtracting the label weights of other emotion true value labels from the label weight of the target emotion true value label, the influence of the co-occurrence migration of the tail tag in the multi-label sample text can be eliminated.

[0127] For the single-label sample text and multi-label sample text corresponding to the tail label, the contribution of the emotion recognition loss of the single-label sample text to the learning of the tail label will be greater than the contribution of the emotion recognition loss of the multi-label sample text to the learning of the tail label. Taking the example that the tail label a has a single-label sample text i and a multi-label sample text x (which contains the head label k in addition to the tail label a), the second loss adjustment weight corresponding to the single-label sample text i is The second loss adjustment weight corresponding to the multi-label sample text x is That is, r1 > r2.

[0128] Schematically, as Figure 8 shown, for 3 emotion true value labels, N1, N2, and N3 respectively represent the number of label samples corresponding to each emotion true value label in the second training dataset. Among them, the label weight corresponding to the first emotion true value label 801 is 1 / N1, the label weight corresponding to the second emotion true value label 802 is 1 / N2, and the label weight corresponding to the third emotion true value label 803 is 1 / N3.

[0129] For the sample text located in region 804, that is, having the first emotion true value label 801 and the third emotion true value label 803, its corresponding second loss adjustment weight is For the sample text located in region 805, that is, having the second emotion true value label 802 and the third emotion true value label 803, its corresponding second loss adjustment weight is For the sample text located in region 806, that is, having the first emotion true value label 801 and the second emotion true value label 802, its corresponding second loss adjustment weight is For the sample text located in region 807, that is, having the first emotion true value label 801, the second emotion true value label 802, and the third emotion true value label 803, its corresponding second loss adjustment weight is

[0130] Step 709, based on the first loss adjustment weight, the second loss adjustment weight, the sample emotion recognition result, and the emotion true value label, determine the second emotion recognition loss of the sample text.

[0131] In some embodiments, after respectively determining the first loss adjustment weight and the second loss adjustment weight, the computer device can then determine the second emotion recognition loss corresponding to the sample text according to the first loss adjustment weight, the second loss adjustment weight, the sample emotion recognition result, and the emotion true value label of each sample text, based on the cross-entropy loss principle.

[0132] Optionally, the second emotion recognition loss can be expressed as Among them, is the annotation record of sample text i on emotion label k. is the sample emotion recognition result, that is, the predicted probability of sample text i on emotion label k, β is the adjustment index, and r is the second loss adjustment weight.

[0133] Optionally, in the case of determining the emotion true label by the dichotomy method, is 0 or 1, and the first emotion recognition loss can also be expressed as:

[0134]

[0135] Step 710, train the emotion recognition model based on the second emotion recognition loss.

[0136] In some embodiments, after determining the second emotion recognition loss of the sample text, the computer device can use the second emotion recognition loss to pre-train the emotion recognition model. That is, based on the second emotion recognition loss, execute the backpropagation algorithm to calculate the parameter gradients of each model parameter in the emotion recognition model, and then re-determine the new model parameters according to the model parameters, parameter gradients, and learning rate, and update the parameters of the emotion recognition model, that is, complete one training of the emotion recognition model. Then, when the decrease in the average emotion recognition loss in the current training round is significantly smaller or no longer decreases, the second-stage training of the emotion recognition model can be ended, thereby obtaining a trained emotion recognition model.

[0137] In the above embodiments, the training process of the emotion recognition model is divided into two stages. First, train the emotion recognition model with sample texts having a single emotion true label, so that the emotion recognition model has a basic emotion recognition ability. Then, after completing the first training stage, train the emotion recognition model with sample texts having at least one emotion true label, so that the emotion recognition model has the ability to recognize multiple emotions. By training the model in stages, the training effect of the model is optimized and the training efficiency of the model is improved.

[0138] In addition, in the process of determining the loss adjustment weight, determine the first loss adjustment weight according to the sample type and the sample emotion recognition result, so as to increase the contribution of the emotion recognition loss of difficult samples to the model training loss, and reduce the contribution of the emotion recognition loss of easy samples to the model training loss, so that the emotion recognition model can focus on learning a small number of difficult samples, and improve the model training quality of the emotion recognition model.

[0139] Moreover, the second loss adjustment weight is determined according to the number of label categories and the number of label samples, which improves the contribution of the sample text with tail labels to the model training loss and reduces the contribution of the sample text with head labels to the model training loss, avoiding the repeated learning of head label information by the emotion recognition model. Meanwhile, in the second training stage, it is possible to control the second loss adjustment weight of the single-label sample text with tail labels to be greater than that of the multi-label sample text with tail labels, thereby reducing the co-occurrence migration problem existing in the imbalanced samples and further improving the model training quality of the emotion recognition model.

[0140] In some embodiments, to improve the output accuracy of the emotion recognition model, the computer device can also respectively set a text feature learning network and an emotion detail learning network in the emotion recognition model. Among them, the text feature learning network is used to perform basic text semantic understanding on the sample text, and the emotion detail learning network is used to further learn delicate emotion information from the text semantic understanding result.

[0141] In a possible implementation manner, the computer device first inputs each sample text in the training dataset into the text feature learning network in the emotion recognition model to obtain the sample semantic understanding result corresponding to each sample text output by the text feature learning network, and then inputs the sample semantic understanding result into the emotion detail learning network in the emotion recognition model to obtain the sample emotion recognition result corresponding to each sample text output by the emotion detail learning network.

[0142] Optionally, the network structure of the text feature learning network can adopt the basic model structure of Chinese-BERT-WWM, and the emotion detail learning network can be implemented by adding multiple layers of Transformer encoders to the basic model structure of Chinese-BERT-WWM.

[0143] Among them, before training the emotion recognition model, for the newly added multiple layers of Transformer encoders, the computer device can perform parameter initialization processing on them using a Gaussian normal distribution between 0 and 1, or can perform parameter initialization processing using the pre-trained parameters with the same structure in the Chinese-BERT-WWM model. The embodiments of the present application do not limit this.

[0144] Schematically, as Figure 9As shown in the figure, the computer device first inputs each sample text in the training data set 901 into the text feature learning network 902 in the emotion recognition model to obtain the sample semantic understanding results corresponding to each sample text output by the text feature learning network 902. Then, the sample semantic understanding results are input into the emotion detail learning network 903 in the emotion recognition model to obtain the sample emotion recognition results 904 corresponding to each sample text output by the emotion detail learning network 903.

[0145] In the above embodiment, based on the existing text feature learning network, an emotion detail learning network is further added to the emotion recognition model, enabling the emotion recognition model to learn more delicate emotion information, optimizing the model performance, and improving the model training effect.

[0146] In some embodiments, after the training of the emotion recognition model is completed, the computer device can use the emotion recognition model to recognize the emotion of any text and obtain the corresponding emotion recognition result.

[0147] In a possible implementation manner, the emotion recognition model can be applied to social network analysis. For example, the emotion recognition model is used to recognize the emotion of the review text published by users on the social network platform, so as to analyze the user's emotional tendency based on the emotion recognition label corresponding to the review text for marketing and product positioning.

[0148] In another possible implementation manner, the emotion recognition model can also be applied to customer feedback analysis. For example, the emotion recognition model is used to recognize the emotion of the customer feedback text, so as to analyze the customer's psychology and understand the customer's satisfaction with the product and service based on the emotion recognition label corresponding to the customer feedback text, in order to improve customer satisfaction.

[0149] In another possible implementation manner, the emotion recognition model can also be applied to advertisement evaluation. For example, the emotion recognition model is used to perform sentiment analysis on the advertisement copy, so as to help the customer understand the advertisement effect and optimize the advertisement strategy based on the emotion recognition label corresponding to the advertisement text.

[0150] In a possible implementation manner, in order to determine whether the emotion design of the script characters is reasonable and evaluate the emotion development of the script characters, the computer device can apply the emotion recognition model to recognize the emotion of the dialogue text in the script, and then perform quality evaluation on the script based on the emotion recognition result of the dialogue text. This process may include the following steps:

[0151] Step 1001: Use the trained emotion recognition model to recognize the emotions of the target dialogue texts of each key character in the script, and determine the target emotion labels corresponding to the target dialogue texts of each key character. The script includes at least one script scene and the target dialogue texts of at least two key characters.

[0152] Optionally, the script includes at least one script scene and the target dialogue texts of at least two key characters. Among them, the script is a text describing the plot of a film or television work, used to guide the shooting of the film or television work. A script contains multiple script scenes, and the script scenes can be divided by scenes or by duration. The embodiments of the present application do not make limitations on this. And each script scene contains at least two key characters, that is, the characters with a relatively large proportion of scenes in the script scene, usually called "protagonists".

[0153] In some embodiments, in order to evaluate the quality of the script, the computer device uses the trained emotion recognition model to recognize the emotions of the target dialogue texts of each key character in the script, so as to respectively determine the target emotion labels corresponding to the target dialogue texts of each key character.

[0154] In a possible implementation manner, considering that the output of the emotion recognition model is the emotion prediction probability of the text corresponding to each emotion label, that is, the possibility that the text contains this emotion, but the emotion recognition model has different learning abilities for different emotion labels. If the emotion label corresponding prediction probability less than 0.5 is directly determined as not containing this emotion, and the emotion label corresponding prediction probability greater than 0.5 is determined as containing this emotion, it may reduce the accuracy of determining the emotion label. Therefore, in order to improve the accuracy of determining the emotion label, before applying the emotion recognition model for emotion recognition, the computer device can also first use the emotion recognition model to recognize the emotions of each verification text in the verification dataset, and then determine the probability decision threshold corresponding to each emotion label according to the emotion recognition result corresponding to the verification text and the optimal value search.

[0155] Optionally, after the computer device outputs the emotion recognition results corresponding to each verification text through the emotion recognition model, it can perform threshold search between 0 and 1 at a step size of 0.05 to determine the optimal probability decision threshold corresponding to each emotion label.

[0156] Furthermore, after using the emotion recognition model to recognize the emotions of the target dialogue texts of each key character in the script and obtaining the target emotion recognition results corresponding to each target dialogue text, the computer device can determine whether the target dialogue text contains the emotion according to the probability decision threshold corresponding to each emotion label. If the target emotion recognition result is greater than the probability decision threshold, it means that the target dialogue text contains the emotion, that is, the target dialogue text has the emotion label; if the target emotion recognition result is less than the probability decision threshold, it means that the target dialogue text does not contain the emotion, that is, the target dialogue text does not have the emotion label, so that the target emotion labels corresponding to the target dialogue texts of each key character can be determined.

[0157] Step 1002, perform a quality assessment on the script based on the target emotion labels corresponding to the target dialogue texts of each key character to obtain a script quality assessment result.

[0158] In some embodiments, after obtaining the target emotion labels corresponding to the target dialogue texts of each key character, the computer device can determine the emotional change trend of the key character in the script according to the target emotion labels and perform a quality assessment on the script, so as to obtain a script quality assessment result.

[0159] In a possible implementation manner, in order to more intuitively analyze the emotional change trend of the key character, the computer device can generate an emotional development curve of each key character in the script scene according to the target emotion labels corresponding to the target dialogue texts of each key character in the same script scene and the emotional prediction probability corresponding to the target emotion label, and then perform a quality assessment on the script based on the emotional development curves corresponding to each key character to obtain a script quality assessment result.

[0160] Optionally, after using the emotion recognition model to recognize the emotion of the target dialogue text, the model output result can be obtained, that is, the emotional prediction probability of the emotion recognition model for each emotion label, which includes the emotional prediction probability corresponding to the target emotion label.

[0161] Regarding the method of generating an emotional development curve, the computer device can convert the emotional prediction probability corresponding to the target emotional label into an emotional score, and generate an emotional development curve based on the emotional score. In one possible implementation, when the target dialogue text has a single target emotional label, the computer device can determine the emotional score of the target dialogue text according to the emotional prediction probability corresponding to the single target emotional label. For example, multiplying the emotional prediction probability by 100 and taking the integer is the emotional score. When the target dialogue text has at least two target emotional labels, the computer device can determine the emotional score of the target dialogue text according to the probability mean value between the emotional prediction probabilities corresponding to the at least two target emotional labels; it can also select the emotional prediction probability corresponding to one of the at least two target emotional labels for calculating the emotional score.

[0162] Optionally, the computer device can determine the emotional significance degree of each target emotional label according to the emotional prediction probability corresponding to each target emotional label, and thus determine the emotional score of the target dialogue text according to the emotional prediction probability corresponding to the target emotional label with the highest emotional significance degree.

[0163] Regarding the method of determining the emotional significance degree of the target emotional label, in one possible implementation, the computer device can perform generalization processing on the emotional prediction probabilities corresponding to each target emotional label to obtain the emotional significance degree of each target emotional label, where the emotional significance degree is positively correlated with the emotional prediction probability, that is, according to the emotional prediction probabilities corresponding to each target emotional label, the target emotional label with the highest emotional prediction probability is determined as the target emotional label with the highest emotional significance degree.

[0164] In another possible implementation, the computer device can also determine the emotional bias score corresponding to each target emotional label according to the emotional prediction probability corresponding to each target emotional label and the probability decision threshold, and then determine the emotional significance degree of each target emotional label according to the emotional bias score corresponding to each target emotional label, where the emotional significance degree is positively correlated with the emotional bias score.

[0165] Optionally, the emotional bias score = (emotional prediction probability - probability decision threshold) / probability decision threshold. The higher the emotional bias score, the greater the possibility that the target dialogue text contains this emotional label and the stronger the emotion.

[0166] Further, after determining the emotional scores of each target dialogue text of the key character, the computer device can generate an emotional development curve of the key character in this script scene with the appearance time of the target dialogue text in the script scene as the abscissa and the emotional score of the target dialogue text as the ordinate according to the appearance order and emotional score of each target dialogue text of the key character in the script scene.

[0167] Optionally, the computer device can also generate an emotional development curve for a key character in a single episode of the script based on the script episode; or it can directly generate an emotional development curve for a key character in the entire script, which is not limited to the embodiments of the present application.

[0168] Optionally, the computer device can also classify emotions into positive emotions and negative emotions, where positive emotions include faith, happiness, expectation, and love, and negative emotions include resentment, fear, worry, and doubt, thereby separately counting the positive emotion scores and negative emotion scores of key figures to generate positive emotion development curves and negative emotion development curves.

[0169] Furthermore, after obtaining the emotion development curves corresponding to the key characters, the computer device can evaluate the quality of the script according to the emotion development trends of the key characters in the script. Optionally, the computer device can divide the character emotion development quality analysis into single-person emotion quality analysis and multi-person emotion quality analysis, and generate a script quality evaluation result based on the single-person emotion evaluation results of each key character and the multi-person emotion evaluation results between the key characters.

[0170] In a possible implementation, the computer device determines the emotional fluctuation amplitude of each key character in the script scene according to the emotional development curve corresponding to each key character, and then performs individual emotional quality analysis on each key character based on the emotional fluctuation amplitude to obtain individual emotional evaluation results of each key character. Among them, for those with a large emotional fluctuation amplitude, it indicates that the character has great emotional elasticity and obvious emotional changes, which is an excellent character emotional shaping; for those with a small emotional fluctuation amplitude, it indicates that the emotional shaping changes are relatively small, which is a relatively ordinary character emotional shaping.

[0171] In a possible implementation, in order to ensure the differentiation of character emotion shaping in the script, avoid the same character emotion development trajectory, or determine whether the difference in emotion change between characters with similar personality traits is too large, the computer device can also determine the emotion comparison result between key characters according to the emotion development curve corresponding to each key character, and the emotion comparison result includes at least one of the similarity of emotion development and the difference of emotion development, so as to perform multi-person emotion quality analysis based on the emotion comparison result between key characters, and obtain the multi-person emotion evaluation result between key characters. Among them, for characters with similar personality but large difference in emotion change, it indicates that the character shaping is illogical; for the main character and the secondary character with similar emotional development trend, it indicates that the character shaping is less differentiated and the character shaping is not distinct enough.

[0172] Indicatively, Figure 11As shown, it shows a schematic diagram of the emotional development curves corresponding to two key characters provided by an exemplary embodiment. Among them, for the first emotional development curve 1101, the amplitude of emotional fluctuations is greater, with obvious cadence in emotional changes, belonging to the emotional shaping of excellent characters; for the second emotional development curve 1102, the amplitude of emotional fluctuations is smaller, and the change in emotional shaping is relatively small, belonging to the emotional shaping of relatively ordinary characters, which needs to be optimized and adjusted.

[0173] In the above embodiment, by applying the emotion recognition model to the field of script analysis, the emotion recognition model outputs the target emotion labels corresponding to each target dialogue text, and then based on the target emotion labels, the emotional development curves of key characters are generated, so that the emotional trends of key characters in the script can be analyzed more intuitively, the quality of character shaping in the script can be evaluated, the script understanding can be assisted, and the efficiency of script creation and modification can be improved.

[0174] Please refer to Figure 12 , which shows a structural block diagram of a training device for an emotion recognition model provided by an exemplary embodiment of the present application. The device includes:

[0175] The first emotion recognition module 1201 is used to perform emotion recognition on the sample texts in the training dataset through the emotion recognition model to obtain the sample emotion recognition results corresponding to each sample text. The training dataset includes the sample texts and the emotion true value labels corresponding to the sample texts;

[0176] The weight determination module 1202 is used to determine the loss adjustment weight corresponding to the sample text according to the sample features of the sample text and the sample emotion recognition result. The sample features include at least one of the sample type and the number of label samples of the emotion true value label corresponding to the sample text. The sample type includes easy samples and difficult samples, and the number of label samples refers to the number of sample texts with the emotion true value label in the training dataset;

[0177] The loss determination module 1203 is used to determine the emotion recognition loss of the sample text based on the loss adjustment weight, the sample emotion recognition result, and the emotion true value label;

[0178] The model training module 1204 is used to train the emotion recognition model based on the emotion recognition loss.

[0179] Optionally, the emotion recognition model is obtained through two-stage training; the model training module 1204 includes:

[0180] A first model training unit, configured to train the emotion recognition model based on a first emotion recognition loss in a first training phase, where the first emotion recognition loss is the emotion recognition loss corresponding to the sample text in a first training dataset, and the sample text in the first training dataset has a single emotion true label;

[0181] A second model training unit, configured to train the emotion recognition model based on a second emotion recognition loss in a second training phase on the basis of the emotion recognition model obtained by training in the first training phase, where the second emotion recognition loss is the emotion recognition loss corresponding to the sample text in a second training dataset, and the sample text in the second training dataset has at least one emotion true label.

[0182] Optionally, in the first training phase, the weight determination module 1202 includes:

[0183] A first weight determination unit, configured to determine a first loss adjustment weight corresponding to the sample text according to the sample type of the sample text and the sample emotion recognition result, where the first loss adjustment weight is used to increase the emotion recognition loss of the difficult samples;

[0184] A second weight determination unit, configured to determine a second loss adjustment weight corresponding to the sample text according to the number of label samples of the emotion true label and the number of first label categories in the first training dataset, where the number of first label categories refers to the total number of label categories included in the first training dataset, and the second loss adjustment weight has a negative correlation with the number of label samples;

[0185] The loss determination module 1203 includes:

[0186] A first loss determination unit, configured to determine the first emotion recognition loss of the sample text based on the first loss adjustment weight, the second loss adjustment weight, the sample emotion recognition result, and the emotion true label.

[0187] Optionally, in the second training phase, the weight determination module 1202 includes:

[0188] A third weight determination unit, configured to determine a first loss adjustment weight corresponding to the sample text according to the sample type of the sample text and the sample emotion recognition result, where the first loss adjustment weight is used to increase the emotion recognition loss of the difficult samples;

[0189] A fourth weight determination unit, configured to determine a second loss adjustment weight corresponding to the sample text according to the number of sample label categories of the sample text, the number of label samples of the emotion true label, and the number of second label categories in the second training dataset;

[0190] The loss determination module 1203 includes:

[0191] A second loss determination unit, configured to determine a second emotion recognition loss of the sample text based on the first loss adjustment weight, the second loss adjustment weight, the sample emotion recognition result, and the emotion true value label.

[0192] Optionally, the fourth weight determination unit is configured to:

[0193] When the number of sample label categories of the sample text is 1, determine the second loss adjustment weight corresponding to the sample text based on the number of label samples of the emotion true value label and the number of second label categories in the second training dataset;

[0194] When the number of sample label categories of the sample text is greater than 1, determine the label weights of each emotion true value label according to the number of label samples of each emotion true value label, and determine the emotion true value label with the least number of label samples as the target emotion true value label; determine the second loss adjustment weight corresponding to the sample text according to the number of second label categories and the weight difference between the label weight of the target emotion true value label and the label weights of other emotion true value labels.

[0195] Optionally, the first emotion recognition module 1201 is configured to:

[0196] Input each sample text in the training dataset into the text feature learning network in the emotion recognition model, and obtain the sample semantic understanding result corresponding to each sample text output by the text feature learning network;

[0197] Input the sample semantic understanding result into the emotion detail learning network in the emotion recognition model, and obtain the sample emotion recognition result corresponding to each sample text output by the emotion detail learning network.

[0198] Optionally, the apparatus further includes:

[0199] A second emotion recognition module, configured to perform emotion recognition on the target dialogue texts of each key character in the script by using the trained emotion recognition model, and determine the target emotion label corresponding to the target dialogue text of each key character, where the script includes at least one script scene and the target dialogue texts of at least two key characters;

[0200] A quality evaluation module, configured to perform quality evaluation on the script based on the target emotion labels corresponding to the target dialogue texts of each key character, and obtain a script quality evaluation result.

[0201] Optionally, the quality assessment module includes:

[0202] A curve generation unit, configured to generate an emotional development curve of each key character in the script scene based on the target emotion labels corresponding to the target dialogue texts of each key character within the same script scene, and the emotion prediction probabilities corresponding to the target emotion labels, where the emotion prediction probabilities are the model output results obtained by performing emotion recognition on the target dialogue texts through the emotion recognition model;

[0203] A quality assessment unit, configured to perform quality assessment on the script based on the emotional development curves corresponding to each key character to obtain the script quality assessment result.

[0204] Optionally, the curve generation unit is configured to:

[0205] In the case where the target dialogue text has a single target emotion label, determine the emotion score of the target dialogue text based on the emotion prediction probability corresponding to the target emotion label;

[0206] In the case where the target dialogue text has at least two target emotion labels, determine the emotional significance degree of each target emotion label based on the emotion prediction probabilities corresponding to each target emotion label; determine the emotion score of the target dialogue text according to the emotion prediction probability corresponding to the target emotion label with the highest emotional significance degree;

[0207] Generate the emotional development curve of the key character in the script scene based on the appearance order of each target dialogue text of the key character in the script scene and the emotion score.

[0208] Optionally, the curve generation unit is further configured to:

[0209] Perform generalization processing on the emotion prediction probabilities corresponding to each target emotion label to obtain the emotional significance degree of each target emotion label, where the emotional significance degree has a positive correlation with the emotion prediction probability; or,

[0210] Determine the emotion offset score corresponding to each target emotion label based on the emotion prediction probabilities corresponding to each target emotion label and a probability decision threshold; determine the emotional significance degree of each target emotion label according to the emotion offset score corresponding to each target emotion label, where the emotional significance degree has a positive correlation with the emotion offset score.

[0211] Optionally, the quality assessment unit is configured to:

[0212] Determine the emotional fluctuation range of each key character in the script scene based on the emotional development curves corresponding to each key character;

[0213] Perform single-person emotion quality analysis on each key person based on the emotional fluctuation amplitude to obtain the single-person emotion evaluation results of each key person;

[0214] Based on the emotional development curves corresponding to each key person, determine the emotional comparison results between key persons, where the emotional comparison results include at least one of emotional development similarity and emotional development difference;

[0215] Perform multi-person emotion quality analysis based on the emotional comparison results between the key persons to obtain the multi-person emotion evaluation results between the key persons;

[0216] Generate the script quality evaluation result based on the single-person emotion evaluation results of each key person and the multi-person emotion evaluation results between the key persons.

[0217] In summary, in the embodiments of the present application, after the emotion recognition model performs emotion recognition on the sample texts in the training dataset to obtain the sample emotion recognition results of each sample text, the emotion recognition loss is not directly determined based on the sample emotion recognition results and the emotion true value labels. Instead, fully considering the sample characteristics of each sample text in the training dataset, first determine the loss adjustment weight according to the sample characteristics of the sample text and the sample emotion recognition result, and then jointly determine the emotion recognition loss by combining the loss adjustment weight, the sample emotion recognition result, and the emotion true value label. Furthermore, train the emotion recognition model based on this emotion recognition loss. By using the solution provided in the embodiments of the present application, by using the loss adjustment weight to determine the emotion recognition loss, it is possible to adjust the emotion recognition loss of easy samples and difficult samples, and in the case where the number of sample texts with different emotion true value labels is unbalanced, adjust the emotion recognition loss of unbalanced label samples, so as to train the emotion recognition model based on the adjusted emotion recognition loss, which can optimize the training effect of the emotion recognition model and improve the accuracy of the output of the emotion recognition model.

[0218] It should be noted that: for the device provided in the above embodiments, only the above division of each functional module is used for illustration. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the device provided in the above embodiments and the method embodiments belong to the same concept, and the implementation process is detailed in the method embodiments, which will not be repeated here.

[0219] It should be noted that before and during the process of obtaining relevant user data such as scripts, online review texts, customer feedback texts, and advertising copy, this application can display a prompt interface, pop-up window, or output voice prompt information. The prompt interface, pop-up window, or voice prompt information is used to prompt the user that their relevant data is currently being collected. This application only starts to execute the relevant steps of obtaining the user's relevant data after obtaining the user's confirmation operation on the prompt interface or pop-up window. Otherwise (that is, when the user's confirmation operation on the prompt interface or pop-up window is not obtained), the relevant steps of obtaining the user's relevant data are terminated, that is, the user's relevant data is not obtained. In other words, the information (including but not limited to user device information, user personal information, etc., and the operation data corresponding to the user), data (including but not limited to data for analysis, stored data, displayed data, etc.), and signals involved in this application are all authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions. For example, data such as scripts, online review texts, customer feedback texts, and advertising copy involved in this application are obtained under full authorization.

[0220] Please refer to Figure 13 , which shows a schematic structural diagram of a computer device provided by an exemplary embodiment of this application. Specifically: The computer device 1300 includes a central processing unit (CPU) 1301, a system memory 1304 including a random access memory 1302 and a read-only memory 1303, and a system bus 1305 connecting the system memory 1304 and the central processing unit 1301. The computer device 1300 may further include a basic input / output system (Input / Output, I / O system) 1306 for facilitating the transmission of information between various components within the computer, and a mass storage device 1307 for storing an operating system 1313, application programs 1314, and other program modules 1315. The application program 1314 may include an application program with the function of training an emotion recognition model.

[0221] In some embodiments, the basic input / output system 1306 includes a display 1308 for displaying information and input devices 1309 such as a mouse, keyboard, etc. for user input of information. Both the display 1308 and the input devices 1309 are connected to the central processing unit 1301 through an input / output controller 1310 connected to the system bus 1305. The basic input / output system 1306 may further include an input / output controller 1310 for receiving and processing inputs from a plurality of other devices such as a keyboard, mouse, or electronic stylus. Similarly, the input / output controller 1310 also provides outputs to a display screen, printer, or other types of output devices.

[0222] The mass storage device 1307 is connected to the central processing unit 1301 through a mass storage controller (not shown) connected to the system bus 1305. The mass storage device 1307 and its associated computer-readable medium provide non-volatile storage for the computer device 1300. That is to say, the mass storage device 1307 may include a computer-readable medium (not shown) such as a hard disk or a drive.

[0223] Without loss of generality, the computer-readable medium may include computer storage media and communication media. Computer storage media includes volatile and non-volatile, removable and non-removable media implemented by any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. Computer storage media includes random access memory (RAM), read-only memory (ROM), flash memory or other solid-state storage technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic tape cartridges, tapes, magnetic disk storage or other magnetic storage devices. Of course, those skilled in the art will know that the computer storage media is not limited to the above several. The above system memory 1304 and mass storage device 1307 may be collectively referred to as memory.

[0224] The memory stores one or more programs, the one or more programs are configured to be executed by one or more central processing units 1301, the one or more programs contain instructions for implementing the above method, and the central processing unit 1301 executes the one or more programs to implement the training method of the emotion recognition model provided by each of the above method embodiments.

[0225] According to various embodiments of the present application, the computer device 1300 may also operate by connecting to a remote computer on a network such as the Internet. That is, the computer device 1300 may be connected to the network 1311 through the network interface unit 1312 connected to the system bus 1305. Or rather, the network interface unit 1312 may also be used to connect to other types of networks or remote computer systems (not shown).

[0226] An embodiment of the present application also provides a computer-readable storage medium, in which at least one instruction is stored, and the at least one instruction is loaded and executed by a processor to implement the training method of the emotion recognition model described in the above embodiment.

[0227] Optionally, the computer-readable storage medium may include: ROM, RAM, solid state drives (SSDs), optical discs, etc. Among them, RAM may include resistive random access memory (ReRAM) and dynamic random access memory (DRAM).

[0228] An embodiment of the present application provides a computer program product, which includes at least one instruction, and the at least one instruction is stored in a computer-readable storage medium. The processor of the computer device reads the at least one instruction from the computer-readable storage medium, and the processor executes the at least one instruction, so that the computer device executes the training method of the emotion recognition model described in the above embodiment.

[0229] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above embodiments can be completed by hardware, or can be completed by a program instructing relevant hardware. The program can be stored in a computer-readable storage medium. The storage medium mentioned above can be a read-only memory, a magnetic disk or an optical disc, etc.

[0230] The above are only optional embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A training method for an emotion recognition model, characterized in that: The method comprises: Performing emotion recognition on sample texts in a training data set by using an emotion recognition model to obtain sample emotion recognition results corresponding to each sample text, wherein the training data set includes the sample texts and emotion truth labels corresponding to the sample texts; According to the sample features of the sample text and the sample emotion recognition result, determine the loss adjustment weight corresponding to the sample text, so that the loss adjustment weight of the difficult sample is greater than the loss adjustment weight of the easy sample, reduce the contribution of the head label to the model loss, and increase the label weight of the tail label in the multi-label sample text, the sample features include at least one of the sample type and the number of label samples of the emotion true value label corresponding to the sample text, the sample type includes the easy sample and the difficult sample, the number of label samples refers to the number of sample texts with the emotion true value label in the training data set, and the number of sample texts corresponding to the head label is more than the number of sample texts corresponding to the tail label; Determining the emotion recognition loss of the sample text based on the loss adjustment weight, the sample emotion recognition result, and the emotion truth label; The emotion recognition model is trained based on the emotion recognition loss.

2. The method according to claim 1, characterized in that The emotion recognition model is obtained through two-stage training; the emotion recognition model is trained based on the emotion recognition loss, including: In a first training stage, the emotion recognition model is trained based on a first emotion recognition loss, where the first emotion recognition loss is the emotion recognition loss corresponding to the sample text in the first training data set, and the sample text in the first training data set has a single emotion truth label; In the second training stage, based on the emotion recognition model trained in the first training stage, the emotion recognition model is trained based on a second emotion recognition loss, where the second emotion recognition loss is the emotion recognition loss corresponding to the sample text in the second training data set, and the sample text in the second training data set has at least one emotion truth label.

3. The method according to claim 2, characterized in that In the first training stage, determining the loss adjustment weight corresponding to the sample text according to the sample features of the sample text and the sample emotion recognition result includes: Determining a first loss adjustment weight corresponding to the sample text according to the sample type of the sample text and the sample emotion recognition result, wherein the first loss adjustment weight is used to increase the emotion recognition loss of the difficult sample; Determining a second loss adjustment weight corresponding to the sample text according to the number of label samples of the true value label of the emotion and the number of first label categories in the first training data set, wherein the first label category number refers to the total number of label categories included in the first training data set, and the second loss adjustment weight is negatively correlated with the number of label samples; The determining the emotion recognition loss of the sample text based on the loss adjustment weight, the sample emotion recognition result and the emotion truth label includes: The first emotion recognition loss of the sample text is determined based on the first loss adjustment weight, the second loss adjustment weight, the sample emotion recognition result, and the emotion truth label.

4. The method according to claim 2, characterized in that: In the second training stage, determining the loss adjustment weight corresponding to the sample text according to the sample features of the sample text and the sample emotion recognition result includes: Determining a first loss adjustment weight corresponding to the sample text according to the sample type of the sample text and the sample emotion recognition result, wherein the first loss adjustment weight is used to increase the emotion recognition loss of the difficult sample; Determine a second loss adjustment weight corresponding to the sample text according to the number of sample label categories of the sample text, the number of label samples of the true emotion label, and the number of second label categories in the second training data set; The determining the emotion recognition loss of the sample text based on the loss adjustment weight, the sample emotion recognition result and the emotion truth label includes: A second emotion recognition loss of the sample text is determined based on the first loss adjustment weight, the second loss adjustment weight, the sample emotion recognition result, and the emotion truth label.

5. The method according to claim 4, characterized in that The determining, according to the number of sample label categories of the sample text, the number of label samples of the true emotion label, and the number of second label categories in the second training data set, a second loss adjustment weight corresponding to the sample text comprises: When the number of sample label categories of the sample text is 1, determining the second loss adjustment weight corresponding to the sample text based on the number of label samples of the true emotion label and the number of the second label categories in the second training data set; When the number of sample label categories of the sample text is greater than 1, the label weight of each emotion truth value label is determined according to the number of label samples of each emotion truth value label, and the emotion truth value label with the least number of label samples is determined as the target emotion truth value label; the second loss adjustment weight corresponding to the sample text is determined according to the second number of label categories and the weight difference between the label weight of the target emotion truth value label and the label weights of other emotion truth value labels.

6. The method according to claim 1, characterized in that The emotion recognition model is used to perform emotion recognition on the sample texts in the training data set to obtain the sample emotion recognition results corresponding to each sample text, including: Inputting each sample text in the training data set into the text feature learning network in the emotion recognition model to obtain a sample semantic understanding result corresponding to each sample text output by the text feature learning network; The sample semantic understanding results are input into the emotion detail learning network in the emotion recognition model to obtain the sample emotion recognition results corresponding to each sample text output by the emotion detail learning network.

7. The method according to claim 1, characterized in that The method further comprises: Using the trained emotion recognition model to perform emotion recognition on target dialogue texts of each key character in the script, and determining target emotion labels corresponding to the target dialogue texts of each key character, wherein the script includes at least one script scene and target dialogue texts of at least two key characters; The script is quality evaluated based on the target emotion labels corresponding to the target dialogue texts of the key characters to obtain a script quality evaluation result.

8. The method according to claim 7, characterized in that The script is evaluated for quality based on the target emotion labels corresponding to the target dialogue texts of the key characters to obtain a script quality evaluation result, including: Generate an emotion development curve of each key character in the same script scene based on the target emotion label corresponding to the target dialogue text of each key character in the same script scene, and the emotion prediction probability corresponding to the target emotion label, wherein the emotion prediction probability is a model output result obtained by performing emotion recognition on the target dialogue text by the emotion recognition model; The script is quality evaluated based on the emotion development curves corresponding to the key characters to obtain the script quality evaluation result.

9. The method according to claim 8, characterized in that The method of generating the emotion development curve of each key character in the same script scene based on the target emotion label corresponding to the target dialogue text of each key character in the same script scene and the emotion prediction probability corresponding to the target emotion label comprises: In the case where the target dialogue text has a single target emotion label, determining the emotion score of the target dialogue text based on the emotion prediction probability corresponding to the target emotion label; In the case where the target dialogue text has at least two target emotion tags, determining the emotion significance of each target emotion tag based on the emotion prediction probability corresponding to each target emotion tag; determining the emotion score of the target dialogue text according to the emotion prediction probability corresponding to the target emotion tag with the highest emotion significance; Based on the order of appearance of each target dialogue text of the key character in the script scene and the emotion score, the emotion development curve of the key character in the script scene is generated.

10. The method according to claim 9, characterized in that The step of determining the emotion significance of each target emotion tag based on the emotion prediction probability corresponding to each target emotion tag includes: Performing generalization processing on the emotion prediction probability corresponding to each target emotion label to obtain the emotion significance of each target emotion label, wherein the emotion significance is positively correlated with the emotion prediction probability; or, Based on the emotion prediction probability and probability judgment threshold corresponding to each target emotion label, the emotion bias score corresponding to each target emotion label is determined; according to the emotion bias score corresponding to each target emotion label, the emotion significance of each target emotion label is determined, and the emotion significance is positively correlated with the emotion bias score.

11. The method according to claim 8, characterized in that The script quality is evaluated based on the emotion development curves corresponding to the key characters to obtain the script quality evaluation result, including: Based on the emotional development curves corresponding to the key characters, determine the emotional fluctuation amplitude of the key characters in the script scenes; Based on the amplitude of the emotion fluctuation, individual emotion quality analysis is performed on each key person to obtain individual emotion evaluation results of each key person; Determine, based on the emotion development curves corresponding to the key characters, an emotion comparison result between the key characters, wherein the emotion comparison result includes at least one of emotion development similarity and emotion development difference; Performing a multi-person emotion quality analysis based on the emotion comparison results between the key persons to obtain a multi-person emotion evaluation result between the key persons; The script quality assessment result is generated based on the single-person emotion assessment results of each key character and the multi-person emotion assessment results among the key characters.

12. A training device for an emotion recognition model, characterized in that: The device comprises: A first emotion recognition module is used to perform emotion recognition on sample texts in a training data set through an emotion recognition model to obtain sample emotion recognition results corresponding to each sample text, wherein the training data set includes the sample texts and emotion truth labels corresponding to the sample texts; A weight determination module, for determining the loss adjustment weight corresponding to the sample text according to the sample features of the sample text and the sample emotion recognition result, so that the loss adjustment weight of the difficult sample is greater than the loss adjustment weight of the easy sample, the contribution of the head label to the model loss is reduced, and the label weight of the tail label in the multi-label sample text is increased, the sample features include at least one of the sample type and the number of label samples corresponding to the emotion true value label of the sample text, the sample type includes the easy sample and the difficult sample, the number of label samples refers to the number of sample texts with the emotion true value label in the training data set, and the number of sample texts corresponding to the head label is more than the number of sample texts corresponding to the tail label; A loss determination module, used to determine the emotion recognition loss of the sample text based on the loss adjustment weight, the sample emotion recognition result and the emotion truth label; A model training module is used to train the emotion recognition model based on the emotion recognition loss.

13. The device according to claim 12, characterized in that The emotion recognition model is obtained through two-stage training; The model training module comprises: A first model training unit, configured to train the emotion recognition model in a first training phase based on a first emotion recognition loss, where the first emotion recognition loss is an emotion recognition loss corresponding to a sample text in a first training data set, and the sample text in the first training data set has a single true emotion label; The second model training unit is used to train the emotion recognition model in a second training stage based on a second emotion recognition loss, based on the emotion recognition model trained in the first training stage, wherein the second emotion recognition loss is the emotion recognition loss corresponding to the sample text in the second training data set, and the sample text in the second training data set has at least one emotion truth label.

14. The device according to claim 13, characterized in that In the first training stage, the weight determination module includes: A first weight determination unit, configured to determine a first loss adjustment weight corresponding to the sample text according to the sample type of the sample text and the sample emotion recognition result, wherein the first loss adjustment weight is used to increase the emotion recognition loss of the difficult sample; a second weight determination unit, configured to determine a second loss adjustment weight corresponding to the sample text according to the number of label samples of the true value label of the emotion and the number of first label categories in the first training data set, wherein the first label category number refers to the total number of label categories included in the first training data set, and the second loss adjustment weight is negatively correlated with the number of label samples; The loss determination module comprises: The first loss determination unit is used to determine the first emotion recognition loss of the sample text based on the first loss adjustment weight, the second loss adjustment weight, the sample emotion recognition result and the emotion truth label.

15. The device according to claim 13, characterized in that In the second training stage, the weight determination module includes: A third weight determination unit, configured to determine a first loss adjustment weight corresponding to the sample text according to the sample type of the sample text and the sample emotion recognition result, wherein the first loss adjustment weight is used to increase the emotion recognition loss of the difficult sample; a fourth weight determination unit, configured to determine a second loss adjustment weight corresponding to the sample text according to the number of sample label categories of the sample text, the number of label samples of the true emotion label, and the number of second label categories in the second training data set; The loss determination module comprises: A second loss determination unit is used to determine a second emotion recognition loss of the sample text based on the first loss adjustment weight, the second loss adjustment weight, the sample emotion recognition result and the emotion truth label.

16. The device according to claim 15, characterized in that The fourth weight determination unit is used to: When the number of sample label categories of the sample text is 1, determining the second loss adjustment weight corresponding to the sample text based on the number of label samples of the true emotion label and the number of the second label categories in the second training data set; When the number of sample label categories of the sample text is greater than 1, determining the label weight of each emotion truth label according to the number of label samples of each emotion truth label, and determining the emotion truth label with the least number of label samples as the target emotion truth label; The second loss adjustment weight corresponding to the sample text is determined according to the second label category number and the weight difference between the label weight of the target emotion truth label and the label weights of other emotion truth labels.

17. The device according to claim 12, characterized in that The first emotion recognition module is used to: Inputting each sample text in the training data set into the text feature learning network in the emotion recognition model to obtain a sample semantic understanding result corresponding to each sample text output by the text feature learning network; The sample semantic understanding results are input into the emotion detail learning network in the emotion recognition model to obtain the sample emotion recognition results corresponding to each sample text output by the emotion detail learning network.

18. The device according to claim 12, characterized in that The device also includes: A second emotion recognition module is used to use the trained emotion recognition model to perform emotion recognition on the target dialogue texts of each key character in the script, and determine the target emotion label corresponding to the target dialogue texts of each key character, wherein the script includes at least one script scene and target dialogue texts of at least two key characters; The quality assessment module is used to perform quality assessment on the script based on the target emotion tags corresponding to the target dialogue texts of the key characters to obtain a script quality assessment result.

19. The device according to claim 18, characterized in that The quality assessment module comprises: A curve generating unit, for generating an emotion development curve of each key character in the same script scene based on a target emotion label corresponding to a target dialogue text of each key character in the same script scene, and an emotion prediction probability corresponding to the target emotion label, wherein the emotion prediction probability is a model output result obtained by performing emotion recognition on the target dialogue text by the emotion recognition model; The quality assessment unit is used to perform quality assessment on the script based on the emotion development curves corresponding to each key character to obtain the script quality assessment result.

20. The device according to claim 19, characterized in that The curve generating unit is used for: In the case where the target dialogue text has a single target emotion label, determining the emotion score of the target dialogue text based on the emotion prediction probability corresponding to the target emotion label; In a case where the target dialogue text has at least two target emotion tags, determining the emotion significance of each target emotion tag based on the emotion prediction probability corresponding to each target emotion tag; Determining the emotion score of the target dialogue text according to the emotion prediction probability corresponding to the target emotion tag with the highest emotion significance; Based on the order of appearance of each target dialogue text of the key character in the script scene and the emotion score, the emotion development curve of the key character in the script scene is generated.

21. The device according to claim 20, characterized in that The curve generating unit is further used for: Generalizing the emotion prediction probability corresponding to each target emotion label to obtain the emotion significance of each target emotion label, wherein the emotion significance is positively correlated with the emotion prediction probability; or, Based on the emotion prediction probability and probability judgment threshold corresponding to each target emotion label, the emotion bias score corresponding to each target emotion label is determined; according to the emotion bias score corresponding to each target emotion label, the emotion significance of each target emotion label is determined, and the emotion significance is positively correlated with the emotion bias score.

22. The device according to claim 19, characterized in that The quality assessment unit is used to: Based on the emotional development curves corresponding to the key characters, determine the emotional fluctuation amplitude of the key characters in the script scenes; Based on the amplitude of the emotion fluctuation, individual emotion quality analysis is performed on each key person to obtain individual emotion evaluation results of each key person; Determine, based on the emotion development curves corresponding to the key characters, an emotion comparison result between the key characters, wherein the emotion comparison result includes at least one of emotion development similarity and emotion development difference; Performing a multi-person emotion quality analysis based on the emotion comparison results between the key persons to obtain a multi-person emotion evaluation result between the key persons; The script quality assessment result is generated based on the single-person emotion assessment results of each key character and the multi-person emotion assessment results among the key characters.

23. A computer device, characterized in that: The computer device includes a processor and a memory; the memory stores at least one instruction, and the at least one instruction is used to be executed by the processor to implement the training method of the emotion recognition model as described in any one of claims 1 to 11.

24. A computer-readable storage medium, characterized in that: The storage medium stores at least one instruction, and the at least one instruction is used to be executed by a processor to implement the training method of the emotion recognition model as described in any one of claims 1 to 11.

25. A computer program product, characterized in that The computer program product includes at least one instruction, which is stored in a computer-readable storage medium; the processor of the computer device reads the at least one instruction from the computer-readable storage medium, and the processor executes the at least one instruction, so that the computer device implements the training method of the emotion recognition model as described in any one of claims 1 to 11.

Citation Information

Patent Citations

  • Training method, system and device for multi-label sentiment classification

    CN116628606A