Emotion Recognition Method, Device, Electronic Device, and Computer-Readable Storage Medium

By establishing a multimodal emotion prediction model, using a conditional variational recurrent neural network with cascade features and a multi-head self-attention mechanism, a multimodal feature vector of the target statement is generated, which solves the problem of failure to effectively identify unknown information in the existing technology, and achieves higher emotional prediction accuracy and efficiency.

CN115099226BActive Publication Date: 2025-07-25NORTHWESTERN POLYTECHNICAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210625611.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-02
Publication Date
2025-07-25
Estimated Expiration
2042-06-02

Smart Images

  • Figure CN115099226B_ABST
    Figure CN115099226B_ABST
Patent Text Reader

Abstract

An embodiment of the present invention discloses an emotion recognition method, device, electronic device, and computer-readable storage medium. The method includes: establishing a multi-modal emotion prediction model, and training the multi-modal emotion prediction model using the multi-modal features of the cascaded historical statement groups and the multi-modal features of the target statement group to obtain a target multi-modal emotion prediction model; processing the target statement to be predicted based on the target multi-modal emotion prediction model to obtain an emotion classification result of the target statement to be predicted. Through the present invention, the technical problem in the related art that most multi-modal emotion prediction methods rely on selecting features and context relationships related to emotion prediction from previous historical information and do not pay attention to the simulation and generation of upcoming dialogue information, making it difficult to effectively improve the recognition accuracy of unknown information, is solved, and the technical effect of effectively improving the accuracy and efficiency of emotion prediction is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of multi-modal sentiment tendency prediction, and particularly to a sentiment recognition method, device, electronic device and computer-readable storage medium. Background Art

[0002] User multi-modal sentiment tendency prediction refers to classifying and predicting unknown emotions that are about to occur through the analysis of known historical information. User multi-modal sentiment tendency prediction is a very popular research field in recent years, with broad development potential and application prospects. For example: autonomous driving fatigue monitoring, airport security monitoring, autism escort and monitoring, smart home care, alarm and monitoring for the elderly and children living alone, public opinion analysis, and mental health analysis.

[0003] In existing multi-modal sentiment prediction technologies, the modalities used for prediction vary according to different research directions, and there are mainly the following four types: visual signals, sound signals, text information, and electroencephalogram signals. Among them, electroencephalogram signals have the relatively highest accuracy, but they must be equipped with corresponding special signal acquisition sensor devices, so it is difficult to widely popularize them conveniently in the daily life field. Therefore, vision, sound, and text are the most common input modalities for multi-modal user sentiment analysis research. In the current related technologies that use these three modalities to predict unknown emotion types that are about to occur, most of them model and identify the unknown dialogue content to be predicted based on the previous historical information. For example, in order to alleviate the cross-modal differences between image visual features and emotional semantic features, an autoencoder is used to learn the joint embedding features of the visual features and semantic features of emotional attributes of the image, narrowing the gap between low-level visual features and high-level semantic features; by introducing an attention model to establish the association between the significant region and the joint embedding features, the significant region related to emotion is determined, the significant region features of the image are extracted, and then the emotion classification of the image is realized based on the significant region features. This algorithm has achieved good emotion classification prediction accuracy on public datasets. However, this method focuses on the extraction of cross-modal key region information and ignores the coherence of context emotional relationships.

[0004] In practical applications, according to the usage method of historical information in the model, it is mainly divided into the following two research directions. One is the extraction and analysis of context sentiment relationships, that is, according to modal features, the transmission characteristics of sentiment among contexts are modeled and analyzed. The other focuses on the analysis of the mutual relationships among multiple modalities, extracts feature information related to the upcoming sentiment, and conducts analysis and prediction. However, the first technology has a good prediction recognition rate when the context mutual relationships are coherent and the feature information noise is small, but in a naturally extracted dialogue database, there are many topic content turns and it cannot maintain such accuracy. The second method has a good tolerance for noise signals in the natural environment, but weakens the extraction of context relationships. At the same time, in the models established by these two research methods, there is no special modeling and fitting for unknown sentiment information, and the information to be predicted is always in an unknown state, only sampling and analyzing historical information.

[0005] Regarding the problem that most of the multi-modal sentiment prediction methods in the above related technologies rely on selecting features and context relationships related to sentiment prediction from previous historical information, and do not pay attention to the simulation and generation of upcoming dialogue information, making it difficult to effectively improve the recognition accuracy of unknown information, no effective solution has been proposed yet. Summary of the Invention

[0006] Embodiments of the present invention provide a sentiment recognition method, device, electronic device, and computer-readable storage medium to at least solve the technical problem that most of the multi-modal sentiment prediction methods in the related technologies rely on selecting features and context relationships related to sentiment prediction from previous historical information, do not pay attention to the simulation and generation of upcoming dialogue information, and are difficult to effectively improve the recognition accuracy of unknown information.

[0007] According to one aspect of the embodiments of the present invention, a sentiment recognition method is provided, including: establishing a multi-modal sentiment prediction model, and training the multi-modal sentiment prediction model using the multi-modal features of the cascaded historical statement group and the multi-modal features of the target statement group to obtain a target multi-modal sentiment prediction model, where the first three sentences of the target statement to be predicted form the historical statement group, the last three sentences of the target statement to be predicted form the target statement group, each sentence of the target statement to be predicted contains multi-modal dialogue information, and the target multi-modal sentiment prediction model is used for the sentiment type recognition of the target statement to be predicted; processing the target statement to be predicted based on the target multi-modal sentiment prediction model to obtain the sentiment classification result of the target statement to be predicted.

[0008] Optionally, the multi-modal sentiment prediction model is trained using the multi-modal features of the concatenated historical statement group and the multi-modal features of the target statement group to obtain a target multi-modal sentiment prediction model, including: extracting the multi-modal features of the historical statement group and the target statement group respectively based on the multi-modal sentiment prediction model to obtain the multi-modal features of the historical statement group and the multi-modal features of the target statement group, and generating a hidden vector of the historical statement group and a hidden vector of the target statement group; determining the multi-modal feature vector of the target statement, wherein, in the training phase, the multi-modal feature vector of the target statement is generated according to the hidden vector of the target statement group and the multi-modal features of the historical statement group; or, in the testing phase, the multi-modal feature vector of the target statement is generated according to the hidden vector of the historical statement group and the multi-modal features of the historical statement group; identifying the multi-modal feature vector of the target statement to obtain a sentiment recognition result, wherein the sentiment recognition result is used as the sentiment classification result of the target statement; optimizing the multi-modal sentiment prediction model based on the backpropagation loss function to obtain the target multi-modal sentiment prediction model.

[0009] Optionally, the multi-modal sentiment prediction model adopts a conditional variational recurrent neural network structure based on a multi-head self-attention mechanism, wherein the conditional variational recurrent neural network structure based on the multi-head self-attention mechanism includes 1 output layer, 1 input layer, and 14 hidden layers, and the hidden layers include 4 recurrent neural network layers, 3 multi-head self-attention layers, 3 activation layers, 3 fully connected layers, and 1 normalization layer.

[0010] Optionally, extracting the multi-modal features of the historical statement group and the target statement group respectively based on the multi-modal sentiment prediction model to obtain the multi-modal features of the historical statement group and the multi-modal features of the target statement group, and generating a hidden vector of the historical statement group and a hidden vector of the target statement group includes: using the input layer to input the multi-modal features of the concatenated historical statement group and the multi-modal features of the target statement group, and the multi-modal features of the historical statement group and the multi-modal features of the target statement group respectively pass through 1 recurrent neural network layer, 1 multi-head attention layer, 1 activation layer, and 1 fully connected layer to complete the generation of the hidden vector, obtaining the hidden vector of the historical statement group and the hidden vector of the target statement group.

[0011] Optionally, determining the multimodal feature vector of the target statement includes: in the training stage, concatenating the hidden vector of the target statement group and the multimodal features of the historical statement group, and inputting the concatenated result into the subsequent fully connected layer and recurrent neural network to generate the multimodal feature vector of the target statement; or, in the testing stage, concatenating the hidden vector of the historical statement group and the multimodal features of the historical statement group, and inputting the concatenated result into the subsequent fully connected layer and recurrent neural network to generate the multimodal feature vector of the target statement.

[0012] Optionally, identifying the multimodal feature vector of the target statement to obtain an emotion recognition result includes: using an identification network to identify the multimodal feature vector of the target statement, and taking the emotion recognition result as the emotion classification result of the target statement, where the identification network includes 1 normalization layer, 1 recurrent neural network layer, 1 multi-head attention layer, 1 activation layer, and 1 fully connected layer.

[0013] Optionally, the loss function at least includes: the relative entropy between the hidden vector of the historical statement group and the hidden vector of the target statement group, the mean absolute error between the multimodal feature vector of the target statement group and the multimodal feature vector of the true target statement group, and the cross entropy between the emotion recognition result and the correct prediction label.

[0014] According to another aspect of the embodiments of the present invention, there is also provided an emotion recognition device, including: a first processing module, configured to establish a multimodal emotion prediction model, and use the multimodal features of the cascaded historical statement group and the multimodal features of the target statement group to train the multimodal emotion prediction model to obtain a target multimodal emotion prediction model, where the first three sentences of the target statement to be predicted constitute the historical statement group, the last three sentences of the target statement to be predicted constitute the target statement group, each sentence of the target statement to be predicted contains multimodal dialogue information, and the target multimodal emotion prediction model is used for emotion type recognition of the target statement to be predicted; a second processing module, configured to process the target statement to be predicted based on the target multimodal emotion prediction model to obtain the emotion classification result of the target statement to be predicted.

[0015] According to another aspect of the embodiments of the present invention, there is also provided an electronic device, including: a processor; a memory for storing processor-executable instructions; wherein, the processor is configured to execute the method steps described in any one of the above.

[0016] According to another aspect of the embodiments of the present invention, there is also provided a computer-readable storage medium, which includes a stored program. When the program runs, it controls the device where the computer-readable storage medium is located to execute the method steps described in any one of the above.

[0017] In the embodiments of the present invention, a multi-modal emotion prediction model is established, and the multi-modal features of the cascaded historical sentence group and the multi-modal features of the target sentence group are used to train the multi-modal emotion prediction model to obtain a target multi-modal emotion prediction model. Among them, the first three sentences of the target sentence to be predicted form the historical sentence group, the last three sentences of the target sentence to be predicted form the target sentence group, and each sentence of the target sentence to be predicted contains multi-modal dialogue information. The target multi-modal emotion prediction model is used for emotion type recognition of the target sentence to be predicted; the target sentence to be predicted is processed based on the target multi-modal emotion prediction model to obtain the emotion classification result of the target sentence to be predicted. That is to say, the embodiments of the present invention make up for the information gap between historical information and unknown information. By establishing a multi-modal emotion prediction model, using the multi-modal features of the cascaded historical sentence group and the multi-modal features of the target sentence group to train the multi-modal emotion prediction model, and then using the trained target multi-modal emotion prediction model to achieve the emotion classification of the target sentence to be predicted, thereby solving the technical problem that most of the multi-modal emotion prediction methods in the related art rely on selecting features and context relationships related to emotion prediction from the previous historical information, and do not pay attention to the simulation and generation of upcoming dialogue information, and it is difficult to effectively improve the recognition accuracy of unknown information, and achieving the technical effect of effectively improving the accuracy and efficiency of emotion prediction. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of this application. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:

[0019] Figure 1 It is a flowchart of an emotion recognition method provided by an embodiment of the present invention;

[0020] Figure 2 It is a schematic diagram of a conditional variational recurrent neural network structure based on a multi-head self-attention mechanism provided by an embodiment of the present invention;

[0021] Figure 3 It is a flowchart of a multi-modal emotion recognition method based on a conditional variational recurrent neural network and a multi-head attention mechanism provided by an embodiment of the present invention;

[0022] Figure 4(a) is a schematic diagram of the pre-distribution of a text, image, and speech information difference input model for a historical statement group and a target statement group provided by an embodiment of the present invention;

[0023] Figure 4(b) is a schematic diagram of the result of classification completed through multi-modal feature extraction and fusion provided by an embodiment of the present invention;

[0024] Figure 5 It is a schematic diagram of an emotion recognition device provided by an embodiment of the present invention. Detailed implementation manners

[0025] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without making creative efforts shall fall within the protection scope of the present invention.

[0026] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the drawings are used to distinguish different objects, rather than to limit a specific order. In addition, the steps shown in the flowchart of the drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0027] Figure 1 It is a flowchart of an emotion recognition method provided by an embodiment of the present invention. As Figure 1 shown, the method includes the following steps:

[0028] Step S102, establish a multi-modal emotion prediction model, and use the multi-modal features of the cascaded historical statement group and the multi-modal features of the target statement group to train the multi-modal emotion prediction model to obtain a target multi-modal emotion prediction model, where the first three sentences of the target statement to be predicted form the historical statement group, the last three sentences of the target statement to be predicted form the target statement group, each sentence of the target statement to be predicted contains multi-modal dialogue information, and the target multi-modal emotion prediction model is used for the emotion type recognition of the target statement to be predicted;

[0029] It should be noted that in multi-modal sentiment prediction, every three dialogue sentences form a sentence group. The first three sentences of the target sentence to be predicted form the historical sentence group, and the last three sentences of the target sentence to be predicted form the target sentence group. The multi-modal dialogue information contained in each sentence is cascaded to generate a multi-modal feature vector, which includes but is not limited to the multi-modal features of the historical sentence group and the multi-modal features of the target sentence group.

[0030] The above multi-modal dialogue information includes but is not limited to images, texts, voices, etc.; the above sentiment types include but are not limited to happiness, excitement, frustration, sadness, anger, and neutrality, etc.

[0031] Step S104, process the target sentence to be predicted based on the target multi-modal sentiment prediction model to obtain the sentiment classification result of the target sentence to be predicted.

[0032] In the embodiment of the present invention, a multi-modal sentiment prediction model is established, and the multi-modal features of the cascaded historical sentence group and the multi-modal features of the target sentence group are used to train the multi-modal sentiment prediction model to obtain the target multi-modal sentiment prediction model. Among them, the first three sentences of the target sentence to be predicted form the historical sentence group, the last three sentences of the target sentence to be predicted form the target sentence group, and each sentence of the target sentence to be predicted contains multi-modal dialogue information. The target multi-modal sentiment prediction model is used for identifying the sentiment type of the target sentence to be predicted; the target sentence to be predicted is processed based on the target multi-modal sentiment prediction model to obtain the sentiment classification result of the target sentence to be predicted. That is to say, the embodiment of the present invention makes up for the information gap between historical information and unknown information. By establishing a multi-modal sentiment prediction model, using the multi-modal features of the cascaded historical sentence group and the multi-modal features of the target sentence group to train the multi-modal sentiment prediction model, and then using the trained target multi-modal sentiment prediction model to achieve the sentiment classification of the target sentence to be predicted, thereby solving the technical problem that most of the multi-modal sentiment prediction methods in the related art rely on selecting features and context relationships related to sentiment prediction from the previous historical information and do not pay attention to the simulation and generation of upcoming dialogue information, and it is difficult to effectively improve the recognition accuracy of unknown information, and achieving the technical effect of effectively improving the accuracy and efficiency of sentiment prediction.

[0033] It should be noted that the application scenarios of the above method include but are not limited to user sentiment tendency prediction.

[0034] In an alternative embodiment, the multi-modal sentiment prediction model is trained using the multi-modal features of the cascaded historical statement group and the multi-modal features of the target statement group to obtain the target multi-modal sentiment prediction model, including: extracting the multi-modal features of the historical statement group and the target statement group respectively based on the multi-modal sentiment prediction model to obtain the multi-modal features of the historical statement group and the multi-modal features of the target statement group, and generating the hidden vector of the historical statement group and the hidden vector of the target statement group; determining the multi-modal feature vector of the target statement, wherein, in the training stage, the multi-modal feature vector of the target statement is generated according to the hidden vector of the target statement group and the multi-modal features of the historical statement group; or, in the test stage, the multi-modal feature vector of the target statement is generated according to the hidden vector of the historical statement group and the multi-modal features of the historical statement group; identifying the multi-modal feature vector of the target statement to obtain the sentiment recognition result, wherein the sentiment recognition result is used as the sentiment classification result of the target statement; optimizing the multi-modal sentiment prediction model based on the backpropagation loss function to obtain the target multi-modal sentiment prediction model.

[0035] In the above embodiment of the present invention, the multi-modal sentiment prediction model combining the multi-head self-attention mechanism and the conditional variational recurrent neural network can effectively generate the multi-modal feature information of the target statement according to the historical statement group and perform sentiment recognition based on the generated multi-modal information. Compared with the traditional feature extraction method, the multi-head self-attention mechanism can better extract the context relationship, and the conditional variational recurrent neural network can transform the unknown prediction problem into the recognition problem of known feature information, improving the overall prediction accuracy. Finally, aiming at the sentiment classification recognition error, the gap between the two groups of hidden vectors, the generated multi-modal information of the target statement and the real multi-modal information of the target statement is reduced through the backpropagation loss function, thereby improving the final sentiment recognition accuracy.

[0036] It should be noted that the above conditional variational recurrent neural network process is divided into two stages: training and testing. In the training stage, the multi-modal features of the real target statement group participate in the training for the network to learn and calculate the loss function. In the testing stage, only the multi-modal features of the historical statements participate in the network calculation to generate the corresponding multi-modal features of the target statement by itself.

[0037] In an alternative embodiment, the above multi-modal sentiment prediction model adopts a conditional variational recurrent neural network structure based on the multi-head self-attention mechanism, wherein the conditional variational recurrent neural network structure based on the multi-head self-attention mechanism includes 1 output layer, 1 input layer and 14 hidden layers. The hidden layers include 4 recurrent neural network layers, 3 multi-head self-attention layers, 3 activation layers, 3 fully connected layers and 1 normalization layer.

[0038] In order to achieve more reasonable and effective multi-modal sentiment prediction, the present invention proposes a novel conditional variational recurrent neural network structure based on the multi-head self-attention mechanism. Based on the required known historical multi-modal information, the conditional variational recurrent neural network is used to extract relevant multi-modal features, and the context relationship is extracted through the multi-head self-attention mechanism. The target sentence to be predicted is modeled and fitted, and then the possible corresponding multi-modal feature distribution of the target sentence is generated. Finally, the recognition network performs targeted sentiment type recognition on the generated multi-modal features to complete the training and optimization of sentiment prediction.

[0039] In addition, to bridge the information gap between historical information and unknown information, through the conditional variational autoencoder and the multi-head self-attention mechanism, according to the multi-modal features and context information of historical information, feature fitting and generation are performed on the possible information to be predicted, converting unknown information into known information, and then targeted sentiment recognition and classification are performed based on the multi-modal features of the generated known information to be predicted. That is, the prediction of unknown information is transformed into two small sub-problems of information generation and sentiment recognition. Using the multi-modal sentiment prediction model based on the conditional variational recurrent neural network structure with the multi-head self-attention mechanism effectively improves the accuracy and efficiency of prediction.

[0040] It should be noted that the above-mentioned conditional variational autoencoder, as a deep latent space generation model, can jointly generate a latent vector that contains both data information and noise from the original data and the specified conditional information. Then, various new data can be generated according to requirements by combining this latent vector with the conditional information. Compared with the ordinary attention mechanism, the above-mentioned multi-head self-attention mechanism has strong parallel computing capabilities and has good effects in dealing with problems with sequence signals, so it is widely used in natural language processing problems. Subsequently, due to the good interpretability of the multi-head self-attention mechanism, the output process of the network is more in line with the intuitive cognition of researchers, so it is increasingly widely used in other neural networks. By combining the conditional variational autoencoder and the multi-head self-attention mechanism, on the basis of better extracting context relationships and multi-modal feature information, it is possible to generate a target sentence feature vector that is closer to the prediction information as much as possible.

[0041] In an alternative embodiment, multi-modal features of the historical statement group and the target statement group are respectively extracted based on the multi-modal sentiment prediction model to obtain the multi-modal features of the historical statement group and the target statement group, and hidden vectors of the historical statement group and the target statement group are generated, including: inputting the multi-modal features of the concatenated historical statement group and the target statement group through the input layer, and respectively passing the multi-modal features of the historical statement group and the target statement group through 1 recurrent neural network layer, 1 multi-head attention layer, 1 activation layer, and 1 fully connected layer to complete the generation of the hidden vectors, thereby obtaining the hidden vectors of the historical statement group and the target statement group.

[0042] In the specific implementation process, the input layer of the above neural network is the multi-modal features of the concatenated historical statement group and the target statement group. The two groups of multi-modal features respectively pass through 1 recurrent neural network layer, 1 multi-head attention layer, 1 activation layer, and 1 fully connected layer to complete the generation of the hidden vectors, and the dimension of the hidden vectors is 64. Among them, the two recurrent neural network layers input separately share parameters.

[0043] In the above embodiment of the present invention, through the fusion and extraction of the multi-modal features, the generation of two groups of hidden vectors for conditional variational is completed.

[0044] In an alternative embodiment, determining the multi-modal feature vector of the target statement includes: in the training stage, concatenating the hidden vector of the target statement group and the multi-modal features of the historical statement group, and inputting them into the subsequent fully connected layer and recurrent neural network to generate the multi-modal feature vector of the target statement, thereby obtaining the multi-modal feature vector of the target statement; or, in the testing stage, concatenating the hidden vector of the historical statement group and the multi-modal features of the historical statement group, and inputting them into the subsequent fully connected layer and recurrent neural network to generate the multi-modal feature vector of the target statement, thereby obtaining the multi-modal feature vector of the target statement.

[0045] In the training stage, the three groups of multi-modal feature vectors of the historical statement group and the target statement group are respectively subjected to feature fusion and extraction through two recurrent neural networks and multi-head self-attention layers sharing parameters to respectively generate the hidden vectors of the historical statement group and the target statement group. Then, the hidden vector of the target statement group is concatenated with the multi-modal feature information of the historical statement group, and the feature information of the target statement group is fitted and generated through a new recurrent neural network. Finally, the generated fitted multi-modal feature information of the target statement group completes the sentiment classification of the target statement through the third recurrent neural network and multi-head attention mechanism;

[0046] In the testing phase, instead of using the information of the target statement group to generate the hidden vector of the target statement, the hidden vectors of the generated historical statement group and the multimodal features of the historical statement group are concatenated and input into the next recurrent neural network for generating the feature information of the target statement and the corresponding sentiment recognition.

[0047] In an alternative embodiment, the multimodal feature vector of the target statement is recognized to obtain the sentiment recognition result, including: using a recognition network to recognize the multimodal feature vector of the target statement and taking the sentiment recognition result as the sentiment classification result of the target statement, where the recognition network includes 1 layer of normalization layer, 1 layer of recurrent neural network layer, 1 layer of multi-head attention layer, 1 layer of activation layer and 1 layer of fully connected layer.

[0048] In the specific implementation process, a sentiment recognition algorithm can be used to process the multimodal feature vector of the target statement to obtain the sentiment recognition result corresponding to the multimodal feature vector of the target statement.

[0049] In an alternative embodiment, the above loss function at least includes: the relative entropy between the hidden vector of the historical statement group and the hidden vector of the target statement group, the mean absolute error between the multimodal feature vector of the target statement group and the multimodal feature vector of the true target statement group, and the cross entropy between the sentiment recognition result and the correct prediction label.

[0050] In order to make the hidden vector of the historical statement group fit the hidden vector of the target statement group as much as possible, in the training phase, the loss function is divided into three parts. One part is the calculation of the relative entropy (information divergence) between the hidden vector of the historical statement group and the hidden vector of the target statement, which is a measure of the asymmetry of the probability distribution difference between the two hidden vectors, in order to reduce the gap between the two through backpropagation; one part is the reconstruction error of the generated multimodal information of the target statement and the true multimodal information of the target statement; the last part is the sentiment classification recognition error. Through the backpropagation loss function of the three parts, the gap between the two groups of hidden vectors, the generated multimodal information of the target statement and the true multimodal information of the target statement is reduced simultaneously, so as to improve the final sentiment recognition accuracy.

[0051] A detailed description of an alternative embodiment of the present invention will be given below.

[0052] To achieve multimodal sentiment prediction, the present invention proposes a multimodal sentiment prediction model based on conditional variational recurrent neural network and multi-head attention mechanism. The conditional variational recurrent neural network is used to extract the required multimodal feature information for the network model, and the multi-head self-attention mechanism is used to strengthen the extraction of context relationships. With the true classification label being six basic sentiment types (happy, excited, frustrated, sad, angry and neutral), the training and optimization of the sentiment prediction of the target statement are completed.

[0053] Figure 2 Schematic diagram of a conditional variational recurrent neural network structure based on the multi-head self-attention mechanism provided by an embodiment of the present invention, as Figure 2 shown. The conditional variational recurrent neural network structure based on the multi-head self-attention mechanism has a total of 16 layers, including 1 output layer, 1 input layer, and 14 hidden layers, including 4 recurrent neural network layers, 3 multi-head self-attention layers, 3 activation layers, 3 fully connected layers, and 1 normalization layer. First, through the fusion and extraction of multi-modal features, two groups of hidden vectors for conditional variation are generated. The neural network input is the multi-modal features of the concatenated historical statement group and the multi-modal features of the target statement group. The two groups of multi-modal features respectively pass through one recurrent neural network layer, one multi-head attention layer, one activation layer, and one fully connected layer to complete the generation of hidden vectors, and the dimension of the hidden vectors is 64. Among them, the two recurrent neural network layers input respectively share parameters. Next, the multi-modal features of the target statement group are generated. In the training stage, the hidden vectors of the target statement group and the multi-modal features of the historical statement group are concatenated and input into the subsequent fully connected layer and recurrent neural network to generate the multi-modal feature vectors of the target statement; while in the test stage, the hidden vectors of the historical statement group and the multi-modal features of the historical statement group are concatenated and input into the subsequent fully connected layer and recurrent neural network to generate the multi-modal feature vectors of the target statement. Finally, the recognition network performs sentiment type recognition on the generated multi-modal features of the target statement. The recognition network includes one normalization layer, one recurrent neural network layer, one activation layer, and one fully connected layer, and the recognition result is used as the final sentiment prediction classification result of the target statement.

[0054] Compared with the existing multi-modal sentiment prediction models, in the multi-modal sentiment prediction model of the present invention, through the multi-head attention mechanism and the conditional variational recurrent neural network structure, the generation of multi-modal feature information of the target statement group is realized, the multi-modal features required for the prediction target are changed from unknown information to known information, and the prediction problem is changed into a generation and recognition problem, so as to realize effective multi-modal sentiment prediction.

[0055] Figure 3 Flowchart of a multi-modal sentiment recognition method based on a conditional variational recurrent neural network and a multi-head attention mechanism provided by an embodiment of the present invention, as Figure 3 shown. The specific implementation manner includes the following steps:

[0056] Step 1: Build a network model based on the conditional variational neural network and the multi-head self-attention mechanism, and train the model using the gradient descent and backpropagation algorithms. The specific process is as follows:

[0057] (1) Construct a network model based on conditional variational neural network and multi-head self-attention mechanism, and initialize all parameters and weights with random numbers. Represent the input of two groups of multimodal information as:

[0058] S_hist = [V h , T h , A h

[0059] S_tar = [V t , T t , A t

[0060] Among them, S_hist is the multimodal feature of the historical statement group, S_tar is the multimodal feature of the target statement group, V is the image information, T is the text information, and A is the voice information.

[0061] (2) Train the model on the multimodal dialogue information database. For the image, text, and voice modal information contained in the information to be recognized, calculate the hidden vector expressions of the two groups of multimodal features obtained after passing through the recurrent neural network layer, multi-head self-attention mechanism, and fully connected layer, which are Z hist and Z tar , respectively. The specific expressions are:

[0062] Z hist = F Latent (S_hist, W h )

[0063] Z tar = F Latent (S_tar, W t )

[0064] Among them, F Latent is the constructed hidden vector generation algorithm model, and W h and W t are the parameter variables involved in the model. Then, in the training stage, the hidden vector Z tar of the target statement group and the multimodal features of the historical statement group are used to generate the multimodal feature vector of the target statement; while in the test stage, the hidden vector Z hist of the historical statement group and the multimodal features of the historical statement group are used to generate the multimodal feature vector of the target statement. The formula is as follows:

[0065] Train: Rec_S_tar = F Generator (Z_tar, S_hist, W G )

[0066] Test: Rec_S_tar = F Generator ​​(Z_hist, S_hist, W G )

[0067] Among them, Rec_S_tar is the multi-modal feature vector of the generated target sentence, and F Generator is the generation algorithm for reconstructing the multi-modal features of the target sentence, and W G is the parameter variable involved in the algorithm. Finally, the sentiment type of the multi-modal features of the reconstructed generated target sentence group is recognized. The formula is as follows:

[0068] Classification = F Class (Rec_S_tar, W C )

[0069] Among them, Classification is the output result of the model, and F Class is the sentiment recognition algorithm, and W C is the parameter variable involved in the algorithm.

[0070] Next, the loss function is calculated. The loss function consists of three parts. One part is the relative entropy Loss_KL of two groups of hidden vectors:

[0071] Loss_KL = KL(Z_hist, Z_tar)

[0072] Among them, KL is the relative entropy calculation formula. The second part of the loss function Loss_Rec is the mean absolute error between the multi-modal features of the reconstructed generated target sentence group and the multi-modal features of the true target sentence group. MAE is the mean absolute error calculation formula:

[0073] Loss_Rec = MAE(Rec_S_tar, S_tar)

[0074] The third part is the cross-entropy between the model output sentiment recognition result Classification and the correct prediction label Label:

[0075] Loss_Class = CrossEntrophy(Classification, Label)

[0076] Among them, CrossEntrophy is the cross-entropy calculation.

[0077] Finally, the loss function Loss is:

[0078] Loss = Loss_KL + Loss_Rec + Loss_Class

[0079] The model is trained by backpropagating the loss function Loss.

[0080] Step 2: Use the data in the dataset that has not been trained as test instances, and use the multi-modal sentiment prediction model based on conditional variational neural network and multi-head self-attention mechanism for calculation to obtain the final classification result.

[0081] The present invention has been verified for effectiveness on the multi-modal sentiment analysis public dataset IEMOCAP. After the IEMOCAP dataset is reorganized by sentences, the training set contains 5,450 samples, and the test set contains 1,530 samples, meeting the training-test ratio of 3:1. Each sample contains text, image, and speech information, and the labels are divided into six categories, namely happy, excited, frustrated, sad, angry, and neutral. The evaluation index is F1-score, and a significance test is carried out through a T-test with a significance level of 0.05. The number of hidden nodes in the recurrent neural network layer of the established model is 512, the number of heads in the multi-head self-attention mechanism is 8, the number of hidden nodes is 1,024, the number of output nodes is 512, the number of hidden nodes in the three fully connected layers are 64, 712, and 6 respectively, and the learning rate is 0.0015.

[0082] Figure 4(a) is a schematic diagram of the distribution of the differences in text, image, and speech information between the historical sentence group and the target sentence group provided by the embodiment of the present invention before inputting into the model. As shown in Figure 4(a), the difference values of the speech, text, and image multi-modal information of the input layer's real historical sentence group and target sentence group are presented. The abscissa is the feature serial number, and the ordinate is the size of the feature value difference. Figure 4(b) is a schematic diagram of the result of completing classification after multi-modal feature extraction and fusion provided by the embodiment of the present invention. As shown in Figure 4(b), in the test stage, after the multi-modal feature information of the historical sentence group passes through the network of this model, the generated multi-modal features of the target sentence group are recognized for the sentiment classification result. It can be seen from the figure that there are considerable differences in the speech, text, and image multi-modal feature information before model training, but after model training, the six emotion types of the unknown target sentence to be predicted can be effectively distinguished. For example, the F1-score is 58.57%, and the classification accuracy is 59.12%, which proves the effectiveness of the above method of the present invention.

[0083] According to another aspect of the embodiment of the present invention, there is also provided an emotion recognition device. Figure 5 It is a schematic diagram of an emotion recognition device provided by the embodiment of the present invention. As Figure 5 shown, the emotion recognition device includes: a first processing module 52 and a second processing module 54. The following is a detailed description of the emotion recognition device.

[0084] The first processing module 52 is configured to establish a multi-modal emotion prediction model, and train the multi-modal emotion prediction model by using the multi-modal features of the cascaded historical statement group and the multi-modal features of the target statement group, so as to obtain a target multi-modal emotion prediction model. Among them, the first three sentences of the target statement to be predicted form the historical statement group, the last three sentences of the target statement to be predicted form the target statement group, each sentence of the target statement to be predicted contains multi-modal dialogue information, and the target multi-modal emotion prediction model is used for emotion type recognition of the target statement to be predicted; the second processing module 54 is connected to the above-mentioned first processing module 52, and is configured to process the target statement to be predicted based on the target multi-modal emotion prediction model to obtain the emotion classification result of the target statement to be predicted.

[0085] In the embodiment of the present invention, the emotion recognition device adopts the method of establishing a multi-modal emotion prediction model, and trains the multi-modal emotion prediction model by using the multi-modal features of the cascaded historical statement group and the multi-modal features of the target statement group, so as to obtain a target multi-modal emotion prediction model. Among them, the first three sentences of the target statement to be predicted form the historical statement group, the last three sentences of the target statement to be predicted form the target statement group, each sentence of the target statement to be predicted contains multi-modal dialogue information, and the target multi-modal emotion prediction model is used for emotion type recognition of the target statement to be predicted; process the target statement to be predicted based on the target multi-modal emotion prediction model to obtain the emotion classification result of the target statement to be predicted. That is to say, the embodiment of the present invention makes up for the information gap between historical information and unknown information. By establishing a multi-modal emotion prediction model, training the multi-modal emotion prediction model by using the multi-modal features of the cascaded historical statement group and the multi-modal features of the target statement group, and then using the obtained target multi-modal emotion prediction model after training to realize the emotion classification of the target statement to be predicted, thereby solving the technical problem that most of the multi-modal emotion prediction methods in the related art rely on selecting features and context relationships related to emotion prediction from the previous historical information, and do not pay attention to the simulation and generation of upcoming dialogue information, and it is difficult to effectively improve the recognition accuracy of unknown information, and achieving the technical effect of effectively improving the accuracy and efficiency of emotion prediction.

[0086] It should be noted here that the above-mentioned first processing module 52 and second processing module 54 correspond to steps S102 to S104 in the method embodiment. The examples and application scenarios implemented by the above-mentioned modules and the corresponding steps are the same, but are not limited to the content disclosed in the above-mentioned method embodiment.

[0087] Optionally, the first processing module 52 described above includes: a generating unit configured to respectively extract multi-modal features of the historical statement group and the target statement group based on the multi-modal sentiment prediction model, obtain the multi-modal features of the historical statement group and the multi-modal features of the target statement group, and generate a hidden vector of the historical statement group and a hidden vector of the target statement group; a determining unit configured to determine a multi-modal feature vector of the target statement, wherein, in the training phase, the multi-modal feature vector of the target statement is generated according to the hidden vector of the target statement group and the multi-modal features of the historical statement group; or, in the testing phase, the multi-modal feature vector of the target statement is generated according to the hidden vector of the historical statement group and the multi-modal features of the historical statement group; an identifying unit configured to identify the multi-modal feature vector of the target statement to obtain a sentiment recognition result, wherein the sentiment recognition result is used as the sentiment classification result of the target statement; an optimizing unit configured to optimize the multi-modal sentiment prediction model based on the backpropagation loss function to obtain a target multi-modal sentiment prediction model.

[0088] Optionally, the multi-modal sentiment prediction model described above adopts a conditional variational recurrent neural network structure based on a multi-head self-attention mechanism. The conditional variational recurrent neural network structure based on the multi-head self-attention mechanism includes 1 output layer, 1 input layer, and 14 hidden layers. The hidden layers include 4 recurrent neural network layers, 3 multi-head self-attention layers, 3 activation layers, 3 fully connected layers, and 1 normalization layer.

[0089] Optionally, the generating unit described above includes: a first generating subunit configured to use the multi-modal features of the historical statement group and the multi-modal features of the target statement group after input layer input concatenation. The multi-modal features of the historical statement group and the multi-modal features of the target statement group respectively pass through 1 recurrent neural network layer, 1 multi-head attention layer, 1 activation layer, and 1 fully connected layer to complete the generation of the hidden vector, obtaining the hidden vector of the historical statement group and the hidden vector of the target statement group.

[0090] Optionally, the determining unit described above includes: a second generating subunit configured to, in the training phase, concatenate the hidden vector of the target statement group and the multi-modal features of the historical statement group, and input the concatenated result into the subsequent fully connected layer and recurrent neural network to generate the multi-modal feature vector of the target statement, obtaining the multi-modal feature vector of the target statement; or, a third generating subunit configured to, in the testing phase, concatenate the hidden vector of the historical statement group and the multi-modal features of the historical statement group, and input the concatenated result into the subsequent fully connected layer and recurrent neural network to generate the multi-modal feature vector of the target statement, obtaining the multi-modal feature vector of the target statement.

[0091] Optionally, the above recognition unit includes: a recognition subunit, configured to recognize the multimodal feature vector of the target statement using a recognition network, and use the sentiment recognition result as the sentiment classification result of the target statement, where the recognition network includes one layer of normalization layer, one layer of recurrent neural network layer, one layer of multi-head attention layer, one layer of activation layer, and one layer of fully connected layer.

[0092] Optionally, the above loss function at least includes: the relative entropy between the hidden vectors of the historical statement group and the hidden vectors of the target statement group, the mean absolute error between the multimodal feature vectors of the target statement group and the multimodal feature vectors of the true target statement group, and the cross entropy between the sentiment recognition result and the correct prediction label.

[0093] According to another aspect of the embodiments of the present invention, there is also provided an electronic device, including: a processor; a memory for storing processor-executable instructions; wherein, the processor is configured to execute the method steps of any one of the above.

[0094] According to another aspect of the embodiments of the present invention, there is also provided a computer-readable storage medium, the computer-readable storage medium including a stored program, wherein, when the program runs, it controls the device where the computer-readable storage medium is located to execute the method steps of any one of the above.

[0095] The above are only the preferred embodiments of the present invention, and are not intended to limit the protection scope of the present invention.

Claims

1. An emotion recognition method, characterized in that, Including: Build a multi-modal sentiment prediction model, and use the multi-modal features of the cascaded historical statement group and the multi-modal features of the target statement group to train the multi-modal sentiment prediction model to obtain a target multi-modal sentiment prediction model. Among them, the first three sentences of the target statement to be predicted constitute the historical statement group, the last three sentences of the target statement to be predicted constitute the target statement group, each sentence of the target statement to be predicted contains multi-modal dialogue information, and the target multi-modal sentiment prediction model is used for the sentiment type recognition of the target statement to be predicted; The step of using the multi-modal features of the cascaded historical statement group and the multi-modal features of the target statement group to train the multi-modal sentiment prediction model to obtain a target multi-modal sentiment prediction model includes: Based on the multi-modal sentiment prediction model, extract the multi-modal features of the historical statement group and the target statement group respectively, obtain the multi-modal features of the historical statement group and the multi-modal features of the target statement group, and generate the hidden vector of the historical statement group and the hidden vector of the target statement group; Determine the multi-modal feature vector of the target statement. Among them, in the training stage, generate the multi-modal feature vector of the target statement according to the hidden vector of the target statement group and the multi-modal features of the historical statement group; or, in the testing stage, generate the multi-modal feature vector of the target statement according to the hidden vector of the historical statement group and the multi-modal features of the historical statement group; Identify the multi-modal feature vector of the target statement to obtain a sentiment recognition result, where the sentiment recognition result is used as the sentiment classification result of the target statement; Optimize the multi-modal sentiment prediction model based on the backpropagation loss function to obtain the target multi-modal sentiment prediction model; The multi-modal sentiment prediction model adopts a conditional variational recurrent neural network structure based on the multi-head self-attention mechanism. Among them, the conditional variational recurrent neural network structure based on the multi-head self-attention mechanism includes 1 output layer, 1 input layer, and 14 hidden layers. The hidden layers include 4 recurrent neural network layers, 3 multi-head self-attention layers, 3 activation layers, 3 fully connected layers, and 1 normalization layer; Based on the target multi-modal sentiment prediction model, process the target statement to be predicted to obtain the sentiment classification result of the target statement to be predicted.

2. The method according to claim 1, characterized in that Based on the multi-modal sentiment prediction model, extract the multi-modal features of the historical statement group and the target statement group respectively, obtain the multi-modal features of the historical statement group and the multi-modal features of the target statement group, and generate the hidden vector of the historical statement group and the hidden vector of the target statement group, including: Use the input layer to input the multi-modal features of the cascaded historical statement group and the multi-modal features of the target statement group. The multi-modal features of the historical statement group and the multi-modal features of the target statement group respectively pass through 1 recurrent neural network layer, 1 multi-head self-attention layer, 1 activation layer, and 1 fully connected layer to complete the generation of hidden vectors, and obtain the hidden vector of the historical statement group and the hidden vector of the target statement group.

3. The method according to claim 1, characterized in that, Determine the multi-modal feature vector of the target statement, including: In the training phase, concatenate the hidden vector of the target statement group and the multi-modal features of the historical statement group, and input them into the subsequent fully connected layer and recurrent neural network to generate the multi-modal feature vector of the target statement, obtaining the multi-modal feature vector of the target statement; or, In the testing phase, concatenate the hidden vector of the historical statement group and the multi-modal features of the historical statement group, and input them into the subsequent fully connected layer and recurrent neural network to generate the multi-modal feature vector of the target statement, obtaining the multi-modal feature vector of the target statement.

4. The method according to claim 1, wherein Identify the multi-modal feature vector of the target statement to obtain an emotion recognition result, including: Use an identification network to identify the multi-modal feature vector of the target statement, and use the emotion recognition result as the emotion classification result of the target statement, where the identification network includes 1 layer of the normalization layer, 1 layer of the recurrent neural network layer, 1 layer of the multi-head self-attention layer, 1 layer of the activation layer, and 1 layer of the fully connected layer.

5. The method according to claim 1, characterized in that, The loss function at least includes: The relative entropy between the hidden vector of the historical statement group and the hidden vector of the target statement group, the mean absolute error between the multi-modal feature vector of the target statement group and the multi-modal feature vector of the true target statement group, and the cross-entropy between the emotion recognition result and the correct prediction label.

6. An emotion recognition device, characterized in that, Include: The first processing module is used to establish a multi-modal emotion prediction model, and use the multi-modal features of the cascaded historical statement group and the multi-modal features of the target statement group to train the multi-modal emotion prediction model, obtaining a target multi-modal emotion prediction model, where the first three sentences of the target statement to be predicted constitute the historical statement group, the last three sentences of the target statement to be predicted constitute the target statement group, each sentence of the target statement to be predicted contains multi-modal dialogue information, and the target multi-modal emotion prediction model is used for emotion type recognition of the target statement to be predicted; The step of using the multi-modal features of the cascaded historical statement group and the multi-modal features of the target statement group to train the multi-modal emotion prediction model to obtain a target multi-modal emotion prediction model includes: Based on the multi-modal emotion prediction model, extract the multi-modal features of the historical statement group and the target statement group respectively, obtaining the multi-modal features of the historical statement group and the multi-modal features of the target statement group, and generate the hidden vector of the historical statement group and the hidden vector of the target statement group; Determine the multi-modal feature vector of the target statement, where, in the training phase, generate the multi-modal feature vector of the target statement according to the hidden vector of the target statement group and the multi-modal features of the historical statement group; or, in the testing phase, generate the multi-modal feature vector of the target statement according to the hidden vector of the historical statement group and the multi-modal features of the historical statement group; Identify the multi-modal feature vector of the target statement to obtain an emotion recognition result, where the emotion recognition result is used as the emotion classification result of the target statement. Optimizing the multimodal sentiment prediction model based on the backpropagation loss function to obtain the target multimodal sentiment prediction model; The multimodal sentiment prediction model adopts a conditional variational recurrent neural network structure based on the multi-head self-attention mechanism. Among them, the conditional variational recurrent neural network structure based on the multi-head self-attention mechanism includes 1 output layer, 1 input layer, and 14 hidden layers. The hidden layer includes 4 recurrent neural network layers, 3 multi-head self-attention layers, 3 activation layers, 3 fully connected layers, and 1 normalization layer; The second processing module is configured to process the target statement to be predicted based on the target multimodal sentiment prediction model to obtain the sentiment classification result of the target statement to be predicted.

7. An electronic device, characterized in that, Comprising: A processor; A memory for storing processor-executable instructions; Wherein, the processor is configured to execute the method according to any one of claims 1 to 5.

8. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored program, wherein when the program runs, it controls the device where the computer-readable storage medium is located to execute the method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Speech recognition and model training method and device

    CN114360511A

  • Cross-lingual classification using multilingual neural machine translation

    US20200342182A1