Dialogue emotion recognition method and device, electronic equipment and storage medium
By using strongly optimized bidirectional encoder and knowledge graph to build a model in dialogue emotion recognition, extracting context and speaker characteristics and fusing them for recognition, the problem of insufficient accuracy of dialogue emotion recognition in the existing technology is solved, and higher recognition accuracy is achieved.
Patent Information
- Application Number
- CN202510214703.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-06-10
AI Technical Summary
The prior art is difficult to achieve accurate recognition in dialogue emotions recognition, especially when relying on global features, it is difficult to capture the complex relationship between the context and the speaker's characteristics.
Using pre-trained powerful optimization of multiple converter layers in the bidirectional encoder, instead of the encoder and decoder in the variational autoencoder, a first feature extraction module is constructed to extract contextual features. At the same time, the automatic knowledge graph construction model is trained using the knowledge graph to obtain a second feature extraction module to extract the speaker's features. Then, the contextual features and the speaker features are fused and input into the emotion classification module for identification.
By extracting the context and speaker characteristics, and fusion of emotions is performed, the accuracy of dialogue emotions is significantly improved, and the emotions contained in each discourse to be identified can be more comprehensive and accurate.
Smart Images

Figure CN120124637A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of emotion recognition, and in particular, to a method, apparatus, electronic device, and storage medium for recognizing dialogue emotions. Background Art
[0002] Emotions contain rich attitudes and views, which affect decision-making and judgment. Dialogue emotion recognition has great application value in fields such as intelligent recommendation, social network analysis, psychological assessment systems, and box office prediction.
[0003] Related technologies extract global features of text information in a dialogue through models such as neural networks, and thus use the global features for emotion recognition. Since the text information in a dialogue is relatively short, it is difficult to achieve accurate emotion recognition only relying on the global features extracted from the text information. Summary of the Invention
[0004] Embodiments of this application provide a method, apparatus, electronic device, and storage medium for recognizing dialogue emotions to accurately recognize the emotions of utterances in a dialogue.
[0005] In a first aspect, embodiments of this application provide a method for recognizing dialogue emotions, including:
[0006] Obtain a dialogue to be recognized; wherein, the dialogue to be recognized includes a plurality of utterances;
[0007] Extract context features from the dialogue to be recognized through a first feature extraction module; wherein, the first feature extraction module is obtained by replacing the encoder and decoder in a variational autoencoder with multiple transformer layers in a pre-trained powerful optimization bidirectional encoder and training the variational autoencoder;
[0008] Extract speaker features from the dialogue to be recognized through a second feature extraction module; wherein, the second feature extraction module is obtained by training an automatic knowledge graph construction model using a preset knowledge graph;
[0009] Determine a fusion feature according to the context features and the speaker features;
[0010] Input the fusion feature into a preset emotion classification module to obtain the emotion of each utterance in the dialogue to be recognized output by the emotion classification module.
[0011] In a possible implementation manner, the context features include topic features and text features;
[0012] The step of extracting context features from the dialogue to be recognized through the first feature extraction module includes:
[0013] Extract a latent vector from the dialogue to be recognized through the encoder of the first feature extraction module;
[0014] Extract the element at the first preset position in the latent vector to obtain the topic feature;
[0015] Extract an output vector from the latent vector through the decoder of the first feature extraction module;
[0016] Extract the element at the second preset position in the output vector to obtain the text feature.
[0017] In a possible implementation manner, before extracting the latent vector from the dialogue to be recognized through the encoder of the first feature extraction module, it further includes:
[0018] Preprocess the dialogue to be recognized;
[0019] The preprocessing includes: reconstructing the dialogue to be recognized according to the digital sequence, label, discourse mask, and the number of true discourses of the dialogue to be recognized.
[0020] In a possible implementation manner, the determining the fusion feature according to the context feature and the speaker feature includes:
[0021] Concatenate the context feature and the speaker feature to obtain a concatenated feature;
[0022] Input the concatenated feature into a preset third feature extraction module to obtain the fusion feature output by the third feature extraction module; wherein, the third feature extraction module is trained according to an encoder composed of a multi-head self-attention layer and a feed-forward network layer.
[0023] In a possible implementation manner, the context feature includes a topic feature and a text feature;
[0024] The concatenating the context feature and the speaker feature to obtain a concatenated feature includes:
[0025] Concatenate the speaker feature between the text feature and the topic feature to obtain the concatenated feature.
[0026] In a possible implementation manner, the inputting the concatenated feature into a preset third feature extraction module to obtain the fusion feature output by the third feature extraction module includes:
[0027] Input the concatenated feature into the multi-head self-attention layer of the third feature extraction module to perform multi-head self-attention calculation to obtain a first feature;
[0028] Perform residual connection and layer normalization on the splicing feature and the first feature to obtain a second feature;
[0029] Input the second feature into the feed-forward network layer of the third feature extraction module to obtain a third feature;
[0030] Perform residual connection and layer normalization on the second feature and the third feature to obtain a fused feature.
[0031] In a possible implementation, the extracting, by the second feature extraction module, the speaker feature from the to-be-recognized dialogue includes:
[0032] Obtain a preset relationship type;
[0033] Input the to-be-recognized dialogue and the relationship type into the second feature extraction module to obtain the speaker feature output by the second feature extraction module.
[0034] In a second aspect, an embodiment of the present application provides a dialogue emotion recognition device, including:
[0035] An acquisition module, configured to acquire a to-be-recognized dialogue; wherein, the to-be-recognized dialogue includes a plurality of utterances;
[0036] A first extraction module, configured to extract a context feature from the to-be-recognized dialogue through a first feature extraction module; wherein, the first feature extraction module is obtained by replacing the encoder and decoder in a variational autoencoder with a plurality of transformer layers in a pre-trained powerful optimization bidirectional encoder and training the variational autoencoder;
[0037] A second extraction module, configured to extract a speaker feature from the to-be-recognized dialogue through a second feature extraction module; wherein, the second feature extraction module is a fusion module obtained by training an automatic knowledge graph construction model using a preset knowledge graph, and is configured to determine a fused feature according to the context feature and the speaker feature;
[0038] A recognition module, configured to input the fused feature into a preset emotion classification module to obtain the emotion of each utterance in the to-be-recognized dialogue output by the emotion classification module.
[0039] In a third aspect, an embodiment of the present application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, and when the processor executes the computer program, the steps of the method described in the first aspect or any possible implementation manner of the first aspect above are implemented.
[0040] Fourthly, an embodiment of the present application provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the steps of the method described in the first aspect above or any possible implementation manner of the first aspect are implemented.
[0041] Fifthly, an embodiment of the present application provides a computer program product, and when the computer program product runs on an electronic device, the electronic device is enabled to execute the steps of the method described in the first aspect above or any possible implementation manner of the first aspect.
[0042] The beneficial effects of the embodiments of the present application compared with the prior art are as follows:
[0043] In the embodiments of the present application, multiple transformer layers in a pre-trained powerful optimization bidirectional encoder are used to replace the encoder and decoder in the variational autoencoder to construct a first feature extraction module, which can optimize the model's ability to retrieve context information; by using the first feature extraction module to extract context features from the dialogue to be recognized, corresponding features can be accurately extracted from the text information of the dialogue, which helps to understand the context and improves the representation effect of the text information of the dialogue to be recognized, thereby improving the dialogue quality and the accuracy of emotion recognition. By training an automatic knowledge graph construction model using a knowledge graph to obtain a second feature extraction module, an external knowledge base can be introduced during the emotion recognition process; by using the second feature extraction module to obtain speaker features from the dialogue to be recognized, information strongly associated with the speaker image can be accurately focused on, realizing the extraction of speaker features, improving the dialogue information, and effectively improving the performance of emotion recognition. By fusing context features and speaker features and inputting the obtained fused features into an emotion classification module for emotion recognition, the influence of context and speaker characteristics on emotion recognition can be taken into account, so as to more comprehensively and accurately recognize the emotion contained in each utterance of the dialogue to be recognized. Description of the Drawings
[0044] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the following drawings are only some embodiments of the present application, and those of ordinary skill in the art can obtain other drawings without creative efforts based on these drawings.
[0045] Figure 1 It is a schematic diagram of a group of dialogues provided by an embodiment of the present application;
[0046] Figure 2 It is a flowchart of the implementation of the dialogue emotion recognition method provided by an embodiment of the present application;
[0047] Figure 3It is a schematic structural diagram of a dialogue emotion recognition model provided by an embodiment of the present application;
[0048] Figure 4 It is a schematic diagram of the common sense knowledge of an event provided by an embodiment of the present application;
[0049] Figure 5 It is a scatter plot of the emotion recognition results of a single relationship type provided by an embodiment of the present application;
[0050] Figure 6 It is a schematic structural diagram of a dialogue emotion recognition device provided by an embodiment of the present application;
[0051] Figure 7 It is a schematic diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0052] In the following description, specific details such as specific system architectures and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of the present application. However, those skilled in the art should clearly understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present application.
[0053] The main goal of emotion recognition in conversation (ERC) is to accurately recognize the emotion expressed by each utterance during a conversation. The inventor of the present application has found that emotion recognition in conversation has the following characteristics: First, utterances are context-dependent. Most daily conversations occur in a specific context, and the context of the conversation contributes significantly to emotion recognition. For example, the emotion of the sentence "You are really amazing" can be either positive or negative. In the Figure 1 shown conversation, influenced by the context of the conversation, "You are really amazing" is marked as "angry". Second, utterances are sensitive to the speaker. Everyone has a unique personality, which affects their feelings and emotional expressions. Speaker characteristics play an important role in emotion recognition. For example, in the Figure 1 conversation, when facing the same thing, Tony exclaims "Pretty good!" and Joey questions "Pretty good?", and the two sentences are marked with different emotions.
[0054] Based on this, the method provided by the present application takes into account the importance of both context and the speaker in emotion recognition. On the basis of obtaining context features, an external knowledge base is also used to focus on information strongly associated with the speaker's image and extract speaker characteristics. By fusing context features and speaker characteristics, the ability of the model to fuse diverse features is enhanced, thereby improving the accuracy of emotion recognition in conversation.
[0055] To make the objectives, technical solutions, and advantages of this application clearer, the following will be described through specific embodiments in conjunction with the accompanying drawings.
[0056] Figure 2 It is a flowchart for implementing the dialogue emotion recognition method provided in the embodiments of this application. Figure 3 It shows a schematic structural diagram of the dialogue emotion recognition model provided in the embodiments of this application. The dialogue emotion recognition model may include a first feature extraction module, a second feature extraction module, an emotion classification module, etc., which are described in detail as follows:
[0057] Step 201, obtain the dialogue to be recognized; where the dialogue to be recognized includes multiple utterances.
[0058] Here, the dialogue to be recognized involves M speakers and consists of N utterances. Each utterance corresponds to a unique speaker, and each utterance expresses a fine-grained emotion.
[0059] The dialogue to be recognized can be represented as U = {u 0 , u 1 , ……, u N}, where u 0 represents, u 1 , ……, u N represents the utterances in the dialogue to be recognized, and N represents the number of utterances in the dialogue to be recognized.
[0060] Step 202, through the first feature extraction module, extract context features from the dialogue to be recognized; where the first feature extraction module is obtained by replacing the encoder and decoder in the variational autoencoder with multiple transformer layers in a pre-trained powerful optimized bidirectional encoder and training the variational autoencoder.
[0061] The context features of the dialogue to be recognized refer to the features reflected by the context information related to the dialogue to be recognized, and these features can reflect information such as the meaning and emotional tendency of the dialogue in a specific context.
[0062] The Variational AutoEncoder (VAE) is a deep learning model mainly composed of an encoder and a decoder, which are used to perform encoding and decoding operations on data. VAE is an improvement of the Autoencoder (AE). AE can be regarded as a technology for compressing and decompressing complex vectors and is widely used in data dimensionality reduction and feature learning. For example, given an input text vector Input, the encoder Encoder compresses Input into a vector Code. Given Code and the decoder Decoder, a vector Output can be generated. During training, the difference between Input and Output is minimized. In practical applications, when AE is trained and the relevant parameters of the decoder are frozen, when generating text data by inputting a random vector Code, the result is usually incomprehensible. To solve the above problems, VAE adds noise to the vector Code to enhance the expressive power of the model.
[0063] Let x represent the input feature, z represent the latent vector, and y represent the output feature; p(z) represents the prior distribution; q(z|x) represents the probability distribution of the encoder, also known as the posterior distribution; p(y|z) represents the probability distribution of the decoder. The training objective of VAE is to maximize the Evidence Lower Bound (ELBO), and the formula is as follows.
[0064] ELBO = E[lnp(y|z)] - KL[q(z|x)||p(z)];
[0065] In the formula, E represents the expectation, and KL represents the Kullback-Leibler divergence.
[0066] In this embodiment, to further improve the effect of feature extraction, the VAE model is improved. Multiple transformer layers in the Robustly Optimized Bidirectional Encoder Representations from Transformers Pretraining Approach (RoBERTa) are used to replace the encoder and decoder of the VAE model. The transformer layers of RoBERTa can be used to improve the model's ability and effect in extracting dialogue context features.
[0067] It should be noted that the transformer layer replacing the encoder of the VAE model and the transformer layer replacing the decoder of the VAE model can be different.
[0068] In addition, by using a large number of dialogue data samples to train the improved variational autoencoder, it learns how to extract useful context features from the dialogue data. During the training process, by continuously adjusting the parameters of the model, the model can better extract the required features.
[0069] Step 203: Extract speaker features from the dialogue to be recognized through the second feature extraction module. Among them, the second feature extraction module is obtained by training an automatic knowledge graph construction model using a preset knowledge graph.
[0070] The speaker features cover various information related to the speaker image in the dialogue to be recognized, such as the speaker's language style (such as word - using habits, expression methods, tone, etc.), emotional tendency (positive, negative, neutral, etc.), knowledge background (common sense, professional knowledge, etc. reflected in the discourse), identity characteristics (characteristics that may be reflected from the language, such as age, occupation, cultural background, etc.).
[0071] A knowledge graph is a structured semantic knowledge base that displays the relationships between entities in a graphical way and contains a large amount of knowledge information. It can contain various entities (such as people, events, concepts, etc.) and the relationships between them (such as "person - occupation", "event - time" and other relationships). By introducing a knowledge graph, an external knowledge base can be used to help understand context information, which can effectively improve the performance of emotion recognition.
[0072] The automatic knowledge graph construction model (Commonsense Transformers for Automatic KnowledgeGraph Construction, COMET) is a model that can automatically extract entities and relationships from data such as text and construct a knowledge graph. It constructs a knowledge graph related to commonsense knowledge based on the Transformer architecture. According to the input text information, it can use the powerful feature extraction and semantic understanding capabilities of Transformer to automatically mine the entities and relationships therein and construct the corresponding knowledge graph.
[0073] Use the preset knowledge graph as training data to train COMET. During the training process, the model will learn how to extract information related to the preset knowledge graph from the input dialogue, thereby constructing knowledge graph - related features about the speaker. By continuously adjusting the model's parameters, the model can better extract the speaker's features from the dialogue, and finally obtain the second feature extraction module that can be used to extract speaker features.
[0074] Then, through the extraction of the dialogue to be recognized, the speaker features can be accurately extracted.
[0075] Optionally, in this embodiment, when extracting speaker features from the dialogue to be recognized through the second feature extraction module, the preset relationship type can be obtained first; then the dialogue to be recognized and the relationship type are input into the second feature extraction module to obtain the speaker features output by the second feature extraction module.
[0076] In this embodiment, the Atomic Commonsense Knowledge Graph (ATOMIC) can be used as an external knowledge base. ATOMIC is a large-scale commonsense knowledge graph that contains rich commonsense knowledge, such as commonsense information about people's intentions, reactions, and the results of events. This commonsense knowledge is very helpful for understanding the semantics in the dialogue and inferring the characteristics of the speaker.
[0077] ATOMIC is a graph atlas based on commonsense reasoning. The nodes are phrases describing events, and the edges are the relationships between two nodes. These relationships have 9 types, namely xIntent, xReact, oReact for describing mental states, xAttr for describing images, and xEffect, oEffect, xNeed, xWant, oWant for describing behaviors. See Figure 4 the schematic diagram of the commonsense knowledge of an event shown, in which some connection relationships are shown.
[0078] ConceptNet is also a knowledge graph that focuses on collecting and organizing the relationships between various concepts, covering a wider range of semantic knowledge, including hyponymy, meronymy, causal relationships, etc. For example, it clarifies the relationship that "apple" is a kind of "fruit", and "branch" is a part of "tree", etc.
[0079] COMET can be an adaptive framework based on ATOMIC and ConceptNet. It can automatically adjust and adapt its knowledge mining and graph construction strategies according to the different types and characteristics of knowledge in ATOMIC and ConceptNet, learn to generate rich commonsense knowledge, and realize the automatic construction of the knowledge base.
[0080] Assume that there is a training knowledge base in the {s, r, o} format, where s is the phrase topic of the tuple, r is the relationship type of the tuple, and o is the phrase object of the tuple. Input the phrase topic and relationship into COMET, and output the phrase object. Take Figure 4 as an example, input the phrase "X repels Y's attack" describing the event and the relationship "xAttr" into COMET, and expect to output another event "X is highly skilled".
[0081] Here, the second feature extraction module can consider the relationship type xAttr that is strongly associated with the speaker's image. Input the dialogue U and the relationship type xAttr into the second feature extraction module to obtain the representation of the speaker's features. Among them, the speaker's features can be denoted as Its dimension can be ds = 768. This dimension is only for illustration and not for limitation.
[0082] Step 204: Determine the fused features according to the context features and the speaker features.
[0083] In this embodiment, integrating or combining the context features and the speaker features can merge two different types of feature information together, which can more comprehensively describe the characteristics of the dialogue to be recognized.
[0084] When analyzing the emotion of a conversation, the fused features can not only consider the context in which the conversation takes place (such as whether it is about something happy or something frustrating), but also consider the personal characteristics of the speaker (such as whether the speaker is an optimistic person or a pessimistic person). In this way, the obtained fused features will be more helpful for more accurately understanding the emotional tendency, semantic connotation, etc. of the conversation.
[0085] Step 205: Input the fused features into a preset emotion classification module to obtain the emotion of each utterance in the dialogue to be recognized output by the emotion classification module.
[0086] In this embodiment, the emotion classification model can be constructed based on machine learning algorithms (such as support vector machines, naive Bayes, etc.) or deep learning models (such as recurrent neural networks, convolutional neural networks, etc.). The model is trained with dialogue data samples with emotion annotations, so as to learn the corresponding relationship between different features and emotions, and obtain the emotion classification module. Among them, emotions can include happiness, sadness, anger, surprise, fear, disgust, etc.
[0087] Optionally, the emotion classification module can include two layers of fully connected networks and a softmax function. After the fused features are sent into the two layers of fully connected networks, emotion classification is performed through the softmax function.
[0088] The softmax function can be used to obtain the probability value of each emotion, and the sum of the probability values of all emotions is 1.
[0089] During training, the cross entropy can be selected as the loss function to represent the gap between the model prediction probability and the actual probability. The smaller the cross entropy, the closer the prediction result is to the actual result. When the cross entropy is less than the preset threshold and the change rate of the cross entropy is less than the preset change threshold, the training of the module is completed.
[0090] The embodiments of the present application adopt multiple Transformer layers in a pre-trained powerful optimization bidirectional encoder to replace the encoder and decoder in the variational autoencoder, so as to construct a first feature extraction module, which can optimize the model's ability to retrieve context information; using the first feature extraction module to extract context features from the dialogue to be recognized can accurately extract corresponding features from the text information of the dialogue, help understand the context, improve the representation effect of the text information of the dialogue to be recognized, improve the dialogue quality, and thus improve the accuracy of emotion recognition. By using the knowledge graph to train the automatic knowledge graph construction model to obtain a second feature extraction module, an external knowledge base can be introduced in the emotion recognition process; by using the second feature extraction module to obtain speaker features from the dialogue to be recognized, information strongly associated with the speaker image can be accurately focused on, the extraction of speaker features can be realized, the dialogue information can be improved, and the performance of emotion recognition can be effectively improved. By fusing the context features and the speaker features and inputting the obtained fused features into the emotion classification module for emotion recognition, the influence of context and speaker characteristics on emotion recognition can be taken into account, so as to more comprehensively and accurately recognize the emotion contained in each utterance in the dialogue to be recognized.
[0091] In some embodiments, the context features include topic features and text features. Among them, the topic features reflect information related to the theme around which the dialogue revolves, such as features related to whether the dialogue is about tourism, food, technology, etc. The text features are features of the language, semantics, etc. of the dialogue text itself, such as word usage characteristics, sentence structure, etc.
[0092] As Figure 3 shown, the first feature extraction module includes an encoder and a decoder. Among them, the encoder and decoder can be respectively composed of multiple Transformer layers of RoBERTa, and the corresponding Transformer layers of the encoder and decoder can be different. For example, the encoder can be composed of 6 Transformer layers in RoBERTa. Here, 6 layers are for illustrative purposes only and are not limitations, and 5 layers or 7 layers, etc. can also be used.
[0093] In this embodiment, through the first feature extraction module, context features are extracted from the dialogue to be recognized. It can be to first extract a latent vector from the dialogue to be recognized through the encoder of the first feature extraction module; extract the element at the first preset position in the latent vector to obtain the topic feature; then extract an output vector from the latent vector through the decoder of the first feature extraction module; finally, extract the element at the second preset position in the output vector to obtain the text feature.
[0094] In this embodiment, topic features and text features, i.e., context features, are obtained after the processing of the encoder and the decoder respectively. Among them, the encoder will deeply analyze each word and sentence in the dialogue, and capture semantic, syntactic, logical and other information in the dialogue through the transformer layer of RoBERTa, and encode it into a latent vector. The latent vector is a compressed representation of the dialogue information.
[0095] Since the latent vector contains various types of information, by extracting the element at the first preset position from the latent vector, the features related to the dialogue topic can be obtained, i.e., the topic features. After extracting the element at the first preset position, further processing (such as normalization, feature fusion, etc.) can be performed.
[0096] Here, the element corresponding to the CLS token in the latent vector can be extracted to obtain the topic features. The topic features can be represented by which can be specifically expressed as Z = CLS[Encoder(U)].
[0097] The decoder will perform a decoding operation on the latent vector, utilize other transformer layers in RoBERTa, restore or further excavate the information related to the original dialogue text, and finally obtain an output vector. The output vector focuses more on reflecting the characteristics of the dialogue text itself, which is different from the focus of the latent vector.
[0098] Then, by extracting the element at the second preset position from the output vector, the key information related to the text can be extracted, and the information describing the dialogue text features can be obtained, such as the features in aspects of the language style and semantic coherence of the text, i.e., the text features.
[0099] Here, the element corresponding to the CLS token in the output vector can also be extracted to obtain the text features. The text features can be represented by which can be specifically expressed as C = CLS[Decoder(Encoder(U))].
[0100] Exemplarily, the dimensions of both the topic features and the text features can be 1024, i.e., dz = dc = 1024. This dimension is only for illustration and is not a limitation.
[0101] Optionally, before extracting the latent vector from the dialogue to be recognized through the encoder of the first feature extraction module, preprocessing is also performed on the dialogue to be recognized; the preprocessing includes: reconstructing the dialogue to be recognized according to the digital sequence, label, discourse mask and the number of true discourses of the dialogue to be recognized.
[0102] The preprocessing can convert the original dialogue to be recognized into a format and form suitable for subsequent model processing, improving the efficiency and accuracy of model processing.
[0103] In natural language processing, computers cannot directly process dialogue data in text form. Therefore, it is necessary to map each language unit such as words or characters in the dialogue into numbers to form a sequence of numbers. For example, using techniques such as word embeddings (such as Word2Vec, GloVe, etc.) or character embeddings, each word is converted into a digital vector with a fixed dimension, and then the digital vectors corresponding to all words in the dialogue are arranged in sequence to obtain a sequence of numbers.
[0104] Labels can represent the emotional tendency of the dialogue, such as positive, negative, neutral, etc., and can provide a reference for subsequent steps.
[0105] The utterance mask is a tool for marking certain parts of the dialogue. In a dialogue, some utterances may be important and require special attention; some utterances may be invalid or interfering and need to be specially processed during processing. The utterance mask can be represented by a binary sequence or other forms, where each element corresponds to an utterance or word in the dialogue. Through the mask, it can be indicated which parts need to be processed and which parts can be ignored.
[0106] Counting the actual number of utterances contained in the dialogue to be recognized can facilitate the subsequent analysis of the dialogue structure and accurately identify the emotions of each utterance.
[0107] In some embodiments, according to the context features and speaker features, the fused feature can be determined. First, the context features and speaker features can be concatenated to obtain a concatenated feature; then the concatenated feature is input into a preset third feature extraction module to obtain the fused feature output by the third feature extraction module; wherein, the third feature extraction module is trained according to an encoder composed of a multi-head self-attention layer and a feed-forward network layer.
[0108] In this embodiment, by concatenating the context features and speaker features, two different types of feature information are combined together, providing a basis for further extracting the fused feature subsequently. Among them, concatenation usually means connecting two or more feature vectors in sequence to form a longer feature vector. For example, if the context feature is a vector of length m and the speaker feature is a vector of length n, then the concatenated feature after concatenation is a vector of length m + n.
[0109] Considering the disadvantage that traditional recurrent neural networks cannot capture long-distance features, in this embodiment, an encoder composed of a multi-head self-attention layer and a feed-forward network layer is used to extract and integrate the concatenated feature to generate the final fused feature.
[0110] Among them, the multi-head self-attention layer enables the model to focus on the relationship between different positions in the input sequence when processing features. By calculating the attention weights between different positions, it highlights important information and better captures the association between features.
[0111] The feedforward network layer is usually composed of multiple fully connected layers, which further transform and nonlinearly map the features processed by the multi-head self-attention layer, and can learn more complex feature representations, and then extract high-quality fusion features from the spliced features.
[0112] Optionally, the context feature includes a topic feature and a text feature. In this embodiment, the context feature and the speaker feature are spliced to obtain a spliced feature. The speaker feature can be spliced between the text feature and the topic feature to obtain a spliced feature.
[0113] Here, both topic features and text features are extracted from the text information of the conversation to be identified. Topic features focus on the characteristic information of the core theme of the conversation to be identified, while text features focus on the language characteristics of the conversation text itself, such as the formality of the words, the complex structure of the sentences, the use of rhetoric, and other features.
[0114] In this embodiment, the speaker features are spliced between the text features and the topic features, and the speaker features can be inserted in the middle of the context features, so that the spliced features can integrate the information of the context features and the speaker features, and closely combine the three aspects of information, namely, topic, text and speaker, so as to subsequently extract high-quality fusion features.
[0115] For example, see Figure 4 As shown, text features, speaker features, and topic features can be spliced in order to accurately mine the correlation features between the speaker and the text, fully integrate the context features and the speaker features, and extract the fusion features that comprehensively and accurately reflect the whole picture of the conversation.
[0116] Optionally, the concatenated features are input into a preset third feature extraction module to obtain a fused feature output by the third feature extraction module. The concatenated features can be first input into the multi-head self-attention layer of the third feature extraction module to perform multi-head self-attention calculation to obtain the first feature; the concatenated features and the first features are residually connected and layer-normalized to obtain the second feature; the second feature is then input into the feedforward network layer of the third feature extraction module to obtain the third feature; finally, the second feature and the third feature are residually connected and layer-normalized to obtain the fused feature.
[0117] The meaning of multi - head is to calculate in parallel using multiple different attention heads, and each attention head focuses on different aspects of the features. In this embodiment, first, the multi - head self - attention layer performs attention calculation on each part of the concatenated features, which can capture the complex relationships and dependencies between different parts of the concatenated features, thereby highlighting important information.
[0118] The attention aggregation function can choose scaled dot - product attention, and its formula is as follows:
[0119]
[0120] Among them, head i represents the output of the i - th attention head, Q represents the query vector in the multi - head self - attention calculation, K represents the key vector in the multi - head self - attention calculation, V represents the value vector in the multi - head self - attention calculation, and d represents the vector length of Q, K, and V. Q, K, and V are obtained by linear transformation of the concatenated features. Its transformation formula can be:
[0121]
[0122] Among them, F represents the concatenated vector, W (q) represents the transformation matrix that converts the concatenated vector into the query vector, W (k) represents the transformation matrix that converts the concatenated vector into the key vector, W (v) represents the transformation matrix that converts the concatenated vector into the value vector, are trainable parameters, dc represents the dimension of the text feature, dz represents the dimension of the topic feature, ds represents the dimension of the speaker feature, and h represents the number of attention heads in the multi - head self - attention calculation.
[0123] After concatenating the results of multi - head attention and passing through another linear transformation, the output after the multi - head attention layer is obtained: M = [head 0 , head 1 , ……, head h-1 W (M) . In the formula, M represents the first feature output by the multi - head attention layer, W (M) represents the linear transformation matrix, which is used to map the concatenated result to the final output space, are trainable parameters.
[0124] In this embodiment, through the residual connection, the original input (concatenated features) can be added to the calculated feature (the first feature), alleviating the problem of gradient disappearance in the deep neural network and helping the model learn richer feature representations. Through layer normalization, the features are normalized, and the values of the features are adjusted to a suitable range, making the model easier to train and converge.
[0125] The feed-forward network layer is usually composed of multiple fully-connected layers, which can further transform and non-linearly map the second feature. Through the calculation of the feed-forward network layer, more complex feature representations can be learned, and deeper information in the second feature can be mined to obtain the third feature. For example, the feed-forward network layer can perform weighted calculations, activation function operations, etc. on the second feature to convert it into a feature representation more suitable for subsequent processing.
[0126] The fused feature obtained based on the above operations synthesizes the information after multiple steps such as the multi-head self-attention calculation and the feed-forward network layer processing of the spliced feature, and is more numerically stable and standardized for subsequent accurate emotion classification, improving the accuracy of dialogue emotion recognition.
[0127] In some embodiments, the MELD dataset and the EMORYNLP dataset are selected for experiments and compared with the following methods, and the comparison results are shown in Table 1.
[0128] (1) Contextual Reasoning Networks for Emotion Recognition in Conversations (DialogueCRN): Construct an inference module using recurrent neural networks and attention mechanisms to iterate the cognitive process and mimic the integration of clues in human thinking; (2) Dialogue Graph Convolutional Network (DialogueGCN): Use graph neural networks to construct contextual context, considering the influence of past and future utterances on the current utterance; (3) Directed Acyclic Graph Network for Conversational Emotion Recognition (DAG-ERC): Use a directed acyclic graph to construct contextual context, only considering the influence of past utterances on the current utterance, and realizing the combination of the advantages of graph neural networks and recurrent neural networks; (4) A Transformer-Based Context-and Speaker-Sensitive Model for Emotion Detection in Conversations (HiTrans): Use BERT to obtain local utterance features, design a Transformer to obtain global utterance features of the conversation. In addition, design an auxiliary task to judge whether two utterances belong to the same speaker; (5) Knowledge-Enriched Transformer for Emotion Detection in Textual Conversations (KET): Use the external knowledge base ConceptNet and the sentiment dictionary NRC_VAD to strengthen the understanding of the context; (6) Knowledge-Interactive Network with sentiment polarity intensity-aware multi-task learning (KI-Net): Use two knowledge sources, ConceptNet and SenticNet, and design the prediction of sentiment polarity intensity as an auxiliary task;(7) Spiking Kernel-based Attention Inference Graph (SKAIG): Obtain utterance-level features through RoBERTa, Maxpooling, and fully connected layers, simulate mental states using four types of commonsense knowledge, namely xWant, oWant, xIntent, and xEffect, and construct a mental-knowledge perception interaction graph. (8) CommonSense knowledge for emotion identification in conversations (COSMIC): Obtain context-independent utterance features by fine-tuning RoBERTa, and use GRU units and five types of knowledge, namely xIntent, xEffect, xReact, oEffect, and oReact, to obtain various state information and participate in the interaction between interlocutors. (9) Event-State Interactions Infused Heterogeneous Graph Neural Network for Conversational Emotion Recognition (ESIHGNN): Use a directed acyclic graph to combine xWant, oWant, xReact, and oReact to dynamically update the representations of events and emotions, improving the coherence and consistency of conversations.
[0129] Table 1 Comparison table of experimental results of each model
[0130]
[0131]
[0132] As can be seen from Table 1, the accuracy and F1 value of the conversational emotion recognition method provided by this application are better than those of each of the above models.
[0133] In the conversational emotion recognition method provided by this application, it is necessary to use ATOMIC as an external knowledge base. To compare the implementation effects of different relationship types, emotion recognition based on a single relationship type of ATOMIC was carried out on the MELD dataset. The experimental results are shown in Table 2. Figure 5 It is the corresponding scatter plot.
[0134] Table 2 Comparison table of emotion recognition results based on a single relationship type of ATOMIC
[0135]
[0136] As can be seen from Table 2, xAttr, xEffect, xReact, and oEffect are all competitive in the two evaluation metrics, and xAttr has the best comprehensive performance.
[0137] Based on the above experimental results, in order to test whether the effect of fusing multiple relation types is better than that of a single relation type, emotion recognition based on multiple relation types of ATOMIC was conducted on the MELD dataset. Table 3 shows the relevant experimental results.
[0138] Table 3 Comparison table of emotion recognition results based on multiple relation types of ATOMIC
[0139]
[0140] As can be seen from Table 3, the effect of fusing multiple relation types is worse than that of a single relation type. The reason may be that incorrect or redundant data brings noise to the model and affects the performance of the model.
[0141] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.
[0142] The following is an apparatus embodiment of the present application. For the details not described in detail herein, reference may be made to the corresponding method embodiments above.
[0143] Figure 6 The structural schematic diagram of the dialogue emotion recognition apparatus provided by the embodiments of the present application is shown. For the sake of convenience of description, only the parts related to the embodiments of the present application are shown and are described in detail as follows:
[0144] As Figure 6 shown, the dialogue emotion recognition apparatus 60 includes:
[0145] An acquisition module 61, configured to acquire a dialogue to be recognized; wherein, the dialogue to be recognized includes multiple utterances;
[0146] A first extraction module 62, configured to extract context features from the dialogue to be recognized through a first feature extraction module; wherein, the first feature extraction module is obtained by replacing the encoder and decoder in the variational autoencoder with multiple transformer layers in a pre-trained powerful optimization bidirectional encoder and training the variational autoencoder;
[0147] A second extraction module 63, configured to extract speaker features from the dialogue to be recognized through a second feature extraction module; wherein, the second feature extraction module is a fusion module 64 obtained by training an automatic knowledge graph construction model using a preset knowledge graph, and is configured to determine fusion features according to the context features and the speaker features;
[0148] An identification module 65, configured to input the fused features into a preset emotion classification module to obtain the emotion of each utterance in the to-be-identified dialogue output by the emotion classification module.
[0149] In a possible implementation manner, the context features include topic features and text features;
[0150] The first extraction module 62 is specifically configured to:
[0151] Extract a latent vector from the to-be-identified dialogue through the encoder of the first feature extraction module;
[0152] Extract the element at the first preset position in the latent vector to obtain the topic feature;
[0153] Extract an output vector from the latent vector through the decoder of the first feature extraction module;
[0154] Extract the element at the second preset position in the output vector to obtain the text feature.
[0155] In a possible implementation manner, the first extraction module 62 is further configured to:
[0156] Preprocess the to-be-identified dialogue;
[0157] The preprocessing includes: reconstructing the to-be-identified dialogue according to the digital sequence, label, utterance mask, and the number of true utterances of the to-be-identified dialogue.
[0158] In a possible implementation manner, the fusion module 64 is specifically configured to:
[0159] Concatenate the context features and the speaker features to obtain concatenated features;
[0160] Input the concatenated features into a preset third feature extraction module to obtain the fused features output by the third feature extraction module; wherein, the third feature extraction module is trained according to an encoder composed of a multi-head self-attention layer and a feed-forward network layer.
[0161] In a possible implementation manner, the context features include topic features and text features;
[0162] The fusion module 64 is specifically configured to:
[0163] Concatenate the speaker features between the text features and the topic features to obtain concatenated features.
[0164] In a possible implementation manner, the fusion module 64 is specifically configured to:
[0165] Input the concatenated features into the multi-head self-attention layer of the third feature extraction module to perform multi-head self-attention calculation to obtain a first feature;
[0166] Perform residual connection and layer normalization on the splicing feature and the first feature to obtain a second feature;
[0167] Input the second feature into the feed-forward network layer of the third feature extraction module to obtain a third feature;
[0168] Perform residual connection and layer normalization on the second feature and the third feature to obtain a fused feature.
[0169] In a possible implementation manner, the second extraction module 63 is specifically configured to:
[0170] Obtain a preset relationship type;
[0171] Input the dialogue to be recognized and the relationship type into the second feature extraction module to obtain the speaker feature output by the second feature extraction module.
[0172] Figure 7 It is a schematic diagram of an electronic device provided by an embodiment of the present application. As Figure 7 shown, the electronic device 70 of this embodiment includes: a processor 71, a memory 72, and a computer program 73 stored in the memory 72 and executable on the processor 71. When the processor 71 executes the computer program 73, the steps in the above-mentioned various embodiments of the dialogue emotion recognition method are implemented, such as Figure 2 the steps 201 to 205 shown. Alternatively, when the processor 71 executes the computer program 73, the functions of each module in the above-mentioned device embodiments are implemented, such as Figure 6 the functions of the modules 61 to 65 shown.
[0173] Exemplarily, the computer program 73 can be divided into one or more modules / units. One or more modules / units are stored in the memory 72 and executed by the processor 71 to complete the present application. One or more modules / units can be a series of computer program instruction segments capable of performing specific functions, and these instruction segments are used to describe the execution process of the computer program 73 in the electronic device 70. For example, the computer program 73 can be divided into Figure 6 the modules 61 to 65 shown.
[0174] The electronic device 70 may include, but is not limited to, a processor 71 and a memory 72. Those skilled in the art can understand that Figure 7 merely an example of the electronic device 70 does not constitute a limitation on the electronic device 70, and it may include more or fewer components than shown in the figure, or combine certain components, or different components. For example, the electronic device may further include input / output devices, network access devices, buses, etc.
[0175] The so-called processor 71 may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0176] The memory 72 may be an internal storage unit of the electronic device 70, such as the hard disk or memory of the electronic device 70. The memory 72 may also be an external storage device of the electronic device 70, such as a plug-in hard disk equipped on the electronic device 70, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. Further, the memory 72 may also include both the internal storage unit of the electronic device 70 and the external storage device. The memory 72 is used to store computer programs and other programs and data required by the electronic device. The memory 72 may also be used to temporarily store data that has been output or is to be output.
[0177] For the convenience and simplicity of description, only the above division of each functional module / unit is used as an example. In actual applications, the above functions may be allocated to different functional modules / units according to needs. The above modules / units may be implemented in the form of hardware, or may be implemented in the form of software, or may be implemented in the form of a combination of hardware and software.
[0178] The embodiment of the present invention also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the methods in the above method embodiments are implemented.
[0179] The embodiment of the present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, the methods in the above method embodiments are implemented.
[0180] Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc.
[0181] In the above embodiments, the descriptions of the respective embodiments have their own emphases. For parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments. Without special instructions and logical conflicts, the terms and / or descriptions between different embodiments are consistent and can be referenced to each other. The technical features in different embodiments can be combined to form new embodiments according to their inherent logical relationships.
[0182] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the protection scope of the present application.
Claims
1. A method for recognizing conversational emotions, characterized in that: include: Acquire a conversation to be recognized; wherein the conversation to be recognized includes multiple utterances; Extracting context features from the to-be-recognized conversation by a first feature extraction module; wherein the first feature extraction module is obtained by replacing the encoder and the decoder in the variational autoencoder with a plurality of transformer layers in a pre-trained powerfully optimized bidirectional encoder, and training the variational autoencoder; Extracting speaker features from the to-be-recognized conversation by a second feature extraction module; wherein the second feature extraction module is obtained by training an automatic knowledge graph construction model using a preset knowledge graph; Determining fusion features according to the context features and the speaker features; The fusion feature is input into a preset emotion classification module to obtain the emotion of each utterance in the to-be-identified conversation output by the emotion classification module.
2. The method for identifying conversational emotions according to claim 1, characterized in that: The context features include topic features and text features; The extracting context features from the to-be-recognized conversation by the first feature extraction module includes: extracting a latent vector from the to-be-recognized conversation by an encoder of the first feature extraction module; Extracting an element at a first preset position in the latent vector to obtain the topic feature; extracting an output vector from the latent vector by a decoder of the first feature extraction module; The element at the second preset position in the output vector is extracted to obtain the text feature.
3. The method for identifying conversational emotions according to claim 2, characterized in that: Before extracting the potential vector from the to-be-recognized conversation by the encoder of the first feature extraction module, the method further comprises: Preprocessing the dialogue to be identified; The preprocessing includes: reconstructing the dialogue to be recognized according to the digital sequence, label, speech mask and number of real speech of the dialogue to be recognized.
4. The method for identifying conversational emotions according to claim 1, characterized in that: The determining of the fusion feature according to the context feature and the speaker feature includes: concatenating the context feature and the speaker feature to obtain a concatenated feature; The spliced features are input into a preset third feature extraction module to obtain a fusion feature output by the third feature extraction module; wherein the third feature extraction module is obtained by training an encoder composed of a multi-head self-attention layer and a feedforward network layer.
5. The method for identifying conversational emotions according to claim 4, characterized in that: The context features include topic features and text features; The step of concatenating the context feature and the speaker feature to obtain a concatenated feature includes: The speaker feature is spliced between the text feature and the topic feature to obtain the spliced feature.
6. The method for identifying conversational emotions according to claim 4, characterized in that: The step of inputting the splicing feature into a preset third feature extraction module to obtain a fusion feature output by the third feature extraction module includes: Inputting the splicing feature into the multi-head self-attention layer of the third feature extraction module, performing multi-head self-attention calculation, and obtaining the first feature; Performing residual connection and layer normalization on the concatenated feature and the first feature to obtain a second feature; Inputting the second feature into the feedforward network layer of the third feature extraction module to obtain the third feature; Residual connection and layer normalization are performed on the second feature and the third feature to obtain a fused feature.
7. The method for recognizing conversational emotions according to any one of claims 1 to 6, characterized in that: The extracting speaker features from the to-be-recognized conversation by the second feature extraction module includes: Get the preset relationship type; The dialogue to be recognized and the relationship type are input into the second feature extraction module to obtain speaker features output by the second feature extraction module.
8. A conversation emotion recognition device, characterized in that: include: An acquisition module, used to acquire a conversation to be recognized; wherein the conversation to be recognized includes multiple utterances; A first extraction module, configured to extract context features from the to-be-recognized conversation through a first feature extraction module; wherein the first feature extraction module is obtained by replacing an encoder and a decoder in a variational autoencoder with a plurality of transformer layers in a pre-trained powerfully optimized bidirectional encoder, and training the variational autoencoder; The second extraction module is used to extract speaker features from the to-be-recognized conversation through a second feature extraction module; wherein the second feature extraction module is obtained by training the automatic knowledge graph construction model using a preset knowledge graph A fusion module, used for determining a fusion feature according to the context feature and the speaker feature; The recognition module is used to input the fusion feature into a preset emotion classification module to obtain the emotion of each utterance in the to-be-recognized dialogue output by the emotion classification module.
9. An electronic device comprising a memory and a processor, wherein the memory is used to store a computer program, and the processor is used to call and run the computer program stored in the memory, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.