Multi-modal dialogue emotion-reason pair extraction method based on speaker relationship

Through the combination of feature extractor and speaker relationship matrix, the problem of insufficient information utilization in the multimodal emotion-cause pair extraction model is solved, achieving higher accuracy and comprehensive emotion-cause pair recognition.

CN120336509AActive Publication Date: 2025-07-18JIANGXI UNIVERSITY OF FINANCE AND ECONOMICS

Patent Information

Application Number
CN202510773643.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-11
Publication Date
2025-07-18
Estimated Expiration
2045-06-11

AI Technical Summary

Technical Problem

The existing multimodal emotion-cause extraction model fails to make full use of multimodal information and speaker relationships, resulting in emotional-cause extraction incomplete and accurate enough.

Method used

A feature extractor is used to combine the speaker encoder, and a multi-modal graph convolution network is used to perform feature interaction and fusion, and an emotion-cause pairing is guided through the speaker relationship matrix, and an expert network and a gated network are constructed for classification prediction.

Benefits of technology

Improved the accuracy of emotion-cause extraction, with F1 values increased by 2.75% and 0.78%, better capture of multimodal information and dynamic social relations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120336509A_ABST
    Figure CN120336509A_ABST
Patent Text Reader

Abstract

The invention provides a multi-modal dialogue emotion-reason pair extraction method based on a speaker relationship. The method comprises the following steps: inputting a multi-modal dialogue into a feature extractor, and carrying out feature extraction in combination with a speaker encoder; carrying out feature interaction and fusion on the multi-modal features by utilizing a multi-modal graph convolutional network; classifying and predicting the final comprehensive feature representation by using a feedforward network; and performing candidate emotion-reason pairing by using the expression of the emotion aspect and the expression of the reason aspect, guiding extraction of the candidate emotion-reason pairing by using the speaker relation matrix, and performing classification prediction by using the expert network and the gating network to obtain a prediction result of the emotion-reason pair. According to the method, the speaker relation matrix is innovatively constructed, the dynamically changing social relation network is included in the consideration range of emotion-reason pair extraction, and the limitation that an existing method adopts simple dot product operation to evaluate utterance pair relevance is overcome.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of natural language processing, and particularly to a multimodal dialogue emotion - reason pair extraction method based on speaker relationships. Background Art

[0002] Emotion - cause pair extraction (ECPE), as a key natural language processing task, is dedicated to identifying emotional expressions in text and their triggering causes. This task is of great significance for enhancing the depth of sentiment analysis and generating empathetic responses, and thus has continuously received extensive attention from the research community. Current research methods mainly focus on text content analysis, especially on identifying emotions and causes at the sentence level. However, this method relying solely on text fails to fully utilize the rich multimodal context information in daily interactions. In actual situations, emotions are often presented through a combination of text and non - text cues. Relying solely on the text modality may lead to an incomplete understanding of emotions and their inducing factors, thereby restricting the overall performance of the ECPE task.

[0003] Based on this understanding, the multimodal emotion - cause pair extraction (MECPE) task is proposed, aiming to improve the extraction effect of emotion - cause pairs by integrating multimodal information. The MECPE task receives multimodal dialogue inputs containing text, vision, and audio, and the goal is to identify the emotions of all utterances and the cause utterances, and construct emotion - cause pairs (ECPs) at the utterance level. This task requires the system to deeply understand the multimodal input information and dialogue context, and at the same time accurately detect emotions, causes, and the relationship between the two, which poses a great challenge to accuracy.

[0004] For this task, existing research has proposed a two - stage method, which detects emotions, causes, and their corresponding relationships by integrating multimodal features and using an interaction matrix. Although this method shows certain effects in practice, we observe that it has two main limitations that may affect the overall performance of the model.

[0005] The primary problem is that most existing multimodal emotion - cause pair extraction models adopt a way of directly aggregating multimodal features to construct utterance - level representations, and then perform cause extraction. This processing method ignores the differential contributions of different modalities in triggering specific emotions.

[0006] Secondly, the current method uses a simple dot - product operation to evaluate the relevance between utterance pairs when extracting emotion - cause pairs. This mechanical processing method has obvious limitations. Especially in a dialogue scenario, the social relationships between speakers form a complex interaction network. The relationships between these speakers have a significant impact on emotion expression and cause attribution, but existing methods fail to effectively take this important dimension into account.

[0007] More importantly, these complex social relationship networks are often dynamically changing in conversations. A certain role may maintain different types of relationships with multiple other roles simultaneously, which further increases the complexity of emotion - reason pair extraction. Simply relying on simple dot - product operations to evaluate the associations between discourse pairs clearly cannot fully capture this complex social relationship dynamics and its impact on the formation of emotion - reason pairs. Summary of the Invention

[0008] In view of the above situation, the main objective of the present invention is to propose a multi - modal dialogue emotion - reason pair extraction method based on speaker relationships to solve the above - mentioned technical problems.

[0009] The present invention proposes a multi - modal dialogue emotion - reason pair extraction method based on speaker relationships, and the method includes the following steps: Step 1: Construct a feature extractor based on a feature extraction mechanism. The feature extractor includes a pre - trained language model, a bidirectional long - short - term memory network, an inflated 3D convolutional neural network model, an embedding layer, and a speaker encoder. Construct a graph convolutional network based on a feature interaction and fusion mechanism, construct a gating network based on a gating mechanism, and construct an expert network based on an expert mixture mechanism. The feature extractor, the graph convolutional network, the feed - forward network, the expert network, and the gating network form a prediction model; Step 2: Input the multi - modal dialogue into the feature extractor, and perform feature extraction on the multi - modal dialogue in combination with the speaker encoder to obtain multi - modal features; Step 3: Perform feature interaction and fusion on the multi - modal features using a multi - modal graph convolutional network to obtain the final comprehensive feature representation; Step 4: Use the feed - forward network to classify and predict the final comprehensive feature representation to obtain the representation in terms of emotion, the prediction result in terms of emotion, the representation in terms of reason, and the prediction result in terms of reason respectively; Step 5: Perform candidate emotion - reason pairing using the representation in terms of emotion and the representation in terms of reason, use the speaker relationship matrix to guide the extraction of candidate emotion - reason pairing, and use the expert network and the gating network for classification and prediction to obtain the prediction result of the emotion - reason pair; Construct an emotion prediction loss based on the prediction result in terms of emotion; Construct a reason recognition loss based on the prediction result in terms of reason; Construct an emotion - reason pairing classification loss based on the final emotion - reason pair; Optimize the prediction model using the emotion prediction loss, the reason recognition loss, and the emotion - reason pairing classification loss to obtain an optimized prediction model; Input the multi - modal dialogue into the optimized prediction model to obtain the final prediction result.

[0010] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. By using a dedicated feature extractor to extract text, audio, and visual modality features respectively, the present invention avoids the limitations of existing methods that directly aggregate multi-modal features, can capture multi-modal clues of emotional expression more comprehensively, and fully utilizes the differential contributions of different modalities in triggering specific emotions. Experimental results show that compared with the BASE method that only relies on a single modality, MM-SRM has improved the F1 value by 2.75%, verifying the importance of effective fusion of multi-modal information for task performance; 2. By innovatively constructing a speaker relationship matrix and incorporating the dynamically changing social relationship network into the consideration of emotion-cause pair extraction, the present invention overcomes the limitation of existing methods that use simple dot product operations to evaluate the relevance of discourse pairs. Through the learnable speaker relationship matrix and the relationship-based expert system, this method can capture the complex emotion-cause association patterns in conversations more accurately. Compared with the state-of-the-art HILO method, MM-SRM has improved the F1 value by 0.78%, demonstrating the positive effect of considering social relationship information between speakers on emotion-cause pair recognition.

[0011] The additional aspects and advantages of the present invention will be partially given in the following description, partially become apparent from the following description, or be understood through the embodiments of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] Figure 1 is a flowchart of the steps of the multi-modal dialogue emotion-cause pair extraction method based on speaker relationship proposed by the present invention.

[0013] Figure 2 is an architecture diagram of the multi-modal emotion cause recognition method of the multi-modal dialogue emotion-cause pair extraction method based on speaker relationship proposed by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0014] The embodiments of the present invention will be described in detail below. The examples of the embodiments are shown in the drawings, where the same or similar reference numerals are the same or similar elements or elements with the same or similar functions from beginning to end. The embodiments described below with reference to the drawings are exemplary and are only used to explain the present invention and should not be construed as a limitation of the present invention.

[0015] Referring to the following description and drawings, these and other aspects of the embodiments of the present invention will be clear. In these descriptions and drawings, some specific embodiments of the embodiments of the present invention are specifically disclosed as some ways to implement the principles of the embodiments of the present invention, but it should be understood that the scope of the embodiments of the present invention is not limited thereto.

[0016] Please refer to Figure 1, an embodiment of the present invention proposes a method for extracting multi-modal dialogue emotion-reason pairs based on speaker relationships, and the method includes the following steps: Step 1, construct a feature extractor based on the feature extraction mechanism. The feature extractor includes a pre-trained language model, a bidirectional long short-term memory network, an inflated 3D convolutional neural network model, an embedding layer, and a speaker encoder. Construct a graph convolutional network based on the feature interaction and fusion mechanism, construct a gating network based on the gating mechanism, and construct an expert network based on the mixture of experts mechanism. The feature extractor, the graph convolutional network, the feed-forward network, the expert network, and the gating network constitute a prediction model.

[0017] Step 2, input the multi-modal dialogue into the feature extractor, and combine with the speaker encoder to perform feature extraction on the multi-modal dialogue to obtain multi-modal features.

[0018] Please refer to Figure 2 , in Step 2, input the multi-modal dialogue into the feature extractor, and combine with the speaker encoder to perform feature extraction on the multi-modal dialogue to obtain multi-modal features, which specifically includes the following steps: Input the multi-modal dialogue, perform sequence splicing on the multi-modal dialogue, and use a delimiter to obtain the input sequence. Input the input sequence into the pre-trained language model, and combine with the speaker encoder to perform feature extraction to obtain the extracted text token vector representation; Perform feature extraction on the multi-modal dialogue using an open-source toolkit to obtain the extracted original audio features. Input the original audio features into the bidirectional long short-term memory network, and combine with the speaker encoder to perform feature extraction to obtain the extracted audio features; Perform feature extraction on the multi-modal dialogue using an inflated 3D convolutional neural network model to obtain the extracted original visual features. Perform feature extraction on the extracted original visual features through a bidirectional long short-term memory network combined with the speaker encoder to obtain the extracted visual representation; Perform vector representation on the multi-modal dialogue to obtain the represented vector. Pass the represented vector through the embedding layer and combine with the speaker encoder to perform a mapping operation to obtain the speaker embedding vector.

[0019] Input the multi-modal utterance, and the following relationship exists in the corresponding process: ; Wherein, represents the multi-modal dialogue, represents the first speaker, represents the first utterance, represents the th speaker, represents the th utterance; In the step of concatenating multi-modal dialogues into a sequence, separating them with delimiters to obtain the input sequence, inputting the input sequence into a pre-trained language model, and extracting features in combination with a speaker encoder to obtain the extracted text token vector representation, the relational expressions in the corresponding process are as follows: ; Among them, represents the entire sequence input into the pre-trained language model, represents the start token in the pre-trained language model, represents the end token in the pre-trained language model, represents the first token in the first text utterance, represents the first token in the second text utterance, represents the first start token of the vector representation, represents the second start token of the vector representation, represents the end token of the vector representation, represents after being processed by the pre-trained language model, represents the th start token of the vector representation; It should be noted that is used to mark the start of a sentence, contains all the concatenated utterance contents and the special delimiters, is used to mark the end of a sentence.

[0020] In the step of extracting features from multi-modal dialogues using an open-source toolkit to obtain the extracted original audio features, inputting the original audio features into a bidirectional long short-term memory network, and extracting features in combination with a speaker encoder to obtain the extracted audio features, the relational expressions in the corresponding process are as follows: ; Among them, represents the original audio features of the th utterance, represents after being extracted by the open-source toolkit features, represents the th utterance of the audio features, represents after being processed by the bidirectional long short-term memory network; It should be noted that is an open-source toolkit for audio feature extraction.

[0021] In the step of extracting features from multi-modal dialogue using an inflated three-dimensional convolutional neural network model to obtain the extracted original visual features, and then extracting features from the extracted original visual features through a bidirectional long short-term memory network combined with a speaker encoder to obtain the extracted visual representation, the relational expressions in the corresponding process are as follows: ; Among them, represents the original visual features of the th utterance, represents after being processed by the inflated three-dimensional convolutional neural network model, represents the th visual feature, represents extracting features from the original visual features; It should be noted that is a deep neural network model for video analysis and visual feature extraction. This model can effectively capture spatio-temporal features in video sequences through a three-dimensional convolutional neural network architecture, and is used to transform the original features into a high-level feature representation with context semantic information.

[0022] In the step of representing the multi-modal dialogue as a row vector to obtain the represented vector, and then performing a mapping operation on the represented vector through an embedding layer and combining it with a speaker encoder to obtain the speaker embedding vector, the relational expressions in the corresponding process are as follows: ; Among them, represents the vector corresponding to the speaker of the th dialogue , represents the real number space of length , represents the sentence speaker embedding vector, represents the embedding matrix.

[0023] It should be noted that has only one element equal to 1, corresponding to the position of the current speaker, and the remaining elements are all 0.

[0024] Step 3: Use a multi-modal graph convolutional network to perform feature interaction and fusion on the multi-modal features to obtain the final comprehensive feature representation.

[0025] In Step 3, using a multi-modal graph convolutional network to perform feature interaction and fusion on the multi-modal features to obtain the final comprehensive feature representation specifically includes the following steps: Define modal nodes for the extracted text token vector representation, the extracted audio features, and the extracted visual representation, respectively obtaining a text modal node, an audio modal node, and a visual modal node; Fully connect all nodes of the same modality in the multimodal dialogue, connect different modality nodes of different utterances to each other, and perform graph construction to obtain the constructed undirected graph; Use a graph convolutional network to iteratively update the node features in the constructed undirected graph to obtain the final node representation; Successively perform feature extraction operations and feature fusion operations on the final node representation to obtain the fusion features of the utterance; Concatenate the fusion features of the utterance, the speaker embedding vector, and the utterance in the multimodal dialogue to obtain the final comprehensive feature representation.

[0026] Fully connect all nodes of the same modality in the multimodal dialogue, connect different modality nodes of different utterances to each other, and perform graph construction to obtain the constructed undirected graph. The relational expressions for the corresponding process are as follows: ; Among them, represents the constructed undirected graph, represents the utterance nodes in the text modal node, the audio modal node, and the visual modal node, represents the edge set containing context and modality dependency relationships; In the step of using a graph convolutional network to iteratively update the node features in the constructed undirected graph to obtain the final node representation, the relational expressions for the corresponding process are as follows: ; Among them, represents the node feature matrix after passing through the -th layer of the graph convolutional network, represents the activation function, and respectively represent two hyperparameters, represents the identity matrix, represents the normalized graph Laplacian matrix of the undirected graph, represents the node feature matrix after passing through the -th layer of the graph convolutional network, represents the original feature representation of the node, represents the learnable weight matrix, represents the layer index of the graph convolutional network, represents the feature representation of the text modal node of the first utterance after passing through the -th layer of the graph convolutional network, represents the feature representation of the audio modal node of the first utterance after passing through the The feature representation after the layer graph convolutional network Indicates the feature representation of the visual modality node of the first utterance after passing through the layer graph convolutional network Indicates the th utterance's text modality node after passing through the layer graph convolutional network Indicates the th utterance's audio modality node after passing through the layer graph convolutional network Indicates the th utterance's visual modality node after passing through the layer graph convolutional network; In the step of successively performing feature extraction operations and feature fusion operations on the final node representation to obtain the fused feature of the utterance, the relational expressions existing in the corresponding process are as follows: ; Among them, Indicates the final feature of the text modality of the th utterance Indicates the final feature of the audio modality of the th utterance Indicates the final feature of the visual modality of the th utterance Indicates the feature concatenation operation Indicates the th utterance's fused feature; In the step of concatenating the fused feature of the utterance, the speaker embedding vector, and the utterance in the multi-modal dialogue to obtain the final comprehensive feature representation, the relational expressions existing in the corresponding process are as follows: ; Among them, Indicates the final comprehensive feature representation Indicates the concatenation operation

[0027] Step 4: Use a feedforward network to classify and predict the final comprehensive feature representation to obtain the representation in terms of emotion, the prediction result in terms of emotion, the representation in terms of reason, and the prediction result in terms of reason respectively.

[0028] In Step 4, use a feedforward network to classify and predict the final comprehensive feature representation to obtain the representation in terms of emotion, the prediction result in terms of emotion, the representation in terms of reason, and the prediction result in terms of reason respectively, which specifically includes the following steps: Perform mechanism processing on the final comprehensive feature representation using a feedforward network to obtain the representation in terms of emotion; Predict the representation of the emotional aspect to obtain the prediction result of the emotional aspect; Perform mechanism processing on the final comprehensive feature representation using a feed-forward network to obtain the representation of the reason aspect; Predict the representation of the reason aspect to obtain the prediction result of the reason aspect.

[0029] Perform mechanism processing on the final comprehensive feature representation using a feed-forward network to obtain the representation of the emotional aspect. The relational formula existing in the corresponding process is as follows: ; Among them, represents the representation of the emotional aspect, represents a two-layer feed-forward network; In the step of predicting the representation of the emotional aspect to obtain the prediction result of the emotional aspect, the relational formula existing in the corresponding process is as follows: ; Among them, represents the emotional prediction probability distribution, represents being subjected to normalization processing, represents the predicted emotional label, represents returning the maximum value in the independent variable, represents the class index; It should be noted that represents the class index, that is, traversing all possible emotional categories to find the target with the highest probability.

[0030] In the step of performing mechanism processing on the final comprehensive feature representation using a feed-forward network to obtain the representation of the reason aspect, the relational formula existing in the corresponding process is as follows: ; Among them, represents the representation of the reason aspect, represents another two-layer feed-forward network; In the step of predicting the representation of the reason aspect to obtain the prediction result of the reason aspect, the relational formula existing in the corresponding process is as follows: ; Among them, represents the predicted reason label probability.

[0031] Step 5: Use the representation of the emotional aspect and the representation of the reason aspect to perform candidate emotion-reason pairing, use the speaker relationship matrix to guide the extraction of candidate emotion-reason pairing, and use the expert network and the gating network to perform classification prediction to obtain the prediction result of the emotion-reason pair; Construct an emotional prediction loss based on the prediction result of the emotional aspect; Construct a cause recognition loss based on the prediction results in terms of causes; Construct a sentiment-cause pairing classification loss based on the final sentiment-cause pairs; Optimize the prediction model using the sentiment prediction loss, the cause recognition loss, and the sentiment-cause pairing classification loss to obtain an optimized prediction model; Input the multi-modal dialogue into the optimized prediction model to obtain the final prediction results.

[0032] In step 5, use the representation in terms of sentiment and the representation in terms of causes to perform candidate sentiment-cause pairing, use the speaker relationship matrix to guide the extraction of candidate sentiment-cause pairs, and use an expert network and a gating network for classification prediction to obtain the prediction results of sentiment-cause pairs, which specifically include the following steps: Construct candidate pairs for the cause discourse candidate representations in the representation in terms of sentiment and the representation in terms of causes to obtain the constructed candidate sentiment-cause pairs; Concatenate the representation in terms of causes and the representation in terms of sentiment to obtain the features of the candidate pairs; Construct a role relationship matrix based on the speaker relationship to obtain the constructed role relationship matrix, and use the constructed role relationship matrix to perform a projection operation on the constructed candidate sentiment-cause pairs to obtain the speaker relationship features; Use a gating mechanism to fuse the features of the candidate pairs and the speaker relationship features to obtain the final fused feature representation; Input the final fused feature representation into the expert network to predict the constructed candidate sentiment-cause pairs to obtain the predicted values output by the expert network; Input the final fused feature representation into the gating network to predict the constructed candidate sentiment-cause pairs to obtain the predicted values output by the gating network; Create a one-hot label based on the sentiment-cause pair relationship and combine the probability distribution output by the gating network Perform re-prediction to obtain the final comprehensive routing probability distribution; Combine the final comprehensive routing probability distribution with the predicted values output by the expert network for classification prediction to obtain the classification prediction results; Perform a threshold judgment on the classification prediction results to obtain the prediction results of sentiment-cause pairs.

[0033] Construct candidate pairs for the representation in terms of sentiment and the representation in terms of causes to obtain the constructed candidate sentiment-cause pairs, and the relational expressions existing in the corresponding process are as follows: ; Among them, represents the sentence and the sentence Candidate pairs formed Indicates the speaker of the sentence of Indicates the speaker of the sentence of Indicates the candidate representation of the causal discourse In the step of splicing the representation of the causal aspect and the representation of the emotional aspect to obtain the features of the candidate pair, the relational expressions existing in the corresponding process are as follows: ; Among them, Indicates the features of the candidate pair; In the step of constructing a role relationship matrix based on the speaker relationship to obtain the constructed role relationship matrix, and using the constructed role relationship matrix to perform a projection operation on the constructed candidate emotion-cause pair to obtain the speaker relationship features, the relational expressions existing in the corresponding process are as follows: ; Among them, Indicates the constructed role relationship matrix, Indicates the sentence Speaker embedding vector, and respectively represent two different learnable weight matrices, Indicates the speaker relationship features; In the step of using a gating mechanism to fuse the features of the candidate pair and the speaker relationship features to obtain the final fused feature representation, the relational expressions existing in the corresponding process are as follows: ; Among them, Indicates the calculated gating value, Represents the first learnable parameter, Represents the second learnable parameter, Indicates the final fused feature representation; In the step of inputting the final fused feature representation into the expert network to predict the constructed candidate emotion-cause pair to obtain the predicted value output by the expert network, the relational expressions existing in the corresponding process are as follows: ; Among them, Indicates the th expert network's predicted value for the candidate pair formed by the sentence and the sentence Output, Represents the third learnable parameter, Represents the fourth learnable parameter, Represents Activation function denotes the 5th learnable parameter denotes the 6th learnable parameter; In the step of inputting the final fused feature representation into the gating network to predict the constructed candidate sentiment - reason pairings to obtain the predicted values output by the gating network, the relational expressions existing in the corresponding process are as follows: ; wherein denotes the distribution over 5 experts denotes the 7th learnable parameter denotes the 8th learnable parameter; It should be noted that the five experts in respectively handle five different corresponding situations. In the actual inference situation, the model cannot determine the label of the candidate pairings, and the gating network is used to adjust the probability of the experts corresponding to the predicted candidate pairings.

[0034] In the step of creating one - hot labels based on the sentiment - reason pair relationship and performing re - prediction by combining the probability distribution output by the gating network to obtain the final comprehensive routing probability distribution, the relational expressions existing in the corresponding process are as follows: ; wherein denotes the final comprehensive routing probability distribution denotes the balance parameter denotes the guiding probability distribution based on the pairing relationship category; In the step of classifying and predicting by combining the final comprehensive routing probability distribution with the predicted values output by the expert network to obtain the classification prediction result, the relational expressions existing in the corresponding process are as follows: ; wherein denotes the sentiment - reason pairing the final classification prediction probability of denotes the routing probability of the th expert denotes the prediction probability of the th expert for the sentiment - reason pairing.

[0035] Construct a sentiment prediction loss based on the prediction result on the sentiment aspect. The relational expressions existing in the corresponding process are as follows: ; wherein denotes the sentiment prediction loss denotes the logarithmic function denotes the The true sentiment label of a piece of discourse; In the step of constructing the cause recognition loss based on the prediction results in terms of causes, the relational expressions existing in the corresponding process are as follows: ; Among them, represents the cause recognition loss, represents the th true cause label of a piece of discourse; In the step of constructing the sentiment-cause pairing classification loss based on the final sentiment-cause pairs, the relational expressions existing in the corresponding process are as follows: ; Among them, represents the sentiment-cause pairing classification loss, represents the true label of the sentiment-cause pairing .

[0036] Furthermore, there is also a total loss function, and the relational expressions existing in the corresponding process are as follows: ; Among them, represents the total loss function.

[0037] It should be understood that each part of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.

[0038] In the description of this specification, the descriptions referring to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.

[0039] The above-described embodiments merely represent several implementation manners of the present invention. The description thereof is relatively specific and detailed, but it should not be construed as a limitation to the scope of the patent of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention shall be subject to the appended claims.

Claims

1. A multimodal dialogue emotion - reason pair extraction method based on speaker relationships, characterized in that, The method includes the following steps: Step 1: Construct a feature extractor based on a feature extraction mechanism. The feature extractor includes a pre-trained language model, a bidirectional long short-term memory network, an inflated 3D convolutional neural network model, an embedding layer, and a speaker encoder. Construct a graph convolutional network based on a feature interaction and fusion mechanism, construct a gating network based on a gating mechanism, and construct an expert network based on a mixture of experts mechanism. The feature extractor, the graph convolutional network, the feed-forward network, the expert network, and the gating network constitute a prediction model; Step 2: Input a multi-modal dialogue into the feature extractor, and perform feature extraction on the multi-modal dialogue in combination with the speaker encoder to obtain multi-modal features; Step 3: Use the multi-modal graph convolutional network to perform feature interaction and fusion on the multi-modal features to obtain a final comprehensive feature representation; Step 4: Use the feed-forward network to classify and predict the final comprehensive feature representation to obtain representations in terms of emotion, prediction results in terms of emotion, representations in terms of reason, and prediction results in terms of reason respectively; Step 5: Use the representations in terms of emotion and the representations in terms of reason to perform candidate emotion-reason pairing, use the speaker relationship matrix to guide the extraction of candidate emotion-reason pairing, and use the expert network and the gating network to perform classification prediction to obtain the prediction results of the emotion-reason pair; Construct an emotion prediction loss based on the prediction results in terms of emotion; Construct a reason recognition loss based on the prediction results in terms of reason; Construct an emotion-reason pairing classification loss based on the final emotion-reason pair; Use the emotion prediction loss, the reason recognition loss, and the emotion-reason pairing classification loss to optimize the prediction model to obtain an optimized prediction model; Input the multi-modal dialogue into the optimized prediction model to obtain the final prediction result.

2. The method for extracting multi-modal dialogue emotion-cause pairs based on speaker relationships according to claim 1, characterized in that, In the said Step 2, input the multi-modal dialogue into the feature extractor, and perform feature extraction on the multi-modal dialogue in combination with the speaker encoder to obtain multi-modal features, which specifically includes the following steps: Input the multi-modal dialogue, perform sequence concatenation on the multi-modal dialogue, and use delimiters to separate it to obtain the input sequence. Input the input sequence into the pre-trained language model, and perform feature extraction in combination with the speaker encoder to obtain the extracted text token vector representation; Perform feature extraction on the multi-modal dialogue using an open-source toolkit to obtain the extracted original audio features. Input the original audio features into the bidirectional long short-term memory network, and perform feature extraction in combination with the speaker encoder to obtain the extracted audio features; Perform feature extraction on the multi-modal dialogue using the inflated 3D convolutional neural network model to obtain the extracted original visual features. Perform feature extraction on the extracted original visual features through the bidirectional long short-term memory network in combination with the speaker encoder to obtain the extracted visual representation; Perform vector representation on the multi-modal dialogue to obtain the represented vector. Pass the represented vector through the embedding layer and perform a mapping operation in combination with the speaker encoder to obtain the speaker embedding vector.

3. The method for extracting multi-modal dialogue emotion-reason pairs based on speaker relationships according to claim 2, characterized in that, Input the multi-modal utterance, and the relational expressions existing in the corresponding process are as follows: ; Among them, represents multimodal dialogue, represents the first speaker, represents the first utterance, represents the th speaker, represents the th utterance; In the step of splicing the multi-modal dialogue into a sequence, separating the input sequence using delimiters, inputting the input sequence into a pre-trained language model, and extracting features in combination with a speaker encoder to obtain the extracted text token vector representation, the relational expressions existing in the corresponding process are as follows: ; Among them, represents the entire sequence input into the pre-trained language model, represents the start token in the pre-trained language model, represents the end token in the pre-trained language model, represents the first token in the first text utterance, represents the first token in the second text utterance, represents the first start token 's vector representation, represents the second start token 's vector representation, represents the end token 's vector representation, represents after being processed by the pre-trained language model, represents the th start token 's vector representation; In the step of extracting features from the multi-modal dialogue using an open-source toolkit to obtain the extracted original audio features, inputting the original audio features into a bidirectional long short-term memory network, and extracting features in combination with a speaker encoder to obtain the extracted audio features, the relational expressions existing in the corresponding process are as follows: ; Among them, represents the original audio features of the th utterance, represents the audio features of the utterance after feature extraction by the open-source toolkit, represents the th utterance, represents the processing by a bidirectional long short-term memory network; In the step of extracting features from the multi-modal dialogue using a dilated 3D convolutional neural network model to obtain the extracted original visual features, and extracting features from the extracted original visual features through a bidirectional long short-term memory network in combination with a speaker encoder to obtain the extracted visual representation, the relational expressions existing in the corresponding process are as follows: ; Among them, represents the original visual features of the th utterance, represents being processed by the dilated 3D convolutional neural network model, represents the th visual feature, represents feature extraction from the original visual features; In the step of performing vector representation on the multi-modal dialogue to obtain the represented vector, passing the represented vector through an embedding layer, and performing a mapping operation in combination with a speaker encoder to obtain the speaker embedding vector, the relational expressions existing in the corresponding process are as follows: ; Among them, represents the vector corresponding to the speaker of the -th conversation, represents the real number space of length , represents the sentence speaker embedding vector, represents the embedding matrix.

4. The method for extracting multi-modal dialogue emotion-cause pairs based on speaker relationships according to claim 3, characterized in that In step 3 above, feature interaction and fusion are performed on the multi-modal features using a multi-modal graph convolutional network to obtain the final comprehensive feature representation, which specifically includes the following steps: Define modal nodes for the extracted text token vector representation, the extracted audio features, and the extracted visual representation, respectively obtaining a text modal node, an audio modal node, and a visual modal node; Fully connect all nodes of the same modality in the multi-modal dialogue, connect different modality nodes of different utterances to each other, and perform graph construction to obtain the constructed undirected graph; Use a graph convolutional network to iteratively update the node features in the constructed undirected graph to obtain the final node representation; Perform feature extraction operations and feature fusion operations on the final node representation in sequence to obtain the fused features of the utterance; Concatenate the fused features of the utterance, the speaker embedding vector, and the utterance in the multi-modal dialogue to obtain the final comprehensive feature representation.

5. The method for extracting multi-modal dialogue emotion-reason pairs based on speaker relationships according to claim 4, wherein Fully connect all nodes of the same modality in the multi-modal dialogue, connect different modality nodes of different utterances to each other, and perform graph construction to obtain the constructed undirected graph, and the relational expressions existing in the corresponding process are as follows: ; Among them, represents the constructed undirected graph, represents the discourse nodes among the text modality nodes, audio modality nodes, and visual modality nodes, represents the edge set containing context and modality dependencies; In the step of using a graph convolutional network to iteratively update the node features in the constructed undirected graph to obtain the final node representation, the relational expressions existing in the corresponding process are as follows: ; Among them, represents the node feature matrix after passing through the -th layer of the graph convolutional network, represents the activation function, and represent two hyperparameters respectively, represents the identity matrix, represents the normalized graph Laplacian matrix of the undirected graph, represents the node feature matrix after passing through the -th layer of the graph convolutional network, represents the original feature representation of the node, represents the learnable weight matrix, represents the layer index of the graph convolutional network, represents the feature representation of the text modality node of the first utterance after passing through the -th layer of the graph convolutional network, represents the feature representation of the audio modality node of the first utterance after passing through the -th layer of the graph convolutional network, represents the feature representation of the visual modality node of the first utterance after passing through the -th layer of the graph convolutional network, represents the -th feature representation of the text modality node of the -th utterance after passing through the -th layer of the graph convolutional network, -th feature representation of the audio modality node of the -th utterance after passing through the -th layer of the graph convolutional network, -th feature representation of the visual modality node of the -th utterance after passing through the -th layer of the graph convolutional network; In the step of performing feature extraction operations and feature fusion operations on the final node representation in sequence to obtain the fused features of the utterance, the relational expressions existing in the corresponding process are as follows: ; Among them, represents the final feature of the text modality of the th utterance, represents the final feature of the audio modality of the th utterance, represents the final feature of the visual modality of the th utterance, represents the feature concatenation operation, represents the fusion feature of the th utterance; In the step of concatenating the fused features of the utterance, the speaker embedding vector, and the utterance in the multi-modal dialogue to obtain the final comprehensive feature representation, the relational expressions existing in the corresponding process are as follows: ; Among them, represents the final comprehensive feature representation, represents the concatenation operation.

6. The method for extracting multi-modal dialogue emotion-reason pairs based on speaker relationships according to claim 5, wherein In step 4 above, use a feedforward network to classify and predict the final comprehensive feature representation to obtain the representation in terms of emotion, the prediction result in terms of emotion, the representation in terms of reason, and the prediction result in terms of reason, which specifically includes the following steps: The final comprehensive feature representation is processed by a feed-forward network to obtain the representation in terms of sentiment; The representation in terms of sentiment is predicted to obtain the prediction result in terms of sentiment; The final comprehensive feature representation is processed by a feed-forward network to obtain the representation in terms of reason; The representation in terms of reason is predicted to obtain the prediction result in terms of reason.

7. The method for extracting multi-modal dialogue emotion-reason pairs based on speaker relationships according to claim 6, wherein The final comprehensive feature representation is processed by a feed-forward network to obtain the representation in terms of sentiment. The relational expression existing in the corresponding process is as follows: ; Among them, represents an emotional expression, represents a two-layer feedforward network; In the step of predicting the representation in terms of sentiment to obtain the prediction result in terms of sentiment, the relational expression existing in the corresponding process is as follows: ; Among them, represents the emotion prediction probability distribution, represents being normalized, represents the predicted emotion label, represents returning the maximum value in the independent variable, represents the class index; In the step of processing the final comprehensive feature representation by a feed-forward network to obtain the representation in terms of reason, the relational expression existing in the corresponding process is as follows: ; Among them, represents the indication in terms of reasons, represents another two-layer feedforward network; In the step of predicting the representation in terms of reason to obtain the prediction result in terms of reason, the relational expression existing in the corresponding process is as follows: ; Among them, represents the predicted cause label probability.

8. The method for extracting multi-modal dialogue emotion-cause pairs based on speaker relationships according to claim 7, characterized in that, In step 5, the representation in terms of sentiment and the representation in terms of reason are used for candidate sentiment-reason pairing. The extraction of candidate sentiment-reason pairing is guided by the speaker relationship matrix, and classification prediction is performed using an expert network and a gating network to obtain the prediction result of the sentiment-reason pair. The specific steps are as follows: Candidate pairs are constructed for the reason discourse candidate representations in the representation in terms of sentiment and the representation in terms of reason to obtain the constructed candidate sentiment-reason pairing; The representation in terms of reason and the representation in terms of sentiment are concatenated to obtain the features of the candidate pairing; A role relationship matrix is constructed based on the speaker relationship to obtain the constructed role relationship matrix. The constructed candidate sentiment-reason pair is subjected to a projection operation using the constructed role relationship matrix to obtain the speaker relationship feature; The features of the candidate pairing and the speaker relationship feature are fused using a gating mechanism to obtain the final fused feature representation; The final fused feature representation is input into the expert network to predict the constructed candidate sentiment-reason pairing to obtain the predicted value output by the expert network; The final fused feature representation is input into the gating network to predict the constructed candidate sentiment-reason pairing to obtain the predicted value output by the gating network; A one-hot label is created based on the sentiment-reason pair relationship and combined with the probability distribution output by the gating network for re-prediction to obtain the final comprehensive routing probability distribution; The final comprehensive routing probability distribution is combined with the predicted value output by the expert network for classification prediction to obtain the classification prediction result; Threshold judgment is performed on the classification prediction result to obtain the prediction result of the sentiment-reason pair.

9. The method for extracting multimodal dialogue emotion - reason pairs based on speaker relationship according to claim 8, wherein, Candidate pairs are constructed for the representation in terms of sentiment and the representation in terms of reason to obtain the constructed candidate sentiment-reason pair. The relational expression existing in the corresponding process is as follows: ; Among them, represents a candidate pair composed of sentence and sentence ; represents the speaker of sentence ; represents the speaker of sentence ; represents a candidate representation of the causal discourse; In the step of concatenating the representation in terms of reason and the representation in terms of sentiment to obtain the features of the candidate pairing, the relational expression existing in the corresponding process is as follows: ; Among them, represents the features of the candidate pairing; In the step of constructing a role relationship matrix based on speaker relationships to obtain the constructed role relationship matrix, and using the constructed role relationship matrix to perform a projection operation on the constructed candidate emotion - cause pairs to obtain speaker relationship features, the relational expressions existing in the corresponding process are as follows: ; Among them, represents the constructed role relationship matrix, represents the sentence speaker embedding vector, and respectively represent two different learnable weight matrices, represents the speaker relationship feature; In the step of using a gating mechanism to fuse the features of candidate pairs and speaker relationship features to obtain the final fused feature representation, the relational expressions existing in the corresponding process are as follows: ; Among them, represents the calculated gating value, represents the first learnable parameter, represents the second learnable parameter, represents the final fused feature representation; In the step of inputting the final fused feature representation into an expert network to predict the constructed candidate emotion - cause pairs to obtain the predicted values output by the expert network, the relational expressions existing in the corresponding process are as follows: ; Among them, represents the th expert network's predicted value for the candidate pair composed of sentence and sentence ; represents the predicted value output; represents the 3rd learnable parameter, represents the 4th learnable parameter, represents the activation function, represents the 5th learnable parameter, represents the 6th learnable parameter; In the step of inputting the final fused feature representation into a gating network to predict the constructed candidate emotion - cause pairs to obtain the predicted values output by the gating network, the relational expressions existing in the corresponding process are as follows: ; Among them, represents the distribution over 5 experts, represents the 7th learnable parameter, represents the 8th learnable parameter; In the step of creating a one - hot label based on the emotion - cause pair relationship and performing re - prediction in combination with the probability distribution output by the gating network to obtain the final comprehensive routing probability distribution, the relational expressions existing in the corresponding process are as follows: ; Among them, represents the final comprehensive routing probability distribution, represents the balance parameter, represents the guiding probability distribution based on the pairing relationship category; In the step of combining the final comprehensive routing probability distribution with the predicted values output by the expert network for classification prediction to obtain the classification prediction result, the relational expressions existing in the corresponding process are as follows: ; Among them, represents the final classification prediction probability of the sentiment - cause pairing , and represents the routing probability of the th expert, and represents the prediction probability of the th expert for the sentiment - cause pairing.

10. The method for extracting multimodal dialogue emotion - reason pairs based on speaker relationships according to claim 9, wherein Constructing an emotion prediction loss based on the prediction result in terms of emotion, the relational expressions existing in the corresponding process are as follows: ; Among them, represents the emotion prediction loss, represents the logarithmic function, represents the true emotion label of the In the step of constructing a cause recognition loss based on the prediction result in terms of cause, the relational expressions existing in the corresponding process are as follows: ; Among them, represents the cause recognition loss, represents the true label of the cause of the th utterance; In the step of constructing an emotion - cause pair classification loss based on the final emotion - cause pairs, the relational expressions existing in the corresponding process are as follows: ; Among them, represents the sentiment - cause pairing classification loss, represents the sentiment - cause pairing true label.

Citation Information

Patent Citations

  • Session sentiment analysis method based on multi-granularity fusion and graph convolutional network

    CN115374281A

  • Multi-modal emotion recognition method based on graph convolutional network

    CN116229225A

  • Online collaborative session emotion recognition method and system for uncertain mode missing

    CN119475088A

  • Method and apparatus for synthesising an emotion conveyed on a sound

    EP1256932A2

  • Method and system for multimodal emotion recognition in conversation (ERC) based on graph neural network (GNN)

    US20240355350A1

Cited By

  • Multi-modal emotion six-tuple extraction method based on time sequence heterogeneous graph and fragment decoupling

    CN122414202A

  • Multi-modal sentiment six-tuple extraction method based on time sequence heterogeneous graph and segment decoupling

    CN122414202B