Conversation emotion recognition method and system based on artificial intelligence

By constructing knowledge-enhanced pathways and heterogeneous graph reasoning pathways, and utilizing preliminary emotion probability distributions for modality alignment and graph neural network reasoning, the problem of insufficient semantic alignment between modalities in multimodal emotion recognition is solved, thereby improving the accuracy and robustness of complex emotion recognition.

CN121705910APending Publication Date: 2026-03-20WUXI WANQING HEALTH CARE TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511537250.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-27
Publication Date
2026-03-20

AI Technical Summary

Technical Problem

Existing multimodal emotion recognition methods lack deep semantic alignment and interaction between modalities, cannot effectively capture subtle emotional representations, and lack robustness and accuracy in handling complex emotions.

Method used

An AI-based dialogue emotion recognition method is adopted. Through knowledge enhancement and heterogeneous graph reasoning pathways, a dynamic heterogeneous graph is constructed. The preliminary emotion probability distribution is used for modality alignment and knowledge injection. Graph neural network reasoning and cross-attention collaborative reasoning are then performed to generate comprehensive discriminative features.

Benefits of technology

It enables deeper contextual semantic mining, improves the ability to recognize complex emotions, and enhances the accuracy and robustness of recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121705910A_ABST
    Figure CN121705910A_ABST
Patent Text Reader

Abstract

The invention provides a dialogue emotion recognition method and system based on artificial intelligence, and relates to the technical field of artificial intelligence, and the method comprises the steps: carrying out the basic feature extraction of dialogue data containing text, audio and video modes, and obtaining text, audio, video, knowledge and intention feature vectors; performing modal alignment and knowledge injection on the feature vector through a knowledge enhancement path to generate initial emotion probability distribution; a dynamic heterogeneous graph is constructed through a heterogeneous graph reasoning path, a graph structure is enhanced by using preliminary emotion probability distribution, and an enhanced heterogeneous graph of emotion perception is formed; performing graph neural network reasoning on the enhanced heterogeneous graph of emotion perception to obtain a graph-level representation vector, and performing cross attention collaborative reasoning by taking the initial emotion probability distribution as a query vector and the graph-level representation vector as a key vector and a value vector to obtain attention features; and splicing the initial emotion probability distribution and the attention feature into a comprehensive discrimination feature, and outputting an emotion label through a final classifier.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of artificial intelligence technology, specifically relating to a dialogue emotion recognition method and system based on artificial intelligence. Background Technology

[0002] In fields such as human-computer interaction, social analytics, and intelligent customer service, accurately identifying users' emotional states is crucial. With the development of multimedia technology, users' modes of expression are becoming increasingly diverse, encompassing multiple modalities such as text, voice, and video. Compared to single-modal data, multimodal data can provide richer and more complementary emotional cues.

[0003] Existing multimodal emotion recognition methods suffer from several limitations: First, most methods employ simple feature splicing or post-fusion strategies, failing to achieve deep semantic alignment and interaction between modalities. This results in insufficient information fusion and difficulty in capturing subtle emotional representations. Second, some solutions attempt to introduce external knowledge or graph structures for modeling, but typically employ linear, parallel processing flows. For example, knowledge injection and graph reasoning are treated as two independent modules, with their outputs then simply fused. This "pipeline" or "parallel splicing" architecture lacks coordination and feedback mechanisms between different processing pathways, making it impossible to verify and correct preliminary judgments using subsequent complex contextual reasoning, and also preventing effective prior guidance for subsequent reasoning processes. Third, for implicit, complex, or dynamically changing emotions, existing methods lack a cognitive process that simulates the human "initial perception - contextual verification - comprehensive decision-making," leading to insufficient accuracy and robustness in recognizing complex emotions. Summary of the Invention

[0004] This application provides a dialogue emotion recognition method and system based on artificial intelligence to solve one of the aforementioned technical problems.

[0005] The technical solution adopted in this application is as follows: This application provides an artificial intelligence-based dialogue emotion recognition method, including: Basic feature extraction is performed on dialogue data containing text, audio, and video modalities to obtain text feature vectors, audio feature vectors, video feature vectors, knowledge feature vectors, and intent feature vectors; The feature vectors are modally aligned and knowledge injected through a knowledge enhancement pathway to generate a preliminary emotion probability distribution. A dynamic heterogeneous graph is constructed through a heterogeneous graph reasoning path, and an enhanced heterogeneous graph for emotion perception is formed by using the preliminary emotion probability distribution to enhance the graph structure. Graph neural network inference is performed on the enhanced heterogeneous graph of emotion perception to obtain a graph-level representation vector. Cross-attention collaborative inference is then performed using the preliminary emotion probability distribution as the query vector and the graph-level representation vector as the key vector and value vector to obtain attention features. The preliminary emotion probability distribution is concatenated with the attention features to form a comprehensive discriminative feature, and the emotion label is output through the final classifier.

[0006] According to one embodiment of this application, the basic feature extraction of dialogue data containing text, audio, and video modalities includes: Extract text feature vectors using a pre-trained language model; Use convolutional neural networks to extract audio and video feature vectors; Knowledge feature vectors are obtained by retrieving and encoding from external knowledge bases; Use an intent recognition model to extract intent feature vectors.

[0007] According to one embodiment of this application, the modality alignment and knowledge injection of the feature vector through a knowledge enhancement pathway includes: The audio and video feature vectors are cross-attentioned with the text feature vectors respectively, and the knowledge feature vectors are injected as bias to generate knowledge-enhanced modal features. The fused features are input into a preliminary classifier to generate the preliminary emotion probability distribution.

[0008] According to one embodiment of this application, the construction of a dynamic heterogeneous graph through a heterogeneous graph inference path includes: An initial heterogeneous graph is constructed using text feature vectors, audio feature vectors, video feature vectors, and intent feature vectors as nodes; The preliminary emotion probability distribution is injected into the initial heterogeneous graph as an additional attribute of the nodes to form an enhanced heterogeneous graph for emotion perception.

[0009] According to one embodiment of this application, the computation process of the cross-attention collaborative reasoning includes: The preliminary emotion probability distribution is linearly transformed into a query vector; The graph-level representation vector is linearly transformed into a key vector and a value vector; Calculate the dot product of the query vector and the key vector, scale them, and then apply the softmax function to obtain the attention weights. Multiplying the attention weights by the value vector yields the attention features.

[0010] According to one embodiment of this application, before outputting the sentiment label through the final classifier, the method further includes: The comprehensive discriminative features are normalized and mapped to the emotion category space through a fully connected layer.

[0011] A second aspect of this application provides an artificial intelligence-based dialogue emotion recognition system, comprising: The feature extraction module is used to perform basic feature extraction on dialogue data containing text, audio, and video modalities, and obtain text feature vectors, audio feature vectors, video feature vectors, knowledge feature vectors, and intent feature vectors. The knowledge enhancement pathway module is used to perform modal alignment and knowledge injection on the feature vector to generate a preliminary emotion probability distribution; The heterogeneous graph reasoning pathway module is used to construct a dynamic heterogeneous graph and utilize the preliminary emotion probability distribution to enhance the graph structure, thereby forming an enhanced heterogeneous graph for emotion perception. The collaborative reasoning module is used to perform graph neural network reasoning on the enhanced heterogeneous graph of emotion perception to obtain a graph-level representation vector, and to perform cross-attention calculation using the preliminary emotion probability distribution as the query vector, the graph-level representation vector as the key vector and the value vector to obtain attention features; The decision module is used to concatenate the preliminary emotion probability distribution with the attention features to form a comprehensive discriminative feature, and output the emotion label through the final classifier.

[0012] According to one embodiment of this application, the feature extraction module includes: Extract text feature vectors using a pre-trained language model; Use convolutional neural networks to extract audio and video feature vectors; Knowledge feature vectors are obtained by retrieving and encoding from external knowledge bases; Use an intent recognition model to extract intent feature vectors.

[0013] A third aspect of this application provides a computer-readable storage medium having a program stored thereon that, when executed by a processor, implements the steps described in the method.

[0014] A fourth aspect of this application provides an electronic device including a memory, a processor, and a program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps of the method as described. Due to the adoption of the above technical solution, the beneficial effects achieved by this application are as follows: This application breaks through the limitations of traditional parallel models by constructing two functionally related pathways: a "knowledge enhancement pathway" and a "heterogeneous graph reasoning pathway," and utilizing a "preliminary emotion probability distribution" as a global coordinator. Specifically, the preliminary emotion probability distribution is used to enhance the heterogeneous graph structure, enabling the graph neural network to focus on key information related to the preliminary emotion judgment when modeling modal relationships, thereby achieving deeper and more targeted contextual semantic mining.

[0015] This application does not present a simple linear process, but rather a cyclical enhancement process of "initial judgment → graph structure enhancement → contextual reasoning → decision correction". By using the preliminary emotion probability distribution output by the knowledge enhancement pathway as input to the heterogeneous graph reasoning pathway (for graph enhancement), and further as query conditions for the collaborative reasoning module, subsequent complex reasoning can perform targeted evidence retrieval and verification based on the initial judgment, simulating the human cognitive process of repeated deliberation, thereby dynamically optimizing and correcting the emotion recognition results.

[0016] By employing the aforementioned collaborative and cyclical reinforcement architecture, this application can more effectively integrate multimodal shallow features, external common sense knowledge, and dynamic contextual dependencies. The comprehensive discriminative features upon which the final decision relies are explicitly composed of a preliminary emotion probability distribution representing the initial judgment and attention features representing contextual evidence. This ensures that the final emotion label is the result of the synergistic effect of initial perception and deep reasoning, thereby significantly improving the ability to identify implicit, contradictory, or dynamically changing emotions in complex scenarios. Attached Figure Description

[0017] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings: Figure 1 A flowchart illustrating the AI-based dialogue emotion recognition method provided in this application embodiment; Figure 2 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application.

[0018] Figure label: 810, Processor; 820, Communication interface; 830, Memory; 840, Communication bus. Detailed Implementation

[0019] To more clearly illustrate the overall concept of this application, a detailed explanation is provided below with reference to the accompanying drawings.

[0020] Many specific details are set forth in the following description to provide a thorough understanding of this application. However, this application may also be implemented in other ways different from those described herein. Therefore, the scope of protection of this application is not limited to the specific embodiments disclosed below. It should be noted that, unless otherwise specified, the embodiments of this application and the features thereof can be combined with each other.

[0021] In this application, unless otherwise expressly specified and limited, the "above" or "below" of the second feature can mean that the first and second features are in direct contact, or that the first and second features are in indirect contact through an intermediate medium. In the description of this specification, references to terms such as "an embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described can be combined in any suitable manner in one or more embodiments or examples. Example

[0022] like Figure 1 As shown, the AI-based dialogue emotion recognition method includes: Basic feature extraction is performed on dialogue data containing text, audio, and video modalities to obtain text feature vectors, audio feature vectors, video feature vectors, knowledge feature vectors, and intent feature vectors.

[0023] As described above, this step aims to extract standardized feature representations that are machine-recognizable and processable from raw, unstructured multimodal dialogue data, laying the data foundation for subsequent collaborative reasoning. Specifically, this process is not a simple data transformation, but rather employs specialized preprocessing and feature extraction strategies tailored to the physical characteristics and semantic connotations of different modalities. For text modalities, the core is to understand its vocabulary, grammar, and contextual semantics, thereby transforming it into a dense vector that retains core semantic information. For audio and video modalities, it is necessary to capture key patterns related to the speaker's emotional state from continuous signals and frame sequences, such as tone and rhythm changes in audio, and facial muscle movements and micro-expressions in video. Furthermore, this method innovatively introduces the extraction of "knowledge feature vectors" and "intent feature vectors," which surpasses traditional feature extraction methods based solely on raw signals. "Knowledge feature vectors" provide the model with background knowledge and potential basis for logical reasoning by retrieving and encoding relevant information from external structured or unstructured knowledge sources (such as common sense knowledge bases and domain dictionaries); "Intent feature vectors" are dedicated to identifying the behavioral purpose or dialogue goal behind user utterances, providing key context for understanding the motivation for emotion generation.

[0024] For example, in a dialogue between a user and e-commerce customer service, the system might receive the following user input: a text complaint like "This battery doesn't last long enough," a rapid and loud voice recording, and a video showing the user frowning and shaking their head. In this example, the basic feature extraction steps will work in parallel: semantic analysis of the text "This battery doesn't last long enough" will extract core concepts such as "battery" and "not lasting long enough," along with their negative emotional tendencies, forming a text feature vector; the audio signal will be analyzed to extract its pitch, speed, and loudness features, forming an audio feature vector representing an "angry" or "excited" tone; the video frame sequence will be analyzed to capture key action units such as frowning and shaking the head, forming a video feature vector representing a "dissatisfied" expression; simultaneously, typical battery life information for this model will be retrieved from the product knowledge base, forming a knowledge feature vector as background reference; and the user's current utterance will be determined to be an "inquiry" rather than a "complaint," forming an intent feature vector. These five feature vectors together constitute a multi-dimensional, structured description of the user's current state.

[0025] It should be noted that, in specific implementation scenarios, based on the above solutions, the extraction of text feature vectors can be achieved through pre-trained models based on the Transformer architecture (such as BERT, ERNIE, etc.), or through recurrent neural networks or convolutional neural networks; the extraction of audio and video feature vectors can be achieved through convolutional neural networks, recurrent neural networks, or dedicated feature extractors (such as Mel filter banks for audio, and optical flow methods for video).

[0026] In specific implementation scenarios, based on the above solutions, the source of the knowledge feature vector is not limited to the product knowledge base, but can be extended to general knowledge graphs (such as CN-DBpedia), domain encyclopedias, or even context-related knowledge constructed in real time from massive texts through entity linking technology.

[0027] In specific implementation scenarios, based on the above solution, the acquisition of the intent feature vector can be an independent classification module, or it can share the underlying model with the text feature extraction process and be output synchronously through multi-task learning. The intent category can be predefined (such as inquiry, complaint, praise, query status, etc.) and can be dynamically expanded according to the application scenario.

[0028] In specific implementation scenarios, this method can also be applied to the problem of out-of-vocabulary (OOV) words in text, based on the above scheme. For example, it can be processed by subword units or character-level encoding. Furthermore, the architecture of this step has the ability to be extended to other modalities (such as physiological signals, haptic feedback, etc.), simply by connecting the corresponding feature extractor.

[0029] The feature vectors are modally aligned and knowledge is injected through a knowledge enhancement pathway to generate a preliminary emotion probability distribution.

[0030] As mentioned above, this step is the primary core pathway in the "dual-path collaborative" architecture of this application. Its purpose is to solve the semantic gap problem of multimodal data and to use external knowledge to provide prior logical support for emotion judgment, thereby generating a preliminary emotion judgment that combines modal consistency and knowledge rationality. Specifically, "modal alignment" is not simply feature splicing, but rather refers to using an interactive attention mechanism to allow features of non-textual modalities (such as audio and video) to align and calibrate towards the semantic anchor of the textual modality. For example, the model learns that the audio feature of "high pitch" should correspond to the emotional intensity expressed by the "exclamation mark" in the text, establishing comparability and correlation between features of different modalities at the semantic level. "Knowledge injection," on the other hand, involves incorporating structured information related to the current dialogue content (such as "short battery life is a negative feature" and "frowning usually indicates dissatisfaction") obtained from an external knowledge base into feature calculation in a learnable way (e.g., as a bias term or through a gating mechanism), thereby giving the model the ability to reason based on common sense beyond the surface patterns of the data. Ultimately, generating a preliminary emotion probability distribution means that this pathway will output a probability vector covering all preset emotion categories (such as joy, anger, sorrow, fear, etc.). This result is not the final decision, but rather serves as a "prior guide" and "query focus" for subsequent pathways to conduct in-depth reasoning, and is the starting point of the entire loop enhancement process.

[0031] For example, continuing from the previous example, a user complains, "This battery doesn't last long at all," accompanied by a hurried voice and a video of them frowning. The specific workflow of the knowledge enhancement pathway here could be as follows: First, modal alignment is performed, calculating the cross-attention between audio and video features and text features, respectively. This results in enhanced audio and video features that contain the original audio / video information but have been "filtered" and "interpreted" by text semantics. Next, knowledge injection is performed. The knowledge feature "short battery life is a product defect," obtained from the knowledge base, is incorporated as a reinforcing signal into the generation of the fused features, guiding the model to associate the current information with negative emotions. Finally, the fused features obtained after alignment and knowledge injection are input into a preliminary classifier (e.g., a fully connected neural network). This classifier, after integrating all information, outputs a preliminary emotion probability distribution, such as anger: 0.75, disappointment: 0.20, neutral: 0.05. This P1 distribution quantifies the initial judgment based on current direct information and common sense, indicating that the user is likely in an "angry" state.

[0032] A dynamic heterogeneous graph is constructed through a heterogeneous graph reasoning path, and an enhanced heterogeneous graph for emotion perception is formed by utilizing the preliminary emotion probability distribution to enhance the graph structure.

[0033] As described above, the initial emotion judgment is verified, corrected, and deepened using global contextual information, resulting in a more robust emotion representation. Specifically, this step comprises two closely connected sub-processes. First, "graph neural network inference on the enhanced heterogeneous graph of emotion perception" refers to using the message passing mechanism of graph neural networks to allow nodes representing different modalities and intentions in the graph to interact with their neighboring nodes in multiple rounds. Through this process, the features of each node are enriched and adjusted by its context. For example, a text node initially judged as "angry" may be influenced by connected audio nodes representing "hesitation" and video nodes representing "confusion," thus refining its representation. Finally, graph readout operations aggregate the updated information of all nodes into a graph-level representation vector that represents the entire global context of the dialogue. Subsequently, the innovative aspect of this step is "cross-attention collaborative inference using the initial emotion probability distribution as the query vector," which constructs a collaborative inference module. This module uses the initial emotion probability distribution representing the initial judgment as the query vector and the graph-level representation vector containing rich contextual relationships as both the key and value vectors. Its computational process simulates human reflective behavior: starting with an initial judgment (query), it actively searches for relevant or contradictory evidence in the global context (key), and extracts key information from the context (value) that supports or corrects the initial judgment based on the weights of this evidence (attention distribution), ultimately outputting a refined attention feature. This process establishes a direct and learnable dialogue channel between the initial judgment and the complex context.

[0034] For example, continuing with the intelligent vehicle scenario, the system has generated a preliminary emotion probability distribution P1, where irritability (0.65) is dominant. Simultaneously, the system's constructed emotion perception enhancement heterogeneous graph includes a text node for "endless rain," an audio node for "sigh," a video node for "frowning," and an "intent: complain" node, with strong connections between these nodes. The graph neural network inference process allows these nodes to exchange information: the "irritability" text node reinforces the negative attribute of "sigh," but simultaneously, the "frowning" node may receive information from the "intent: complain" node, making its representation more biased towards "dissatisfaction" rather than "anger." The final aggregated graph-level representation vector is a nuanced global emotion context that integrates all modal details and intents. Subsequently, the collaborative inference module begins its work: it uses P1 (irritability: 0.65) as the query to examine the entire graph context (keys and values). Through computation, it may find that the contextual evidence strongly supports "irritability," but also contains some elements of "helplessness." Therefore, the attention features it generates are a new and more precise emotional representation that amplifies the "annoyed" subject while incorporating subtle differences in "helplessness," providing a key basis for the final decision.

[0035] Graph neural network inference is performed on the enhanced heterogeneous graph of emotion perception to obtain a graph-level representation vector. Cross-attention collaborative inference is then performed using the preliminary emotion probability distribution as the query vector and the graph-level representation vector as the key vector and value vector to obtain attention features.

[0036] As described above, this step is the core component for realizing the process from local perception to global reasoning and performing iterative optimization. Its purpose is to leverage the rich contextual relationships carried by the graph structure to deeply verify and refine the initial emotion judgment. Specifically, the process first uses a graph neural network for reasoning on the "enhanced heterogeneous graph of emotion perception" constructed in the previous steps. This heterogeneous graph contains nodes of various types, including text, audio, video, and intent, and the connections between nodes have been enhanced by the initial emotion probability. The graph neural network, through a multi-round message passing mechanism, allows the features of each node in the graph to interact and merge with its neighboring nodes, thereby integrating isolated modal and intent information into a coherent semantic network that can represent the global dialogue context. By aggregating all updated node features (e.g., using global pooling or attention pooling operations), a "graph-level representation vector" that condenses the information of the entire graph is obtained. This vector carries the comprehensive emotion semantics modulated by contextual relationships. Subsequently, the crucial "cross-attention collaborative reasoning" is executed: the innovation here lies in using the "preliminary emotional probability distribution" representing the initial intuitive judgment as the query vector, while the "graph-level representation vector" containing complex contextual evidence is used as both the key vector and the value vector. This design simulates the human thought process of "boldly hypothesizing and carefully verifying"—starting from the initial judgment (query), actively retrieving relevant or contradictory clues in the global context (key), and extracting key information from the context (value) that can support, correct, or refute the initial judgment based on the importance of the retrieved clues (i.e., attention weights), ultimately outputting an "attention feature" refined from contextual evidence.

[0037] For example, in an online education scenario, a student's answer to a question is brief (text: "I don't know"), but audio features detect a slight tremor in their voice, video features show shifty eyes, and the initial emotion probability distribution P1 indicates "nervousness" as the highest probability. When the graph neural network infers from the enhanced heterogeneous graph in this context, the "brief text" node receives signals reinforcing "nervousness" from the "trembling audio" and "shifty eyes" nodes, and may also receive information from the "teacher-student relationship" knowledge node that "the student may feel stressed in this context." The resulting graph-level representation vector is a semantic carrier containing the comprehensive context of "possibly unable to answer due to nervousness." The subsequent cross-attention collaborative inference uses P1 ("nervousness" as the primary query) to comprehensively examine this graph-level context (as key and value). Through computation, the collaborative inference module may find contextual evidence strongly supporting the "nervousness" judgment, but also weak signals of "confusion" (e.g., captured from a modal interaction not emphasized by P1). Thus, the attentional characteristics it generates are a new emotional representation with greater depth, which both reinforces the subjective judgment of "tension" and incorporates the subtle difference of "confusion".

[0038] The preliminary emotion probability distribution is concatenated with the attention features to form a comprehensive discriminative feature, and the emotion label is output through the final classifier.

[0039] As described above, the two complementary types of information generated by the aforementioned pathways—preliminary judgments representing direct perception and refined evidence representing contextual reasoning—are organically integrated to make a final emotion determination. Specifically, "concatenating the preliminary emotion probability distribution with attention features into a comprehensive discriminative feature" means connecting these two vectors, originating from different processing paths and carrying different semantic information, end-to-end along the feature dimension to create a new, more comprehensive feature vector. The preliminary emotion probability distribution carries the "first impression" after knowledge enhancement and modality alignment, while the attention features contain "key evidence" extracted from the global context by the graph neural network, capable of verifying or correcting this first impression. This concatenation operation preserves the complete decision-making chain information from intuitive perception to deep reasoning to the greatest extent possible. Subsequently, "outputting emotion labels through the final classifier" means inputting the aforementioned comprehensive discriminative feature into a trained classification network. This network, through its internal hierarchical nonlinear computation, learns to map the final, most reliable emotion category from this composite feature. The design of this final classifier takes into account the complex, non-linear interaction between the initial judgment and contextual evidence, thereby deriving a more robust final sentiment judgment that integrates all available information.

[0040] For example, following the aforementioned intelligent vehicle scenario, the system has already obtained a preliminary emotion probability distribution P1 (e.g., tension: 0.70, confusion: 0.25, other: 0.05) and attention features obtained through collaborative inference (the latter may reinforce the clue of "confusion due to misunderstanding of the problem"). In this step, the system concatenates these two vectors to form a longer comprehensive discriminative feature vector. This new vector simultaneously contains the initial intuition that "the driver is likely to be tension" and the deep discovery that "there is strong evidence pointing to confusion in the context." Subsequently, this comprehensive feature is fed into the final classifier. Through its complex internal parameter calculations, the classifier may determine that "confusion" has a higher evidence weight, or determine it as a complex emotion of "a mixture of tension and confusion," thus outputting a more accurate emotion label different from the preliminary judgment P1, such as "confusion," or, in a classification system that supports multiple labels, outputting two labels, "tension" and "confusion," along with their confidence levels.

[0041] According to one embodiment of this application, the basic feature extraction of dialogue data containing text, audio, and video modalities includes: Extract text feature vectors using a pre-trained language model; Use convolutional neural networks to extract audio and video feature vectors; Knowledge feature vectors are obtained by retrieving and encoding from external knowledge bases; Use an intent recognition model to extract intent feature vectors.

[0042] As mentioned above, for the text modality, a pre-trained language model is used to extract text feature vectors. This process inputs the original text into a deep learning model that has been pre-trained on massive amounts of text data. Utilizing the model's deep semantic understanding capabilities, the text is converted into a fixed-length, dense numerical vector that effectively captures key semantic information and contextual relationships within the text.

[0043] For audio modalities, convolutional neural networks are used to extract audio feature vectors. First, the original audio waveform signal is converted into a time-frequency representation such as a spectrogram. Then, this is used as a two-dimensional input, and the convolutional neural network automatically learns and extracts emotion-related audio features, such as pitch, rhythm, and energy changes. Finally, the feature vector representing the audio characteristics is output.

[0044] For video modalities, convolutional neural networks are also used to extract video feature vectors. This process processes the frame-by-frame images in the video sequence, extracting visual features through convolutional neural networks. These features can capture facial expressions, body postures, and visual scene information related to emotional expression, and then aggregate them to form feature vectors representing visual content.

[0045] Simultaneously, knowledge feature vectors are retrieved and encoded from an external knowledge base. Key entities or concepts are identified based on the dialogue content, and queries and retrievals are performed in a structured external knowledge base accordingly. The acquired structured knowledge is converted into numerical vectors by an encoder, thereby introducing external common sense and logical background into emotion recognition.

[0046] Furthermore, an intent recognition model is used to extract intent feature vectors. Natural language processing techniques are used to analyze the dialogue text to identify the user's potential intent or behavioral purpose in the current conversation, such as inquiries, complaints, or suggestions. This intent category information is encoded into a feature vector, providing supplementary information for understanding the motivations behind emotions.

[0047] According to one embodiment of this application, the modality alignment and knowledge injection of the feature vector through a knowledge enhancement pathway includes: The audio and video feature vectors are cross-attentioned with the text feature vectors respectively, and the knowledge feature vectors are injected as bias to generate knowledge-enhanced modal features. The fused features are input into a preliminary classifier to generate the preliminary emotion probability distribution.

[0048] As described above, firstly, the audio and video feature vectors are cross-attentionally calculated with the text feature vectors, respectively. During this calculation, the text feature vector is used as the reference query, and attention interactions are performed with both the audio and video feature vectors to determine key information related to text semantics within the non-text modality features, thus achieving alignment of different modalities in the semantic space.

[0049] Building upon this modality alignment, knowledge feature vectors are injected as bias terms into the attention calculation process. By integrating knowledge feature vectors into the generation of attention weights or the feature fusion stage in a learnable manner, external knowledge can guide the modality alignment process, strengthen knowledge-related feature representations, and generate knowledge-enhanced audio and video modality features.

[0050] Subsequently, the knowledge-enhanced audio modal features and video modal features are fused with the original text feature vectors to form a unified fused feature representation.

[0051] Finally, the fused features are input into a preliminary classifier. This preliminary classifier maps the fused features to the emotion category space through a fully connected layer and a normalization operation, and outputs a probability distribution representing the preliminary emotion judgment, which is the preliminary emotion probability distribution.

[0052] According to one embodiment of this application, the construction of a dynamic heterogeneous graph through a heterogeneous graph inference path includes: An initial heterogeneous graph is constructed using text feature vectors, audio feature vectors, video feature vectors, and intent feature vectors as nodes; The preliminary emotion probability distribution is injected into the initial heterogeneous graph as an additional attribute of the nodes to form an enhanced heterogeneous graph for emotion perception.

[0053] As described above, firstly, an initial heterogeneous graph is constructed using text feature vectors, audio feature vectors, video feature vectors, and intent feature vectors as different nodes in the heterogeneous graph. The connections between nodes in the graph are established based on the semantic and temporal relationships between modalities, forming the graph structure foundation that reflects multimodal interaction relationships.

[0054] Then, the preliminary emotion probability distribution is injected into the initial heterogeneous graph as an additional attribute of the nodes. Specifically, the preliminary emotion probability value corresponding to each node is concatenated or weighted and fused with the original feature vector of that node, so that each node in the graph carries emotion perception information, thereby forming an enhanced heterogeneous graph of emotion perception.

[0055] This enhancement allows the graph structure to simultaneously consider modal features and preliminary sentiment judgments in subsequent processing, providing a sentiment-oriented graph structure foundation for deep reasoning based on graph neural networks.

[0056] According to one embodiment of this application, the computation process of the cross-attention collaborative reasoning includes: The preliminary emotion probability distribution is linearly transformed into a query vector; The graph-level representation vector is linearly transformed into a key vector and a value vector; Calculate the dot product of the query vector and the key vector, scale them, and then apply the softmax function to obtain the attention weights. Multiplying the attention weights by the value vector yields the attention features.

[0057] As described above, firstly, the preliminary emotion probability distribution is linearly transformed through a learnable weight matrix and mapped to the query vector space to form a query vector.

[0058] Simultaneously, the graph-level representation vector is linearly transformed through two different learnable weight matrices to generate corresponding key vectors and value vectors.

[0059] Next, the dot product of the query vector and the key vector is calculated. The result is then scaled by dividing the result by the square root of the vector dimension. Finally, the scaled result is normalized using the softmax function to obtain the attention weight distribution.

[0060] Finally, the obtained attention weights and value vectors are weighted and summed to generate the final attention features.

[0061] According to one embodiment of this application, before outputting the sentiment label through the final classifier, the method further includes: The comprehensive discriminative features are normalized and mapped to the emotion category space through a fully connected layer.

[0062] As described above, the comprehensive discriminative features are first normalized by adjusting the mean and variance of the feature distribution through layer normalization or batch normalization methods to improve the stability of the feature representation.

[0063] The normalized features are then input into a fully connected layer. Through a combination of linear transformation and nonlinear activation function, the high-dimensional features are mapped to a low-dimensional space corresponding to the number of emotion categories.

[0064] Finally, the output of the fully connected layer is processed by a normalized exponential function to obtain the predicted probability of each emotion category, thus completing the mapping from the feature space to the emotion category space.

[0065] A second aspect of this application provides an artificial intelligence-based dialogue emotion recognition system, comprising: The feature extraction module is used to perform basic feature extraction on dialogue data containing text, audio, and video modalities, and obtain text feature vectors, audio feature vectors, video feature vectors, knowledge feature vectors, and intent feature vectors. The knowledge enhancement pathway module is used to perform modal alignment and knowledge injection on the feature vector to generate a preliminary emotion probability distribution; The heterogeneous graph reasoning pathway module is used to construct a dynamic heterogeneous graph and utilize the preliminary emotion probability distribution to enhance the graph structure, thereby forming an enhanced heterogeneous graph for emotion perception. The collaborative reasoning module is used to perform graph neural network reasoning on the enhanced heterogeneous graph of emotion perception to obtain a graph-level representation vector, and to perform cross-attention calculation using the preliminary emotion probability distribution as the query vector, the graph-level representation vector as the key vector and the value vector to obtain attention features; The decision module is used to concatenate the preliminary emotion probability distribution with the attention features to form a comprehensive discriminative feature, and output the emotion label through the final classifier.

[0066] According to one embodiment of this application, the feature extraction module includes: Extract text feature vectors using a pre-trained language model; Use convolutional neural networks to extract audio and video feature vectors; Knowledge feature vectors are obtained by retrieving and encoding from external knowledge bases; Use an intent recognition model to extract intent feature vectors.

[0067] A third aspect of this application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the method described in any of the first aspects above.

[0068] Figure 2 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 2 As shown, the electronic device may include: a processor 810, a communication interface 820, a memory 830, and a communication bus 840, wherein the processor 810, the communication interface 820, and the memory 830 communicate with each other via the communication bus 840. The processor 810 may call logical instructions in the memory 830 to execute the method in any of the embodiments of the first aspect described above, the method including: Basic feature extraction is performed on dialogue data containing text, audio, and video modalities to obtain text feature vectors, audio feature vectors, video feature vectors, knowledge feature vectors, and intent feature vectors; The feature vectors are modally aligned and knowledge injected through a knowledge enhancement pathway to generate a preliminary emotion probability distribution. A dynamic heterogeneous graph is constructed through a heterogeneous graph reasoning path, and an enhanced heterogeneous graph for emotion perception is formed by using the preliminary emotion probability distribution to enhance the graph structure. Graph neural network inference is performed on the enhanced heterogeneous graph of emotion perception to obtain a graph-level representation vector. Cross-attention collaborative inference is then performed using the preliminary emotion probability distribution as the query vector and the graph-level representation vector as the key vector and value vector to obtain attention features. The preliminary emotion probability distribution is concatenated with the attention features to form a comprehensive discriminative feature, and the emotion label is output through the final classifier.

[0069] Furthermore, the logical instructions in the aforementioned memory 830 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory, random access memory, magnetic disks, or optical disks.

[0070] On the other hand, the present invention also provides a computer program product, the computer program product comprising a computer program, the computer program being able to be stored on a non-transitory computer-readable storage medium, and when the computer program is executed by a processor, the computer being able to perform the methods provided by the above methods, the method comprising: Basic feature extraction is performed on dialogue data containing text, audio, and video modalities to obtain text feature vectors, audio feature vectors, video feature vectors, knowledge feature vectors, and intent feature vectors; The feature vectors are modally aligned and knowledge injected through a knowledge enhancement pathway to generate a preliminary emotion probability distribution. A dynamic heterogeneous graph is constructed through a heterogeneous graph reasoning path, and an enhanced heterogeneous graph for emotion perception is formed by using the preliminary emotion probability distribution to enhance the graph structure. Graph neural network inference is performed on the enhanced heterogeneous graph of emotion perception to obtain a graph-level representation vector. Cross-attention collaborative inference is then performed using the preliminary emotion probability distribution as the query vector and the graph-level representation vector as the key vector and value vector to obtain attention features. The preliminary emotion probability distribution is concatenated with the attention features to form a comprehensive discriminative feature, and the emotion label is output through the final classifier.

[0071] In another aspect, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to perform the methods provided by the above methods, the method comprising: Basic feature extraction is performed on dialogue data containing text, audio, and video modalities to obtain text feature vectors, audio feature vectors, video feature vectors, knowledge feature vectors, and intent feature vectors; The feature vectors are modally aligned and knowledge injected through a knowledge enhancement pathway to generate a preliminary emotion probability distribution. A dynamic heterogeneous graph is constructed through a heterogeneous graph reasoning path, and an enhanced heterogeneous graph for emotion perception is formed by using the preliminary emotion probability distribution to enhance the graph structure. Graph neural network inference is performed on the enhanced heterogeneous graph of emotion perception to obtain a graph-level representation vector. Cross-attention collaborative inference is then performed using the preliminary emotion probability distribution as the query vector and the graph-level representation vector as the key vector and value vector to obtain attention features. The preliminary emotion probability distribution is concatenated with the attention features to form a comprehensive discriminative feature, and the emotion label is output through the final classifier.

[0072] For any parts not mentioned in this application, existing technologies may be used or referenced.

[0073] The various embodiments in this specification are described in a progressive manner. The same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on describing the differences from other embodiments.

[0074] The above description is merely an embodiment of this application and is not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.

Claims

1. A dialogue emotion recognition method based on artificial intelligence, characterized in that, include: Basic feature extraction is performed on dialogue data containing text, audio, and video modalities to obtain text feature vectors, audio feature vectors, video feature vectors, knowledge feature vectors, and intent feature vectors; The feature vectors are modally aligned and knowledge injected through a knowledge enhancement pathway to generate a preliminary emotion probability distribution. A dynamic heterogeneous graph is constructed through a heterogeneous graph reasoning path, and an enhanced heterogeneous graph for emotion perception is formed by using the preliminary emotion probability distribution to enhance the graph structure. Graph neural network inference is performed on the enhanced heterogeneous graph of emotion perception to obtain a graph-level representation vector. Cross-attention collaborative inference is then performed using the preliminary emotion probability distribution as the query vector and the graph-level representation vector as the key vector and value vector to obtain attention features. The preliminary emotion probability distribution is concatenated with the attention features to form a comprehensive discriminative feature, and the emotion label is output through the final classifier.

2. The method according to claim 1, characterized in that, The basic feature extraction of dialogue data containing text, audio, and video modalities includes: Extract text feature vectors using a pre-trained language model; Use convolutional neural networks to extract audio and video feature vectors; Knowledge feature vectors are obtained by retrieving and encoding from external knowledge bases; Use an intent recognition model to extract intent feature vectors.

3. The method according to claim 1, characterized in that, The modality alignment and knowledge injection of the feature vector through a knowledge enhancement pathway includes: The audio and video feature vectors are cross-attentioned with the text feature vectors respectively, and the knowledge feature vectors are injected as bias to generate knowledge-enhanced modal features. The fused features are input into a preliminary classifier to generate the preliminary emotion probability distribution.

4. The method according to claim 1, characterized in that, The construction of a dynamic heterogeneous graph through a heterogeneous graph inference path includes: An initial heterogeneous graph is constructed using text feature vectors, audio feature vectors, video feature vectors, and intent feature vectors as nodes; The preliminary emotion probability distribution is injected into the initial heterogeneous graph as an additional attribute of the nodes to form an enhanced heterogeneous graph for emotion perception.

5. The method according to claim 1, characterized in that, The computational process of the cross-attention collaborative reasoning includes: The preliminary emotion probability distribution is linearly transformed into a query vector; The graph-level representation vector is linearly transformed into a key vector and a value vector; Calculate the dot product of the query vector and the key vector, scale them, and then apply the softmax function to obtain the attention weights. Multiplying the attention weights by the value vector yields the attention features.

6. The method according to claim 1, characterized in that, Before outputting the sentiment label through the final classifier, the following steps are also included: The comprehensive discriminative features are normalized and mapped to the emotion category space through a fully connected layer.

7. An artificial intelligence-based dialogue emotion recognition system, characterized in that, include: The feature extraction module is used to perform basic feature extraction on dialogue data containing text, audio, and video modalities, and obtain text feature vectors, audio feature vectors, video feature vectors, knowledge feature vectors, and intent feature vectors. The knowledge enhancement pathway module is used to perform modal alignment and knowledge injection on the feature vector to generate a preliminary emotion probability distribution; The heterogeneous graph reasoning pathway module is used to construct a dynamic heterogeneous graph and utilize the preliminary emotion probability distribution to enhance the graph structure, thereby forming an enhanced heterogeneous graph for emotion perception. The collaborative reasoning module is used to perform graph neural network reasoning on the enhanced heterogeneous graph of emotion perception to obtain a graph-level representation vector, and to perform cross-attention calculation using the preliminary emotion probability distribution as the query vector, the graph-level representation vector as the key vector and the value vector to obtain attention features; The decision module is used to concatenate the preliminary emotion probability distribution with the attention features to form a comprehensive discriminative feature, and output the emotion label through the final classifier.

8. The system according to claim 7, characterized in that, The feature extraction module includes: Extract text feature vectors using a pre-trained language model; Use convolutional neural networks to extract audio and video feature vectors; Knowledge feature vectors are obtained by retrieving and encoding from external knowledge bases; Use an intent recognition model to extract intent feature vectors.

9. A computer-readable storage medium having a program stored thereon, characterized in that, When the program is executed by the processor, it implements the steps of the method as described in any one of claims 1-6.

10. An electronic device comprising a memory, a processor, and a program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the method as described in any one of claims 1-6.