A multi-modal fusion emotion and intent joint recognition method and system
Patent Information
- Application Number
- CN202611069329.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-17
- Publication Date
- 2026-09-25
AI Technical Summary
[0005]鉴于此,本发明实施例提供了一种多模态融合的情感与意图联合识别方法及系统,以解决现有技术存在的多模态异构特征缺乏稳定统一的对齐机制、跨任务级联交互局限于浅层静态单向检索而无法实现深层双向逻辑验证、以及模型决策过程缺乏可解释性且缺乏统一时间基准的可视化分析手段的问题
[0016]本发明通过获取多模态对话数据并对其进行特征提取与维度对齐,解决了多模态异构特征难以统一处理的问题,使得文本、音频和视觉三种模态的信息能够在统一的语义空间中进行深度融合,有效消除了单一模态在反讽、双关等复杂语境下的语义歧义。在此基础上,本发明在初级跨模态交互之后进一步采用交换查询深层交互机制,情感分支和意图分支分别使用对方生成的初级交互特征作为查询向量进行深层注意力计算,打破了传统方法始终以自身初始特征为出发点的单向检索模式,实现了两个任务在深层语义空间中的真正双向交叉验证与逻辑纠偏,从而在复杂对话场景下显著提升了情感与意图联合识别的准确性。同时,本发明通过自适应门控调节对深层交互特征和初级交互特征进行动态加权,能够有效过滤强制交互过程中可能引入的不相关噪声,增强了模型在多模态信息互补与冲突并存情况下的鲁棒性,尤其在处理数据长尾分布中的少数类样本时,少数类别的召回率获得明显提升。此外,本发明通过对交换查询深层交互过程中多头注意力权重矩阵的提取、时域重采样和高斯核平滑处理,能够生成在统一时间轴上的动态热力图,直观展示模型在不同模态之间的注意力转移过程,为模型决策逻辑提供了可验证的可视化证据。
Smart Images

Figure CN122817835A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and in particular to a multimodal fusion method and system for joint recognition of emotion and intent. Background Technology
[0002] In the field of artificial intelligence and natural human-computer interaction, emotion recognition and intent recognition are two core tasks for realizing intelligent dialogue systems. Real human dialogue integrates multiple information channels, including textual semantics, voice tone, and facial expressions, with complementarity and redundancy existing between different modalities. Emotion is often the outward manifestation of intent, while intent is the internal driving force of emotion; the two are closely related in cognitive logic. However, traditional methods typically model emotion recognition and intent recognition as two independent tasks, ignoring the strong correlation between the tasks. This can easily lead to contradictory results, such as correctly judging intent but misjudging emotion or vice versa, severely damaging the semantic consistency of the dialogue system.
[0003] With the development of multi-task learning, researchers have begun to explore the joint modeling of emotion and intention. However, existing solutions still have significant shortcomings. At the modality fusion level, the physical properties of text, audio, and visual signals differ greatly, and their feature dimensions, sampling frequencies, and semantic abstraction levels are inconsistent. Simple splicing or averaging fusion is insufficient to retain effective information and easily leads to cross-modal semantic misalignment. At the cross-task interaction level, although the most representative joint recognition framework currently introduces cascading interaction, its deep interaction stage still uses the initial static features of each branch as query vectors. The emotion branch always retrieves intention information based on its own features, and the intention branch does the same. This self-centered unidirectional retrieval mode fails to achieve deep bidirectional cross-validation and logical correction, making it difficult for the model to capture the deep causal coupling between emotion and intention in complex contexts such as irony and puns, resulting in insufficient accuracy and robustness of joint recognition. In addition, existing models generally exhibit black-box characteristics, only outputting classification labels without presenting the decision-making basis. When multimodal information is conflicting or complementary, researchers cannot know which modality the model focuses on. Even if we try to extract attention weights for analysis, the time series lengths and sampling frequencies of each modality are uneven, and direct interpretation will lead to temporal semantic misalignment, making it impossible to form visual evidence with physical reference.
[0004] In summary, existing technologies mainly face three challenges: the lack of a stable and unified alignment mechanism for multimodal heterogeneous features; cross-task cascading interactions remaining at a shallow and static retrieval level; and the lack of interpretability of the model decision-making process and the lack of visualization methods with a unified time benchmark. Summary of the Invention
[0005] In view of this, embodiments of the present invention provide a multimodal fusion-based joint recognition method and system for emotion and intent, in order to solve the problems of existing technologies, such as the lack of a stable and unified alignment mechanism for multimodal heterogeneous features, the limitation of cross-task cascaded interaction to shallow static one-way retrieval which cannot achieve deep two-way logical verification, and the lack of interpretability of the model decision-making process and the lack of a unified time benchmark for visualization analysis.
[0006] On one hand, the present invention provides a multimodal fusion method for joint recognition of emotion and intent, the method comprising: Acquire multimodal dialogue data, wherein the multimodal dialogue data includes at least text modal data, audio modal data, and visual modal data; Feature extraction and dimension alignment are performed on the multimodal dialogue data to obtain multimodal alignment features, which are then used as initial features for the sentiment branch and the intent branch. Based on the initial features of the emotion branch and the initial features of the intention branch, a primary cross-modal interaction is performed to obtain the primary emotion interaction features and the primary intention interaction features. Based on the primary emotional interaction features and the primary intention interaction features, a deep interaction query is performed to obtain deep emotional interaction features and deep intention interaction features; wherein, the deep interaction query includes: using the primary intention interaction features as a query vector and the primary emotional interaction features as keys and values, performing attention calculation to obtain the deep emotional interaction features; and using the primary emotional interaction features as a query vector and the primary intention interaction features as keys and values, performing attention calculation to obtain the deep intention interaction features; Gating is applied based on the deep emotional interaction features and the primary emotional interaction features to obtain an emotional decision vector, which is then input into an emotional classifier to obtain an emotional recognition result. Gating is applied based on the deep intention interaction features and the primary intention interaction features to obtain an intention decision vector, which is then input into an intention classifier to obtain an intention recognition result.
[0007] In some embodiments of the present invention, feature extraction and dimension alignment are performed on the multimodal dialogue data to obtain multimodal aligned features, and the multimodal aligned features are used as initial features for the sentiment branch and initial features for the intent branch, including: The visual modality data is subjected to frame extraction processing, and the resulting image sequence is input into a visual feature extraction network to obtain initial visual features; the audio modality data is input into an audio feature extraction network to obtain initial audio features; the text modality data is input into a text feature extraction network to obtain initial text features. A fixed random projection matrix based on a normal distribution is used to perform dimensionality reduction mapping on the initial visual features and the initial audio features respectively, so that the dimensions are aligned with the dimensions of the initial text features, resulting in aligned visual features and aligned audio features. The aligned visual features, the aligned audio features, and the initial text features are concatenated in dialogue time order to obtain the multimodal aligned features.
[0008] In some embodiments of the present invention, a fixed random projection matrix based on a normal distribution is used to perform dimensionality reduction mapping on the initial visual features and the initial audio features, respectively; wherein, the dimensionality reduction mapping is achieved by performing matrix multiplication operations between the initial visual features or the initial audio features and the fixed random projection matrix.
[0009] In some embodiments of the present invention, primary cross-modal interaction is performed based on the initial features of the emotion branch and the initial features of the intention branch to obtain primary emotion interaction features and primary intention interaction features, including: Using the initial features of the sentiment branch as the query vector and the initial features of the intent branch as the key and value, cross-attention calculation is performed to obtain the primary interaction features of the sentiment. Using the initial features of the intent branch as the query vector and the initial features of the sentiment branch as the key and value, cross-attention calculation is performed to obtain the initial interaction features of the intent.
[0010] In some embodiments of the present invention, the primary emotional interaction features are calculated using the following formula: ; The primary interaction features of the intent are calculated using the following formula: ; in, This represents the primary interactive features of the emotion; This represents the cross-attention calculation function, in which... Represents the query vector. Represents the key vector. Represents a value vector; This indicates the initial characteristics of the emotional branch; This indicates the initial feature of the intent branch; This indicates the primary interactive features of the stated intent.
[0011] In some embodiments of the present invention, the deep emotional interaction features are calculated using the following formula: ; The intent deep interaction features are calculated using the following formula: ; in, This represents the deep emotional interaction characteristics; This represents the cross-attention calculation function, in which... Represents the query vector. Represents the key vector. Represents a value vector; This indicates the primary interactive features of the stated intent; This represents the primary interactive features of the emotion; This indicates the deep interaction features of the stated intent.
[0012] In some embodiments of the present invention, gating adjustment is performed based on the deep emotional interaction features and the primary emotional interaction features to obtain an emotional decision vector, which is achieved through the following formula: ; Based on the deep interaction features and primary interaction features of the intent, gating is performed to obtain the intent decision vector, which is achieved through the following formula: ; in, Represents the emotional decision vector; This represents the deep emotional interaction characteristics; This represents the Sigmoid activation function; This represents the primary interactive features of the emotion; This indicates the initial characteristics of the emotional branch; This represents the intention decision vector; This indicates the deep interactive features of the stated intent; This indicates the primary interactive features of the stated intent; This represents the initial feature of the intent branch.
[0013] In some embodiments of the present invention, the method further includes: During the deep interaction of the exchange query, the original attention weight matrix generated by the multi-head attention layer is extracted; The original attention weight matrices of multiple attention heads are fused by mean fusion to obtain the global attention weight matrix; The global attention weight matrix is resampled in the temporal domain to map the attention weights of each modality to a unified time axis; The attention weights after temporal resampling are smoothed using Gaussian kernels to generate a dynamic heatmap, which is used to display the decision attention distribution of the emotion recognition results and the intent recognition results.
[0014] On the other hand, the present invention provides a multimodal fusion emotion and intention joint recognition system, including a processor, a memory, and a computer program or instructions stored in the memory, wherein the processor is used to execute the computer program or instructions, and when the computer program or instructions are executed, the system implements the steps of any of the methods submitted above.
[0015] On the other hand, the present invention also provides a computer-readable storage medium having a computer program or instructions stored thereon that, when executed by a processor, implement the steps of the method as described in any of the above submissions.
[0016] This invention addresses the challenge of uniformly processing heterogeneous multimodal features by acquiring multimodal dialogue data and extracting and aligning its features. This allows for deep fusion of text, audio, and visual information within a unified semantic space, effectively eliminating semantic ambiguity in complex contexts such as irony and puns. Building upon this foundation, the invention employs a deep interaction mechanism with exchange queries after the initial cross-modal interaction. The sentiment and intent branches use the initial interaction features generated by the other as query vectors for deep attention calculations, breaking away from the traditional one-way retrieval model that always starts with its own initial features. This achieves true bidirectional cross-validation and logical correction between the two tasks in the deep semantic space, significantly improving the accuracy of joint sentiment and intent recognition in complex dialogue scenarios. Furthermore, the invention dynamically weights deep and initial interaction features through adaptive gating, effectively filtering irrelevant noise introduced during forced interaction. This enhances the model's robustness in situations where multimodal information complementarity and conflict coexist, particularly significantly improving recall for minority class samples in long-tailed data distributions. Furthermore, by extracting the multi-head attention weight matrix, performing temporal resampling, and applying Gaussian kernel smoothing during the deep interaction of exchange queries, this invention can generate a dynamic heatmap on a unified time axis, intuitively demonstrating the attention transfer process between different modalities of the model, and providing verifiable visual evidence for the model's decision-making logic.
[0017] Additional advantages, objects, and features of the invention will be set forth in part in the description which follows, and will also become apparent in part to those skilled in the art upon studying the description, or may be learned by practice of the invention. The objects and other advantages of the invention can be realized and obtained by means of the structures specifically pointed out in the description and drawings.
[0018] Those skilled in the art will understand that the objectives and advantages achievable with the present invention are not limited to those specifically described above, and that the above and other objectives achievable with the present invention will become clearer from the following detailed description. Attached Figure Description
[0019] The accompanying drawings, which are included to provide a further understanding of the invention and form part of this application, are not intended to limit the scope of the invention. In the drawings: Figure 1 This is a schematic diagram illustrating the steps of a multimodal fusion-based joint recognition method for emotion and intent in one embodiment of the present invention.
[0020] Figure 2 This is a schematic diagram illustrating the principle of deep interaction in query exchange in one embodiment of the present invention.
[0021] Figure 3 This is a schematic diagram of the attention visualization mapping process in one embodiment of the present invention. Detailed Implementation
[0022] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the embodiments and accompanying drawings. Here, the illustrative embodiments and descriptions of this invention are used to explain the invention, but are not intended to limit the invention.
[0023] It should also be noted that, in order to avoid obscuring the invention with unnecessary details, only the structures and / or processing steps closely related to the solution according to the invention are shown in the accompanying drawings, while other details that are not closely related to the invention are omitted.
[0024] It should be emphasized that the term "including / comprises" as used herein refers to the presence of a feature, element, step, or component, but does not exclude the presence or addition of one or more other features, elements, steps, or components.
[0025] It should also be noted that, unless otherwise specified, the term "connection" in this article can refer not only to a direct connection, but also to an indirect connection involving an intermediary.
[0026] In the following description, embodiments of the invention will be illustrated with reference to the accompanying drawings. In the drawings, the same reference numerals represent the same or similar parts, or the same or similar steps.
[0027] It should be emphasized here that the step markers mentioned below are not a limitation on the order of the steps, but should be understood as meaning that the steps can be executed in the order mentioned in the embodiments, or in a different order than in the embodiments, or several steps can be executed simultaneously.
[0028] To address the shortcomings of existing technologies, such as the lack of a stable and unified alignment mechanism for multimodal heterogeneous features, the limitation of cross-task cascading interactions to shallow, static, unidirectional retrieval preventing deep, bidirectional logical verification, and the lack of interpretability and unified time-based visualization analysis methods in the model decision-making process, this invention provides a multimodal fusion-based joint recognition method for emotion and intent. Figure 1 As shown, the method includes the following steps S101~S105: Step S101: Obtain multimodal dialogue data. The multimodal dialogue data includes at least text modal data, audio modal data, and visual modal data.
[0029] Step S102: Perform feature extraction and dimension alignment on the multimodal dialogue data to obtain multimodal alignment features, and use the multimodal alignment features as the initial features for the sentiment branch and the initial features for the intent branch.
[0030] Step S103: Perform primary cross-modal interaction based on the initial features of the sentiment branch and the initial features of the intent branch to obtain the primary interaction features of sentiment and the primary interaction features of intent.
[0031] Step S104: Perform deep interaction exchange based on the primary emotion interaction features and the primary intent interaction features to obtain deep emotion interaction features and deep intent interaction features. The deep interaction exchange includes: using the primary intent interaction features as the query vector and the primary emotion interaction features as the key and value, performing attention calculation to obtain deep emotion interaction features; and using the primary emotion interaction features as the query vector and the primary intent interaction features as the key and value, performing attention calculation to obtain deep intent interaction features.
[0032] Step S105: Gating is performed based on deep emotional interaction features and primary emotional interaction features to obtain an emotional decision vector, and the emotional decision vector is input into an emotional classifier to obtain an emotional recognition result; gating is performed based on deep intention interaction features and primary intention interaction features to obtain an intention decision vector, and the intention decision vector is input into an intention classifier to obtain an intention recognition result.
[0033] In step S101, the raw input data required for joint emotion and intent recognition is acquired. Multimodal dialogue data includes at least text modal data, audio modal data, and visual modal data. It is understood that the multimodal nature referred to in this invention is not limited to the above three modalities, and may include data from more modalities in other embodiments.
[0034] Multimodal dialogue data originates from the raw recordings of user interactions with the intelligent dialogue system. Specifically, visual modal data consists of a continuous video stream containing facial images of the dialogue participants; audio modal data consists of speech signals recorded synchronously with the video stream; and text modal data consists of dialogue text obtained by transcribing speech through automatic speech recognition (ASR), or text messages directly input during the dialogue.
[0035] The data from the three modalities mentioned above are aligned in the time dimension, meaning that the text semantics, speech tone, and facial expressions at the same moment together constitute a complete discourse expression unit.
[0036] During the model deployment phase, the aforementioned data is acquired in real time through the system's multimodal data acquisition module; during the model training and testing phase, it is read from existing multimodal dialogue datasets. If some application scenarios can only acquire data for a subset of the three modalities, this invention also supports recognition on a subset of the aforementioned three modalities, but complete multimodal input can achieve the best recognition results.
[0037] This step provides a raw, unprocessed source of multimodal information for all subsequent computational processing. Compared to recognition methods that rely solely on a single text modality, the introduction of multimodal data allows the model to simultaneously utilize multi-channel information such as text semantics, speech prosody, and facial expressions, providing an informational foundation for eliminating potential semantic ambiguities that may exist in a single modality.
[0038] In step S102, the original multimodal data obtained in step S101 is converted into a high-dimensional feature vector representation that can be processed by the deep learning model, and the problem of inconsistent feature dimensions caused by the differences in physical properties of the three modalities is solved. Finally, the initial input features shared by the sentiment branch and the intent branch are generated.
[0039] First, feature extraction is performed using separate deep coding networks for each of the three modalities of data.
[0040] For visual modality data, the original video stream is processed by frame extraction at a fixed frame rate. Preferably, this invention uses a frame extraction rate of 3 frames per second (3FPS), which strikes a good balance between preserving the temporal changes in facial expressions and controlling computational overhead. The extracted RGB image sequence is input into a pre-trained visual feature extraction network to extract initial high-dimensional visual features for each frame. Preferably, this invention uses a ResNet-50 network pre-trained on a large-scale image dataset as the visual feature extraction network. This network, through its residual structure, can effectively extract deep semantic features related to facial expressions in the image. The ResNet-50 network outputs a fixed-length feature vector for each frame. Arranging the feature vectors of all frames in chronological order constitutes the initial visual feature sequence.
[0041] For audio modal data, the original one-dimensional audio waveform signal is input into a pre-trained audio feature extraction network to extract acoustic feature sequences. This invention preferably uses the Wav2Vec 2.0 model as the audio feature extraction network. Wav2Vec 2.0 is a large-scale speech pre-training model based on self-supervised learning, capable of converting the original audio waveform into a high-dimensional feature representation rich in speech prosody, emotion, and semantic information. It has excellent capture capabilities for acoustic features closely related to emotional expression in speech, such as intonation, speech rate, and pitch. The feature sequence output by this model is the initial audio feature sequence, with each frame corresponding to a time window in the original audio.
[0042] For text modal data, each word or subword in the dialogue text is input into a pre-trained text feature extraction network to extract contextual semantic features at the word level. This invention preferably uses the RoBERTa-Base model as the text feature extraction network. RoBERTa-Base is a masked language model based on the Transformer architecture, which, after pre-training on large-scale text corpora, can generate word embeddings rich in contextual semantic information. The model outputs a feature sequence corresponding to the length of the input text, with the feature vector at each position fusing the semantic information of the word and its context. Unlike visual and audio features, which require dimensionality reduction, text features directly generate feature representations of the target dimension during the extraction process, without the need for additional dimensionality reduction.
[0043] After extracting features from each modality, it is necessary to address the issue of inconsistent feature vector dimensions across different modalities. Since initial visual and audio features typically have high feature dimensions and varying numerical values, while the Transformer architecture requires input features to have a uniform dimension for subsequent attention calculations, it is necessary to reduce the dimensionality of visual and audio features and map them to the same target alignment dimension as the text features.
[0044] In some embodiments, to achieve the above objectives, the present invention employs a fixed random projection matrix based on a normal distribution to perform dimensionality reduction mapping on the initial visual features and the initial audio features, respectively. Specifically, a fixed random projection matrix is constructed. The dimension of this matrix is ,in, The original dimension of the feature to be reduced. Align the dimensions to the target. Projection matrix. Each element in Independent sampling from normal distribution .
[0045] In some embodiments, the calculation formula for the dimensionality reduction mapping is shown in Equation (1): ; (1) in, This represents the alignment feature obtained after dimensionality reduction mapping; Represents initial visual or initial audio features (depending on the specific dimensionality reduction object); This represents matrix multiplication. The random projection method, based on the Johnson-Lindenstrauss lemma, can preserve the relative distance relationships between samples in the original high-dimensional feature space with a high probability during dimensionality reduction, thus maintaining the core discriminative structure of the original features to the greatest extent. Simultaneously, due to the projection matrix... Once generated during model initialization, the model remains fixed and does not participate in subsequent gradient backpropagation updates. Therefore, compared with using learnable fully connected layers for dimension mapping, this method has advantages such as high parameter efficiency, low computational cost, and less tendency to overfit under small sample conditions.
[0046] Through the above operations, the initial visual features are projected onto the dimension. The aligned visual features are obtained; the initial audio features are also projected onto the dimension. This yields the aligned audio features.
[0047] After aligning the dimensions of each modality's features, the alignment problem of multimodal data over time can be addressed. Since the visual modality is represented by a sequence of extracted images, the audio modality by a sequence of feature frames within a fixed time window, and the text modality by a sequence of words, the sequence lengths and the time span corresponding to each feature vector on the physical time axis are all different. Directly combining these three features in an unordered manner will cause the subsequent cross-modal attention calculation to lose its physical temporal reference.
[0048] In some embodiments, to address the aforementioned temporal alignment problem, this invention introduces a global hash index tree mechanism. Specifically, firstly, the temporal information from the multimodal dialogue data obtained in step S101—namely, the visual features of each frame, the audio features of each frame, and the start and end times corresponding to each word—is stored as the key in the global hash index tree. Then, the three modal features, after dimensional alignment, are arranged according to their absolute temporal order in the dialogue. The visual features, audio features, and text features within the same time period are concatenated along the feature dimension to form a joint feature vector sequence that integrates the information from the three modalities. The resulting product is the multimodal aligned feature.
[0049] After obtaining the multimodal alignment features, they are copied into two copies, which serve as the initial features for the sentiment branch and the initial features for the intent branch, respectively. In this invention, the sentiment branch and the intent branch share the exact same input features in the initial stage to maintain equal access to the original information for both tasks. Subsequently, they gradually differentiate their respective task-specific representations through different interaction paths.
[0050] In step S103, the first round of cross-modal information exchange takes place between the sentiment branch and the intent branch. Specifically, each branch uses its own initial features as a query vector to retrieve information from the initial features of the other branch, generating primary interaction features carrying information from the other branch. This step can be considered a shallow information fusion stage, providing basic input for subsequent deeper exchange queries.
[0051] The core computational mechanism in this step is the Cross-Attention mechanism. Cross-Attention is the core computational unit in the Transformer architecture. Its basic function is as follows: given a query vector sequence (Query, Q), a key vector sequence (Key, K), and a value vector sequence (Value, V), the similarity between Q and K is used as a weight to weight and aggregate the information in V, thereby generating an output feature sequence that is guided by Q and incorporates relevant information from V.
[0052] In this invention, all attention calculations employ a multi-head attention mechanism, which involves running multiple independent attention heads in parallel, each focusing on a different feature subspace. Finally, the outputs of all heads are concatenated and fused to enhance the model's expressive power.
[0053] In this specific step, the emotion branch and the intention branch are calculated independently and symmetrically for cross-attention.
[0054] The initial interaction calculation for the sentiment branch is as follows: using the initial features of the sentiment branch as the query vector and the initial features of the intent branch as the key vector and value vector, the results are fed into the cross-attention function for calculation to obtain the initial interaction features of the sentiment.
[0055] In some embodiments, the primary interaction calculation of the emotion branch is as shown in formula (2): ; (2) in, This represents the primary interactive characteristics of emotions; This represents the cross-attention calculation function, in which... Represents the query vector. Represents the key vector. Represents a value vector; Indicates the initial characteristics of the sentiment branch; Indicates the initial features of the branch.
[0056] Formula (2) means that the sentiment branch takes its own initial features as the starting point for information retrieval, searches for the most relevant information fragments in the initial features of the intention branch, and weights and aggregates these relevant information into its own feature representation, so that the output sentiment primary interaction features retain the original information of the sentiment branch and incorporate the intention information related to sentiment judgment from the intention branch.
[0057] Similarly, the initial interaction calculation of the intent branch adopts symmetrical logic: the initial features of the intent branch are used as the query vector, and the initial features of the sentiment branch are used as the key vector and value vector, which are fed into the cross-attention function for calculation to obtain the initial interaction features of the intent.
[0058] In some embodiments, the initial interaction calculation of the intent branch is as shown in formula (3): ; (3) in, This indicates the initial interactive features of the intent. Other annotations can be found in formula (2), and will not be elaborated here.
[0059] Formula (3) means that the intention branch takes its own features as the starting point, retrieves information related to intention understanding from the features of the emotion branch, and integrates this information into its own feature representation. The output intention primary interaction features retain the original information of the intention branch and also incorporate the emotional information related to intention judgment from the emotion branch.
[0060] Through the two sets of cross-attention calculations described above, the sentiment branch and the intent branch each obtain a set of primary interaction features carrying information about the other. It is understandable that the above calculation process is performed under a multi-head attention mechanism, where each attention head independently calculates its own attention score matrix, and the outputs of multiple heads are concatenated along the feature dimension to form the final interaction features.
[0061] In step S104, unlike the traditional method that still uses the initial features of each branch as the query vector in the deep interaction stage, the present invention forcibly breaks the inherent retrieval pattern. The sentiment branch must use the primary interaction features generated by the intent branch as the query vector, and the intent branch must use the primary interaction features generated by the sentiment branch as the query vector, thereby realizing the two tasks in the deep semantic space for bidirectional cross-validation and logical correction, i.e., the exchange query mechanism.
[0062] like Figure 2 The diagram shown illustrates the principle of deep interaction in query exchange.
[0063] This step builds upon step S103, taking as input the primary emotion interaction features and primary intent interaction features generated in step S103. Crucially, this step completely abandons the traditional paradigm of using the features of each branch as query vectors, and reconstructs the underlying attention retrieval path.
[0064] In the second stage of traditional cascaded interactions, the sentiment branch again uses its initial features as a query vector to retrieve its own primary interaction features. However, in this invention, the sentiment branch is forced to use the primary interaction features generated by the intent branch as a completely new query vector, while using its own primary interaction features as the key and value vectors. Both are then fed into a cross-attention function for calculation to obtain the deep sentiment interaction features.
[0065] In some embodiments, the deep interaction of the sentiment branch exchange query is as shown in formula (4): ; (4) in, This indicates the characteristics of deep emotional interaction. Other annotations have been explained above, so they will not be repeated here.
[0066] Formula (4) means that the emotional task no longer seeks intention information from its own perspective, but rather, from the standpoint of having already initially understood the intention, it re-examines its own emotional characteristics under the direct guidance of the intention information. This mechanism, similar to perspective-taking, forces the emotional branch and the intention branch to perform logical alignment and cross-verification in the deep semantic space, which can effectively capture complex causal relationships such as "because a certain intention is expressed, the emotional expression should be like this" or "although the emotional signal is strong, the intention may be obscure".
[0067] Similarly, the intent branch employs a completely symmetrical logic. It forces the intent branch to use the primary interaction features generated by the sentiment branch as the query vector, while using the intent branch's own primary interaction features as the key and value vectors. Both are then fed into a cross-attention function for computation to obtain the deep interaction features of the intent.
[0068] In some embodiments, the deep interaction of the intent branch's exchange query is as shown in formula (5): ; (5) in, This indicates the intent behind deep interactive features. Other annotations have already been explained, so they will not be repeated here.
[0069] Similarly, this design allows intentional tasks to re-examine their intentional characteristics under the direct guidance of emotional information, achieving reverse verification from emotion to intention.
[0070] It should be noted that the attention calculation for both sets of "exchange queries" above uses the same multi-head attention mechanism as in step S103. Among them, The standard cross-attention calculation function works as follows: First, it calculates the dot product similarity between the query vector and the key vector. After scaling and softmax normalization, it obtains the attention weight matrix. Then, it uses this weight matrix to perform a weighted summation of the value vectors to obtain the final output. In this step, the multi-head attention layer generates raw attention weight matrices that record the degree of correlation between different modalities and tasks. These matrices will be used for visualization analysis in subsequent steps.
[0071] In step S105, the deep interaction features generated in step S104 undergo quality control. An adaptive gating mechanism filters out noise information that may be introduced during the deep forced interaction process. Then, the decision vectors for emotion and intent are fed into their respective classifiers, outputting the final emotion recognition result and intent recognition result. This step also includes the extraction and visualization mapping of attention weights to show the attention distribution during model decision-making.
[0072] In the deep interaction of the exchange query in step S104, since the sentiment branch and the intent branch are forced to exchange information bidirectionally, although this forced mechanism brings the benefit of deep logical verification, it also carries the risk of introducing irrelevant noise. That is, not all intent information is helpful for sentiment judgment, and vice versa. To address this, the present invention deploys an adaptive gating regulator at the end of each branch to evaluate the quality of this deep interaction and dynamically adjust the feature weights.
[0073] Taking the emotion branch as an example, the gating regulator takes the deep emotion interaction features generated in step S104 and the primary emotion interaction features generated in step S103 as inputs, and generates dynamic weight coefficients through the gating unit.
[0074] In some embodiments, the calculation formula is as shown in formula (6): ; (6) in, This represents the final emotional decision vector for the emotional branch; represents the Sigmoid activation function, whose output range is (0, 1), which maps the input to a probability value between 0 and 1; the addition between vectors is element-wise addition (Hadamard addition), that is, the elements at corresponding positions in the two feature vectors are added respectively.
[0075] The core principle of this gating mechanism lies in the fact that the Sigmoid function spontaneously evaluates the quality of a deep interaction based on the combined information of deep interaction features and primary features. Specifically, when the information introduced by the deep interaction is consistent with and mutually reinforces the effective information in the primary features, the Sigmoid output value approaches 1, indicating that the deep interaction feature is amplified and propagated. Conversely, when the deep interaction introduces noise that is irrelevant to or even conflicts with the effective information in the primary features, the Sigmoid output value approaches 0, indicating that the deep interaction feature is suppressed. Through this data-driven adaptive weighting, the model can dynamically adjust its level of trust in the deep interaction results.
[0076] Similarly, the gating adjustment of the intention branch adopts a completely symmetrical calculation method, as shown in formula (7): ; (7) in, This represents the intention decision vector. Other notes have already been explained above, so they will not be repeated here.
[0077] After obtaining the sentiment decision vector and intent decision vector, they are subjected to dimensionality permutation and flattening operations to adapt to the input format requirements of the fully connected classifier. Then, the processed sentiment decision vector is fed into a sentiment classifier, such as a Softmax classifier, which outputs the probability distribution of sentiment categories. The category with the highest probability is selected as the final sentiment recognition result. Similarly, the processed intent decision vector is fed into an intent classifier, such as a Softmax classifier, which outputs the probability distribution of intent categories. The category with the highest probability is selected as the final intent recognition result. Both the sentiment and intent classifiers are fully connected networks, and their output dimensions correspond to the preset number of sentiment categories and the preset number of intent categories, respectively.
[0078] In some embodiments, to visually demonstrate the internal attention distribution of the model when making the aforementioned recognition decisions, and to verify whether the model truly understands the complementary logic among multimodalities, this invention also provides a visualization mechanism that maps the decision logic backward from the algorithm's underlying layer. For example... Figure 3 As shown, the specific implementation of this mechanism includes the following steps S201~S205: Step S201: Extract the original attention weight matrix.
[0079] During the deep interaction of the exchange query in step S104, the multi-head attention layer calculates... This will generate an original attention score matrix that records the similarity between Q and K. For the The original weight matrix of each attention head is calculated as shown in formula (8): ; (8) in, Indicates the first The original weight matrix of each attention head. Each element value in the model represents the degree of attention the model pays to a particular modal information at a given moment. This represents the query vector matrix after the query is swapped, i.e., in the sentiment branch. Or in the intent branch ; This represents the target feature matrix used as the key, i.e., in the sentiment branch. Or in the intent branch ; Indicates the feature dimension; This indicates transpose.
[0080] Step S202: Multi-head mean fusion. After extracting the original weight matrices of all attention heads, normalize and superimpose them, and then fuse them using the mean to obtain the global attention weight matrix. This eliminates local random noise that may be contained in a single attention head, making the final attention distribution more stable and reliable, and reflecting the overall attention tendency of the model.
[0081] Step S203: Temporal resampling mapping.
[0082] Because the sequence lengths of text, audio, and visual modalities differ in the original data, directly presenting attention weights at their respective original resolutions makes comparison and interpretation impossible under a unified physical time reference. To address this issue, this invention introduces a temporal resampling algorithm to forcibly normalize the attention weight results of different modalities to a unified [0,1] standard time axis using interpolation. This time axis uses the start and end times of the entire dialogue segment as the normalization benchmark, enabling modalities of different lengths to be compared under the same temporal reference.
[0083] Step S204: Gaussian kernel smoothing.
[0084] The original attention weights may exhibit drastic jumps or high-frequency noise between adjacent time points. These fluctuations are partly due to the randomness of model calculations and partly due to the sparsity of input data, which is not conducive to reflecting the overall stable attention pattern of the model. To suppress the aforementioned high-frequency noise, this invention introduces a Gaussian kernel to perform convolutional smoothing on the instantaneous attention weights, making the attention distribution curve more continuous and smooth, showing clear upward and downward trends, which is easy for observers to interpret intuitively.
[0085] Step S205: Generating a dynamic heatmap.
[0086] The weighted and smoothed attention weights are mapped to a predefined color lookup table, such as a gradient from blue (low attention) through green (medium attention) to red (high attention), to generate a dynamic heatmap. The horizontal axis of this heatmap is a normalized standard time axis, and the vertical axis represents different modalities, such as text, audio, and visual. The color intensity indicates the degree of attention the model pays to the corresponding modality at a given time.
[0087] By observing heatmaps, one can visually see how the model's attention dynamically shifts between different modalities when processing a multimodal dialogue. For example, does the model focus more on the audio and visual modalities at emotionally charged sentences, and more on the text modalities at sentences with high information density?
[0088] Corresponding to the above method, the present invention also provides a multimodal fusion emotion and intention joint recognition system. The electronic device includes a computer device, which includes a processor and a memory. The memory stores computer instructions, and the processor is used to execute the computer instructions stored in the memory. When the computer instructions are executed by the processor, the electronic device implements the steps of the method described above.
[0089] This invention also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of the aforementioned method. The computer-readable storage medium may be a tangible storage medium, such as random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, register, floppy disk, hard disk, removable storage disk, CD-ROM, or any other form of storage medium known in the art.
[0090] Those skilled in the art will understand that the exemplary components, systems, and methods described in conjunction with the embodiments disclosed herein can be implemented in hardware, software, or a combination of both. Whether implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this invention. When implemented in hardware, it can be, for example, electronic circuits, application-specific integrated circuits (ASICs), appropriate firmware, plug-ins, function cards, etc. When implemented in software, the elements of this invention are programs or code segments used to perform the desired tasks. The programs or code segments can be stored in a machine-readable medium or transmitted over a transmission medium or communication link via data signals carried in a carrier wave.
[0091] It should be clarified that the present invention is not limited to the specific configurations and processes described above and shown in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of the present invention is not limited to the specific steps described and shown. Those skilled in the art can make various changes, modifications, and additions, or change the order of steps, after understanding the spirit of the present invention.
[0092] In this invention, features described and / or illustrated for one embodiment may be used in the same or similar manner in one or more other embodiments, and / or combined with or in place of features of other embodiments.
[0093] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, various modifications and variations of the embodiments of the present invention are possible. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A multimodal fusion method for joint recognition of emotion and intent, characterized in that, The method includes: Acquire multimodal dialogue data, which includes text modal data, audio modal data, and visual modal data; Feature extraction and dimension alignment are performed on the multimodal dialogue data to obtain multimodal alignment features, which are then used as initial features for the sentiment branch and the intent branch. Based on the initial features of the emotion branch and the initial features of the intention branch, a primary cross-modal interaction is performed to obtain the primary emotion interaction features and the primary intention interaction features. Based on the primary emotional interaction features and the primary intention interaction features, a deep interaction query is performed to obtain deep emotional interaction features and deep intention interaction features; wherein, the deep interaction query includes: using the primary intention interaction features as a query vector and the primary emotional interaction features as keys and values, performing attention calculation to obtain the deep emotional interaction features; and using the primary emotional interaction features as a query vector and the primary intention interaction features as keys and values, performing attention calculation to obtain the deep intention interaction features; Gating is applied based on the deep emotional interaction features and the primary emotional interaction features to obtain an emotional decision vector, which is then input into an emotional classifier to obtain an emotional recognition result; and gating is applied based on the deep intention interaction features and the primary intention interaction features to obtain an intention decision vector, which is then input into an intention classifier to obtain an intention recognition result.
2. The multimodal fusion-based joint recognition method for emotion and intent according to claim 1, characterized in that, Feature extraction and dimension alignment are performed on the multimodal dialogue data to obtain multimodal aligned features, which are then used as initial features for the sentiment branch and the intent branch, including: The visual modality data is subjected to frame extraction processing, and the resulting image sequence is input into a visual feature extraction network to obtain initial visual features; the audio modality data is input into an audio feature extraction network to obtain initial audio features; the text modality data is input into a text feature extraction network to obtain initial text features. A fixed random projection matrix based on a normal distribution is used to perform dimensionality reduction mapping on the initial visual features and the initial audio features respectively, so that the dimensions are aligned with the dimensions of the initial text features, resulting in aligned visual features and aligned audio features. The aligned visual features, the aligned audio features, and the initial text features are concatenated in dialogue time order to obtain the multimodal aligned features.
3. The multimodal fusion-based joint recognition method for emotion and intent according to claim 2, characterized in that, A fixed random projection matrix based on a normal distribution is used to perform dimensionality reduction mapping on the initial visual features and the initial audio features, respectively; wherein, the dimensionality reduction mapping is achieved by performing matrix multiplication operation between the initial visual features or the initial audio features and the fixed random projection matrix.
4. The multimodal fusion-based joint recognition method for emotion and intent according to claim 1, characterized in that, Based on the initial features of the emotion branch and the initial features of the intention branch, a primary cross-modal interaction is performed to obtain primary emotion interaction features and primary intention interaction features, including: Using the initial features of the sentiment branch as the query vector and the initial features of the intent branch as the key and value, cross-attention calculation is performed to obtain the primary interaction features of the sentiment. Using the initial features of the intent branch as the query vector and the initial features of the sentiment branch as the key and value, cross-attention calculation is performed to obtain the initial interaction features of the intent.
5. The multimodal fusion-based joint recognition method for emotion and intent according to claim 4, characterized in that, The primary emotional interaction features are calculated using the following formula: ; The initial interaction features of the intent are calculated using the following formula: ; in, This represents the primary interactive features of the emotion; This represents the cross-attention calculation function, in which... Represents the query vector. Represents the key vector. Represents a value vector; This indicates the initial characteristics of the emotional branch; This indicates the initial feature of the intent branch; This indicates the primary interactive features of the stated intent.
6. The multimodal fusion-based joint recognition method for emotion and intent according to claim 1, characterized in that, The deep emotional interaction features are calculated using the following formula: ; The intent deep interaction features are calculated using the following formula: ; in, This represents the deep emotional interaction characteristics; This represents the cross-attention calculation function, in which... Represents the query vector. Represents the key vector. Represents a value vector; This indicates the primary interactive features of the stated intent; This represents the primary interactive features of the emotion; This indicates the deep interaction features of the stated intent.
7. The multimodal fusion-based joint recognition method for emotion and intent according to claim 1, characterized in that, Based on the deep emotional interaction features and the primary emotional interaction features, gating is applied to obtain the emotional decision vector, which is achieved through the following formula: ; Based on the deep interaction features and primary interaction features of the intent, gating is performed to obtain the intent decision vector, which is achieved through the following formula: ; in, This represents the emotion decision vector; This represents the deep emotional interaction characteristics; This represents the Sigmoid activation function; This represents the primary interactive features of the emotion; This indicates the initial characteristics of the emotional branch; This represents the intention decision vector; This indicates the deep interactive features of the stated intent; This indicates the primary interactive features of the stated intent; This represents the initial feature of the intent branch.
8. The multimodal fusion-based joint recognition method for emotion and intent according to claim 1, characterized in that, The method further includes: During the deep interaction of the exchange query, the original attention weight matrix generated by the multi-head attention layer is extracted; The original attention weight matrices of multiple attention heads are fused by mean fusion to obtain the global attention weight matrix; The global attention weight matrix is resampled in the temporal domain to map the attention weights of each modality to a unified time axis; The attention weights after temporal resampling are smoothed using Gaussian kernels to generate a dynamic heatmap, which is used to display the decision attention distribution of the emotion recognition results and the intent recognition results.
9. A multimodal fusion system for jointly recognizing emotion and intent, comprising a processor, a memory, and a computer program or instructions stored in the memory, characterized in that, The processor is configured to execute the computer program or instructions, and when the computer program or instructions are executed, the system implements the steps of the method as described in any one of claims 1 to 8.
10. A computer-readable storage medium having a computer program or instructions stored thereon, characterized in that, When the computer program or instructions are executed by a processor, they implement the steps of the method as described in any one of claims 1 to 8.