A robust multi-modal dialogue understanding method based on heterogeneous graph denoising

CN121658598BActive Publication Date: 2026-09-22UNIV OF ELECTRONICS SCI & TECH OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511732405.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-11-24
Publication Date
2026-09-22
Estimated Expiration
2045-11-24

AI Technical Summary

Technical Problem

[0004](1) 噪声传播问题突出:现有方法通常使用零向量填充缺失模态,并将其与可用模态的特征拼接以构建图节点

Benefits of technology

[0060]本发明的有益效果为:本发明可以应对实际场景中多模态对话数据存在的模态缺失、噪声干扰等问题,提升多模态信息融合与语义理解的鲁棒性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121658598B_ABST
    Figure CN121658598B_ABST
Patent Text Reader

Abstract

The present application belongs to the technical field of dialogue understanding and computer, and particularly relates to a robust multi-modal dialogue understanding method based on heterogeneous graph denoising. The present application aims to solve the problems of missing modalities and noise interference in multi-modal dialogue data in actual scenarios, so as to improve the robustness of multi-modal information fusion and semantic understanding. The method firstly constructs a heterogeneous graph structure that fuses multi-modal features such as text, speech and vision, which is used to represent the multi-modal information in the dialogue and the cross-modal and cross-temporal correlation relationship. Then, a heterogeneous graph denoising network based on diffusion mechanism is used to model and repair the missing modalities or noise features, so as to suppress the noise propagation and enhance the feature consistency. On this basis, through dynamic updating of nodes and edges, efficient fusion of multi-modal information and modeling of context semantics are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the fields of dialogue understanding and computer technology, and specifically relates to a robust multimodal dialogue understanding method based on heterogeneous graph denoising. Background Technology

[0002] Multimodal dialogue understanding aims to comprehensively analyze multiple modalities, including linguistic, visual, and acoustic information, in dialogue to identify key elements such as inherent semantic content and emotional state. Existing research in this field largely relies on the ideal assumption that all modal data is complete and available. However, real-world applications often suffer from missing modal data due to factors such as sensor malfunctions, information occlusion, environmental noise, or privacy restrictions, posing a significant challenge to accurately understanding dialogue semantics. To address the problem of incomplete modal data, several processing methods have been proposed, but these methods are mostly limited to isolated statements or single-speaker scenarios, failing to fully consider the complex interactive dynamics between multiple participants in multi-turn dialogues. Since semantic information in real-world scenarios often evolves gradually with continuous dialogue, these limitations severely restrict the practical application effectiveness of existing technologies.

[0003] In recent years, researchers have begun to explore the problem of modality loss in dialogue contexts and have developed a typical processing paradigm: first, available modal features are concatenated; then, graph neural networks are used for context modeling; and finally, a reconstruction loss function is used to recover missing data. Despite these methods achieving initial progress, the following key shortcomings remain:

[0004] (1) The problem of noise propagation is prominent: Existing methods usually use zero-padded vectors to fill missing modalities and concatenate them with features of available modalities to construct graph nodes. Such zero-padded vectors lack effective semantics and introduce a lot of artificial noise. Since they neither provide cross-modal associations nor disrupt temporal continuity, noise in the graph accumulates continuously during information propagation, ultimately weakening the semantic understanding quality of the model.

[0005] (2) Insufficient multimodal and temporal fusion: Current methods generally rely on early fused features for context modeling, failing to explicitly strengthen fine-grained intermodal associations, thus limiting their full exploitation of multimodal correlations. At the same time, existing models lack an effective mechanism for jointly learning cross-modal interactions and temporal contexts, resulting in poor performance in dialogue scenarios with complex multimodal cues that evolve dynamically. Summary of the Invention

[0006] To address the aforementioned issues, this invention proposes a robust multimodal dialogue understanding method based on heterogeneous graph denoising, targeting multimodal dialogue understanding tasks under incomplete modal conditions, and aiming to achieve robust multimodal and temporal dynamic modeling. Specifically, this invention constructs a heterogeneous graph that optimizes modality-specific representations to preserve the unique features of each modality, while also learning modality-invariant representations to capture shared semantics across modalities. To achieve cross-modal and cross-temporal information propagation, this invention designs edge connections that encode temporal continuity within modalities and semantic relevance between modalities. Based on this graph structure, effective flow of multimodal and temporal information is achieved through graph convolution operations. Furthermore, to suppress noise introduced by incomplete inputs, a diffusion-based module and semantic denoising technique is introduced. By optimizing the graph topology and refining node representations, the robustness of the learned features is improved.

[0007] The technical solution adopted in this invention is:

[0008] A robust multimodal dialogue understanding method based on heterogeneous graph denoising includes the following steps:

[0009] S1. Obtain multimodal training data, specifically:

[0010] Get the dialogue sequence ,in It is the total number of statements. It is the first Statements, express exist The true labels for each category; for each statement Extract a set of multimodal features including language Visual Harmony Modality , , These represent the dimensions of language, visual, and acoustic features, respectively. A modality-deficient scenario is simulated by randomly removing certain modalities, while ensuring that at least one modality is retained for each sample, thus obtaining the multimodal features of each sentence. Represented as , Missing indicator Representing modes exist, This indicates its absence, and ,in ;

[0011] S2. Construct a multimodal dialogue understanding network, including multimodal graph construction, graph denoising, reconstruction, and classification;

[0012] The process of constructing the multimodal graph is as follows: using the obtained multimodal features Construct an undirected heterogeneous graph heterogeneous node set , For a modality-specific set of nodes, Given a modality-invariant node set, unimodal information is encoded through a fully connected network to generate modality-specific features. :

[0013] ,

[0014] in, , These are learnable weights. For the feature dimensions of a graph network, use Initialize modality-specific node embedding This makes each statement have One node;

[0015] Multimodal features spliced ​​as Then the spliced ​​feature sequence Inputting the bidirectional long short-term memory (BiLSTM) network to encode context dependencies yields mode-invariant nodes. :

[0016] ,

[0017] in, , These are trainable parameters, used Initialize modally invariant node embedding Each statement generates a modally invariant node;

[0018] Edge set Includes cross-modal edge sets same modal edge set Cross-modal edge sets The method of establishing is to set each node Other available modalities that can be connected to the same statement:

[0019] ,

[0020] in Simultaneously, cross-modal edge weights are adaptively assigned based on the semantic similarity between node features:

[0021] ,

[0022] in, , Represents the cosine similarity function. It is the sigmoid activation function;

[0023] Same-modal edge set The method of establishment is in each mode Inside, each node Connect to all other available statements in the same dialog:

[0024] ,

[0025] Define the weights of edges with the same modality as:

[0026] ,

[0027] in, Represents a node and Edge weights between them Representing nodes respectively Time and location, It is a hyperparameter;

[0028] Integrate cross-modal and same-modal edge weights into the adjacency matrix. Among them This represents the number of nodes in the graph, and each element of the matrix is ​​represented as:

[0029] ;

[0030] The image denoising process is as follows:

[0031] In the constructed graph Execution graph convolution is used to propagate multimodal and temporal information:

[0032] ,

[0033] in, It is the first Layer input, It is a learnable weight matrix. It is the node degree matrix, which is the result of convolutional stacking of the graph. Layers enable neighborhood aggregation;

[0034] For the constructed graph Structural denoising is performed, specifically by denoising the edge set. Optimize: Define Represents a generic node of any type in the graph. From The derived initial binary adjacency matrix, It is an ideal binary adjacency matrix, representing a connection pattern without structured noise; The aim is to approximate the matrix, first by gradually moving towards Injecting isotropic Gaussian noise to simulate the forward failure process:

[0035] ,

[0036] in, It is the diffusion time The noisy adjacency matrix at that location, Control Time Noise level at:

[0037] ,

[0038] in, and It is a hyperparameter that controls the intensity of noise injection; with Approaching 1, the system systematically destroys the graph structure with increasing noise levels. To reverse this process, a neural network is trained based on... and diffusion time step To predict injected noise The network is characterized by node features , and As input, and outputting an edge validity score, the node is estimated. and The probability of the existence of an edge between them, and the training objective is:

[0039] ,

[0040] in, It is diffusion loss. It is injected into the node and The noise components on the edges between them; during inference, the denoised adjacency matrix is ​​obtained by thresholding the predicted edge probabilities. :

[0041] ,

[0042] in, It is the sigmoid function. It is an indicator function. Represents the nodes after structural denoising and The absence or presence of edges between them; using the denoised adjacency matrix Optimized edge set Thus, the optimized notation is obtained as follows: ;

[0043] Semantic denoising is performed based on the results of neighborhood aggregation and structural denoising. Specifically, the graph convolutional layer... The output is passed through a graph diffusion layer to suppress noise in the node feature space, making The semantics of nodes are gradually refined during the graph diffusion process as follows:

[0044] ,

[0045] in, From The normalized graph Laplacian matrix of the generated encoding structure connections. Controlling the diffusion rate, It is a diffusion step The characteristic matrix at the location; after After the diffusion step, a denoised image output is obtained. Statement The learned features of modality-specific and modality-invariant nodes are represented as follows: ;

[0046] The reconstruction and classification process is as follows: For each statement, the final representation is obtained by fusing the graph output and the initial modality-invariant information.

[0047] ,

[0048] A modality-specific reconstruction module is used to reconstruct the complete data, and a linear transformation is employed to map the final features to the input space:

[0049] ,

[0050] in, It is the complete data of the reconstruction. These are learnable weights. It is modal Feature dimensions, The final representation of the above statement is as follows: For learnable bias;

[0051] Will Input a softmax layer to produce the final predicted class distribution. :

[0052] ,

[0053] in, It is an estimated probability. It is the number of tags. These are learnable weights;

[0054] S3. Train the constructed multimodal dialogue understanding network using multimodal training data, with the optimization objective being:

[0055] ,

[0056] ,

[0057] ,

[0058] in, It is classification loss. It is the reconstruction loss. It's a real label. Identify each statement Available modes, It is a mask indicating the missing location; after training under the set conditions, a trained multimodal dialogue understanding network is obtained.

[0059] S4. Obtain the dialogue sequence to be understood, extract the multimodal features of each sentence, input them into the trained multimodal dialogue understanding network, and obtain the prediction results based on the network's output.

[0060] The beneficial effects of this invention are: it can address the problems of modality loss and noise interference in multimodal dialogue data in real-world scenarios, and improve the robustness of multimodal information fusion and semantic understanding. Attached Figure Description

[0061] Figure 1 This is a schematic diagram of the overall framework of the method of the present invention.

[0062] Figure 2 This is a schematic diagram of structural denoising in the method of the present invention. Detailed Implementation

[0063] To make the objectives, technical solutions, and advantages of this application clearer, the specific embodiments of the present invention will be described in further detail below with reference to the accompanying drawings.

[0064] Consider a dialogue sequence ,in It is the total number of statements. It is the first Statements, express exist The true labels for each category. This invention uses functions. Statement The index is mapped to its speaker. For each statement Extract a set of multimodal features including language ( ), visual ( ) and acoustics ( Modality. , , These represent the dimensions of language, visual, and acoustic features, respectively. This invention simulates real-world modality-deficient scenarios by randomly removing certain modalities, while ensuring that each sample retains at least one modality. Specifically, this invention applies this principle to each sentence... Each mode Introduce a missing indicator . Representing modes exist, This indicates its absence. Therefore, the multimodal features observed for each statement... Represented as To ensure that each statement retains at least one modality, constraints are enforced. ,in The overall structure of the method of this invention is as follows: Figure 1 As shown, the proposed framework comprises three components: multimodal graph construction, graph denoising, and reconstruction and classification, which are described in detail below:

[0065] Multimodal graph construction:

[0066] A dialogue consisting of multiple statements is modeled as an undirected heterogeneous graph. The graph construction of this invention explicitly models the complex temporal and multimodal relationships between statements, while also considering scenarios where these interactions are only partially observable due to modal missingness, in order to robustly handle incomplete input conditions.

[0067] 1) Node Construction: Incomplete multimodal data presents two key challenges: preserving the unique features of each available modality, and extracting context-aware shared semantics that can generalize to missing or noisy inputs. To address these challenges, this invention constructs a heterogeneous node set. Each statement contains two types of nodes: modality-specific nodes. and modally invariant nodes .

[0068] Modality-specific nodes: The multimodal features observed for each statement are represented as follows: To preserve the inherent properties of each modality, this invention first represents each available modality in the statement as a node. Specifically, this invention uses a fully connected network to encode single-modal information to generate modality-specific features:

[0069] (1)

[0070] in , These are learnable weights. This invention uses... Initialize modality-specific node embedding This makes each statement have Each node.

[0071] Modality-invariant nodes: While the aforementioned modality-specific nodes can preserve unimodal cues, they cannot model shared semantics across modalities or contexts, especially under incomplete input conditions. Therefore, this invention introduces additional modality-invariant nodes to learn shared representations that bridge cross-modal and cross-temporal information. Specifically, the input features are first concatenated as follows: Then the spliced ​​feature sequence Input the bidirectional long short-term memory network BiLSTM to encode context dependencies:

[0072] (2)

[0073] here , These are trainable parameters. This invention uses... Initialize modally invariant node embedding Each statement generates a modally invariant node.

[0074] 2) Edge Construction: Based on heterogeneous node representations, the relational structure between statements can be modeled. In the context of incomplete multimodal data, a key challenge lies in constructing an information-rich and robust graph topology, as missing modalities can disrupt intermodal interactions and intramodal temporal continuity. To address this issue, this invention designs a set of weighted edges. It includes two edge types: (1) cross-modal edges (1) Capture complementary interactions between different modalities within the same statement; (2) Same modal edges To maintain temporal continuity across statements within the same modality.

[0075] Cross-modal edges: This invention first introduces cross-modal edges to capture semantic interactions between different modalities. Specifically, this invention will... (in Connecting to other available modalities within the same statement, including observed modality-specific nodes and always-present modality-invariant nodes. The established cross-modal edge set can be represented as:

[0076] (3)

[0077] This design ensures that each node aggregates valid intermodal information while satisfying missing constraints. It also guarantees that modality-invariant nodes act as central hubs for integrating available modalities, enhancing robustness in incomplete scenarios.

[0078] After establishing cross-modal connections in the structure, this invention next designs an adaptive weighting mechanism. The method of this invention does not treat all connections equally, but rather adaptively allocates edge weights based on the semantic similarity between node features. Mathematically, the edge weight is defined as:

[0079] (4)

[0080] here . Represents the cosine similarity function. The sigmoid activation function is used. By applying the sigmoid function to the cosine similarity scores, this invention obtains a soft-edge weighting scheme that transforms the original semantic similarity into smooth and bounded importance scores. In this way, the inter-modal interaction strength is differentiable and adapts to the degree of semantic alignment between modalities. Therefore, graph learning can selectively emphasize semantically aligned modalities while reducing the weights of conflicting or noisy modalities. This is beneficial when the amount of available modal information varies, as it suppresses destructive signals and promotes efficient fusion.

[0081] Same-modal edges: Besides intermodal relationships, maintaining temporal continuity within each modality is also important, especially in scenarios where modalities are missing. Therefore, this invention further constructs same-modal edges that explicitly encode temporal coherence. Specifically, in each modality... Inside, each node Connect to all other available statements in the same dialogue. Formally, a set of same-modal edges can be represented as:

[0082] (5)

[0083] Modal absence often disrupts the natural temporal continuity of single-modal sequences because crucial contextual signals may be missing or unevenly distributed over time. Therefore, this work proposes a time- and context-sensitive edge weighting mechanism to guide intramodal information flow. By incorporating exponential decay of temporal distance, the model strengthens connections between temporally and semantically similar statements, helping each modality maintain internal coherence even in the absence of certain context. Specifically, the intramodal edge weights are calculated as follows:

[0084] (6)

[0085] in Represents a node and The edge weights between them. Representing nodes respectively The time and location. It is a hyperparameter. The weighting mechanism proposed in this invention integrates semantic similarity and time-sensitive constraints. (First term) It quantifies the semantic similarity between statements, allowing semantically related pairs to be assigned higher importance. Meanwhile, the exponential decay term... It strengthens interactions between temporally adjacent data and penalizes statements that are temporally distant, thereby encouraging the model to capture local patterns that are less sensitive to missing data.

[0086] Then, this invention integrates cross-modal and same-modal edge weights into the adjacency matrix. Among them This represents the number of nodes in the graph. Each element of the matrix can be represented as:

[0087] (7)

[0088] Image denoising:

[0089] After the graph is constructed, this invention performs graph learning. Although various strategies are employed during graph construction to mitigate the impact of missing modalities, the constructed graph may still suffer from structural noise. This noise may originate from cross-modal semantic inconsistencies, insufficient contextual information, or other factors caused by incomplete input. This noise may propagate through the graph and affect node semantic learning and final performance. To address this issue, this invention proposes a graph learning module that first performs neighborhood aggregation to propagate multimodal and temporal information across the graph, then applies structural and semantic denoising to optimize noisy connections, and further refines node representations by suppressing feature noise in the optimized graph.

[0090] Neighborhood aggregation: This invention constructs a graph... Graph convolution is performed to propagate multimodal and temporal information. Mathematically,

[0091] (8)

[0092] in, It is the first The input of the layer. It is a learnable weight matrix. It is an adjacency matrix that encodes edge weights. It is the node degree matrix. This invention stacks graph convolutions. layer.

[0093] Structural Denoising: Although the edge construction strategy of this invention introduces information-rich inter-modal and intra-modal connections, redundant edges may still exist due to semantic inconsistencies or weak temporal relationships. Naive graph message passing methods, as shown in Equation 8, are effective in aggregating local information, but they may also propagate structural noise into node semantics, prompting this invention to propose a structural and semantic denoising strategy. The core idea of ​​this invention here is to evaluate edge effectiveness based on the original input features (before graph convolution) and optimize the graph structure by removing potentially noisy edges. This denoised graph is not used to influence GCN training but serves as the basis for the final semantic denoising layer.

[0094] This invention describes structural denoising as a random diffusion process on the adjacency matrix, such as... Figure 2 As shown. Specifically, let This represents a generic node of any type in the graph (before any graph convolution is performed). From The derived initial binary adjacency matrix. It is an ideal binary adjacency matrix that represents connection patterns without structured noise. The aim is to approximate this matrix. This invention first approaches it step by step... Injecting isotropic Gaussian noise to simulate the forward failure process:

[0095] (9)

[0096] in It is the diffusion time The noisy adjacency matrix at that location, Control Time Noise level at the location. Noise dispatching. Defined as:

[0097] (10)

[0098] in and It is a hyperparameter controlling the noise injection intensity. This diffusion process follows... Approaching 1, the system systematically disrupts the graph structure with increasing noise levels.

[0099] To reverse this diffusion process, this invention trains a neural network based on and diffusion time step To predict injected noise The network is characterized by node features. , and As input, and outputting an edge validity score, the node is estimated. and The probability of the existence of an edge between them. The training objective is:

[0100] (11)

[0101] here It is injected into the node and The noise components on the edges between them. During the inference process, this invention obtains the denoised adjacency matrix by thresholding the predicted edge probabilities. :

[0102] (12)

[0103] in It is the sigmoid function. It is an indicator function. Represents the nodes after structural denoising and The absence or presence of edges between them. This produces a sparse matrix that selectively preserves edges rich in structural information while filtering out noisy edges. The optimized graph notation is as follows. ,in It is the optimized edge set.

[0104] Semantic Denoising: Next, this invention proposes a graph diffusion-based semantic denoising layer to further refine node representations. Unlike structural denoising that optimizes graph topology, this layer iteratively smooths node features through graph propagation while adaptively preserving key signals through residual weighting, thereby implicitly suppressing semantic noise. Specifically, the graph convolutional layer... The output is passed through a graph diffusion layer to suppress noise in the node feature space. Let The semantics of nodes are gradually refined during the graph diffusion process as follows:

[0105] (13)

[0106] in From The normalized graph Laplacian matrix of the generated encoding structure connections. Control the diffusion rate. It is a diffusion step The feature matrix at the given location. This process achieves adaptive feature optimization by iteratively propagating neighborhood information while preserving the original node features, thereby enhancing robustness against noise or inconsistent semantic relevance.

[0107] go through After the diffusion step, the image denoising output is obtained. Statement The learned features of modality-specific and modality-invariant nodes are represented as follows: .

[0108] Reconstruction and Classification:

[0109] For each statement, this invention obtains the final representation by fusing the graph output and the initial mode-invariant information:

[0110] (14)

[0111] here This is a modality-invariant feature derived from Equation 2. This invention observes that summation-based fusion performs better than splicing-based fusion.

[0112] Then, a modality-specific reconstruction module is used to reconstruct the complete data. This invention employs a linear transformation to map the final features to the input space:

[0113] (15)

[0114] here It is the complete data of the reconstruction. These are learnable weights. It is modal The feature dimensions.

[0115] Finally, the present invention will Input a softmax layer to produce the final predicted class distribution. :

[0116] (16)

[0117] in It is an estimated probability. It represents the number of tags. These are learnable weights.

[0118] Joint optimization

[0119] The optimization objective consists of three parts: classification loss. Reconstruction loss and diffusion loss This invention follows previous work by using cross-entropy loss as the classification loss:

[0120] (17)

[0121] in It's a real label.

[0122] In addition, the reconstruction loss between the original features and the reconstructed features is calculated at the missing locations:

[0123] (18)

[0124] in Identify each statement Available modes, It is a mask indicating the missing locations. This is combined with the training objective of the structural denoising process. The overall optimization objective of the method of this invention is:

[0125] (19)

[0126] The proposed model is evaluated on three multimodal datasets: IEMOCAP, CMU-MOSI, and CMU-MOSEI. The performance of the proposed method is validated at different missing rates, defined as follows: .here, It is the first The number of available modalities for each statement. It is the number of statements. It refers to the number of modes. This invention achieves this by setting... and This ensures that at least one modality is available for each sample. For a total of three modalities, The range is from 0.0 to 0.7, where 0.7 is approximately equal to... This invention maintains the same missing rate during training, validation, and testing. The invention selects a weighted average F1 score as the evaluation metric. The model is trained using the Adam optimizer with a learning rate of 0.001. For all datasets, this invention tests the diffusion step in the range {3, 5, 10, 20}. , convolutional layer Within the range of 1 to 4. Dropout rate set to 0.5, ζ set to 0.5. Set to 0.1, Set to 20, Set it to 0.1.

[0127] The results of comparing the model proposed in this invention with various baseline methods are shown in Tables 1 and 2 (for random missing values) and Table 3 (for fixed missing values). The present invention yields the following observations:

[0128] Table 1. Performance comparison on IEMOCAP (quadratic classification) and IEMOCAP (six-class classification) at different missing data rates. The metric is the weighted F1 score (%). Best performance is indicated in bold.

[0129]

[0130] Table 2 compares the performance on the CMU-MOSI and CMU-MOSEI datasets under different missing data rates. The metric is weighted F1 score / accuracy score (%). Best performance is indicated in bold.

[0131]

[0132] Table 3 shows the results for six possible modality missing scenarios. For example, “{a}” indicates that the audio modality is available, while the video and text modalities are missing. “AVG.” indicates the average performance across the six possible scenarios. † indicates a scenario where F1 score or accuracy is not reported. Bold indicates the best results.

[0133]

[0134] As shown in Tables 1 and 2, the proposed method consistently achieves state-of-the-art performance across all benchmark datasets and missing rate settings. On IEMOCAP (quad-class), IEMOCAP (six-class), CMU-MOSI, and CMU-MOSEI, WAF1 scores are improved by 1.71%, 2.76%, 1.0%, and 4.4%, respectively. These results demonstrate the effectiveness of the proposed method under incomplete modal inputs.

[0135] In addition to its strong overall performance, the method of this invention also exhibits good robustness with increasing missing rates, as shown in Tables 1 and 2. While all methods show performance degradation with increasing missing rates, the model of this invention shows a slower performance degradation. For example, on IEMOCAP (six-class classification), WAF1 decreased by 4.64% from a missing rate of 0.0 to 0.7, while the baseline method decreased by 5.18% to 33.58%. This trend is even more pronounced on CMU-MOSEI, where the WAF1 of this invention decreased by 6.3%, while the baseline method decreased by 10.5% to 14.7%. This invention attributes this robustness to the ability to more effectively reveal the underlying relationships between incomplete inputs through the denoising process.

[0136] The method of this invention is based on complete multimodal data (i.e. The invention also demonstrates competitive performance on four datasets. Specifically, it achieves WAF1 improvements of 1.72%, 2.16%, 1.1%, and 2.0%, respectively. This indicates that the model is not only resilient to perturbations caused by missing data but also effective in capturing multimodal temporal information from complete inputs.

[0137] Table 3 further examines the performance under six fixed missing modality settings. It can be seen that the model proposed in this invention outperforms previous methods in most unimodal and bimodal scenarios, while achieving a new state-of-the-art level on average across all datasets. This validates the powerful ability of the proposed method to utilize unimodal cues and integrate complementary information between modalities, highlighting its flexibility and adaptability under various incomplete input scenarios.

[0138] Ablation studies:

[0139] To quantitatively assess the impact of each component of HGDNet, this invention conducted an ablation study on IEMOCAP, and the results are presented in Table 4.

[0140] Table 4. Ablation experiments on the IEMOCAP (four-class classification) and IEMOCAP (six-class classification) datasets. This invention reports the weighted F1 score (%), with best performance indicated in bold.

[0141]

[0142] 1) Effect of Modality-Invariant Nodes: This invention first removes modality-invariant nodes from the graph model to study its effect. It can be seen that the results show a significant decrease across all missing rates in both datasets, with average WAF1 decreasing by 3.50% (four-class classification) and 1.09% (six-class classification), respectively. This indicates that modality-invariant representation plays a crucial role in bridging cross-modal information and promoting effective fusion.

[0143] 2) Effect of Structural Denoising: The present invention then removes the structural denoising process during graph learning. As can be observed from Table 4, a significant decrease in results is seen in most cases, and this component appears to be more valuable for fine-grained six-class classification settings. This indicates that denoising noisy structural links in the graph contributes to more stable performance.

[0144] 3) Effect of Semantic Denoising: This invention further examines the case where both structural and semantic denoising modules are removed simultaneously. When both modules are removed, the performance degradation is more significant than removing structural denoising alone. This indicates that semantic denoising complements structural denoising by enhancing feature-level consistency and mitigating semantic noise.

[0145] 4) Effect of Edge Weight Initialization: This invention investigates the impact of the specially designed edge weights (Equations 4 and 6) on graph construction. The invention presents results obtained from variations of randomly initialized edge weights. This invention observes a performance decrease across all missing rates, highlighting the effectiveness of the edge weight design in this invention.

Claims

1. A robust multimodal dialogue understanding method based on heterogeneous graph denoising, characterized in that, Includes the following steps: S1. Obtain multimodal training data, specifically: Get the dialogue sequence ,in It is the total number of statements. It is the first Statements, express exist The true labels for each category; for each statement Extract a set of multimodal features including language Visual Harmony Modality , , These represent the dimensions of linguistic, visual, and acoustic features, respectively. By randomly removing certain modalities to simulate a modality-deficient scenario, while ensuring that each sample retains at least one modality, multimodal features of each sentence are obtained. Represented as , Missing indicator Representing modes exist, This indicates its absence, and ,in ; S2. Construct a multimodal dialogue understanding network, including multimodal graph construction, graph denoising, reconstruction, and classification; The process of constructing the multimodal graph is as follows: using the obtained multimodal features Construct an undirected heterogeneous graph heterogeneous node set , For a modality-specific set of nodes, Given a modality-invariant node set, unimodal information is encoded through a fully connected network to generate modality-specific features. : , in, , These are learnable weights. For the feature dimensions of a graph network, use Initialize modality-specific node embedding This makes each statement have One node; Multimodal features spliced ​​as Then the spliced ​​feature sequence Inputting the bidirectional long short-term memory (BiLSTM) network to encode context dependencies yields mode-invariant nodes. : , in, , These are trainable parameters, used Initialize modally invariant node embedding Each statement generates a modally invariant node; Edge set Includes cross-modal edge sets same modal edge set Cross-modal edge sets The method of establishing is to set each node Other available modalities that can be connected to the same statement: , in Simultaneously, cross-modal edge weights are adaptively assigned based on the semantic similarity between node features: , in, , Represents the cosine similarity function. It is the sigmoid activation function; Same-modal edge set The method of establishment is in each mode Inside, each node Connect to all other available statements in the same dialog: , Define the weights of edges with the same modality as: , in, Represents a node and Edge weights between them Representing nodes respectively Time and location, It is a hyperparameter; Integrate cross-modal and same-modal edge weights into the adjacency matrix. Among them This represents the number of nodes in the graph, and each element of the matrix is ​​represented as: ; The image denoising process is as follows: In the constructed graph Execution graph convolution is used to propagate multimodal and temporal information: , in, It is the first Layer input, It is a learnable weight matrix. It is the node degree matrix, which is the result of convolutional stacking of the graph. Layers enable neighborhood aggregation; For the constructed graph Structural denoising is performed, specifically by denoising the edge set. Optimize: Define Represents a generic node of any type in the graph. From The derived initial binary adjacency matrix, It is an ideal binary adjacency matrix, representing a connection pattern without structured noise; The aim is to approximate the matrix, first by gradually moving towards Injecting isotropic Gaussian noise to simulate the forward failure process: , in, It is the diffusion time The noisy adjacency matrix at that location, Control Time Noise level at: , in, and It is a hyperparameter that controls the intensity of noise injection; with Approaching 1, the system systematically destroys the graph structure with increasing noise levels. To reverse this process, a neural network is trained based on... and diffusion time To predict injected noise The network is based on nodes , and As input, and outputting an edge validity score, the node is estimated. , The probability of the existence of an edge between them, and the training objective is: , in, It is diffusion loss. It is injected into the node , The noise components on the edges between them; during inference, the denoised adjacency matrix is ​​obtained by thresholding the predicted edge probabilities. : , in, It is the sigmoid function. It is an indicator function. Represents the nodes after structural denoising and The absence or presence of edges between them; using the denoised adjacency matrix Optimized edge set Thus, the optimized notation is obtained as follows: ; Semantic denoising is performed based on the results of neighborhood aggregation and structural denoising. Specifically, the graph convolutional layer... The output is passed through a graph diffusion layer to suppress noise in the node feature space, making The semantics of nodes are gradually refined during the graph diffusion process as follows: , in, From The normalized graph Laplacian matrix of the generated encoding structure connections. Controlling the diffusion rate, It is a diffusion step The characteristic matrix at the location; after After the diffusion step, a denoised image output is obtained. Statement The learned features of modality-specific and modality-invariant nodes are represented as follows: ; The reconstruction and classification process is as follows: For each statement, the final representation is obtained by fusing the graph output and the initial modality-invariant information. , A modality-specific reconstruction module is used to reconstruct the complete data, and a linear transformation is employed to map the final features to the input space: , in, It is the complete data of the reconstruction. These are learnable weights. It is modal Feature dimensions, The final representation of the above statement is as follows: For learnable bias; Will Input a softmax layer to produce the final predicted class distribution. : , in, It is an estimated probability. It is the number of tags. These are learnable weights; S3. Train the constructed multimodal dialogue understanding network using multimodal training data, with the optimization objective being: , , , in, It is classification loss. It is the reconstruction loss. It's a real label. Identify each statement Available modes, It is a mask indicating the missing location; after training under the set conditions, a trained multimodal dialogue understanding network is obtained. S4. Obtain the dialogue sequence to be understood, extract the multimodal features of each sentence, input them into the trained multimodal dialogue understanding network, and obtain the prediction results based on the network output.

Citation Information

Patent Citations

  • Multi-modal dialogue emotion recognition method based on plural neural networks

    CN118568664A

  • Live broadcast room content identification and intelligent distribution method and system based on multi-modal fusion

    CN119377895A