Multi-channel dynamic hypergraph sentiment analysis method and analysis network fusing time sequence consistency

By employing multi-channel feature extraction, local temporal context fusion, and dynamic hypergraph construction, this study addresses the issues of modal feature homogenization, insufficient temporal modeling, and information leakage in multimodal sentiment analysis, thereby improving the accuracy and stability of sentiment recognition.

CN120929912APending Publication Date: 2025-11-11HARBIN INST OF TECH
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202511033944.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-25
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

Existing multimodal sentiment analysis technologies suffer from limitations such as homogenization of features within each modality, insufficient temporal modeling capabilities, limitations in cross-modal interaction modeling, and the risk of information leakage from static graph structures, which restrict the accuracy and generalization ability of sentiment recognition.

Method used

We employ a multi-channel feature extraction, local temporal context fusion, dynamic hypergraph construction, and multi-level supervision approach. We extract multi-channel features through a heterogeneous pre-trained model, combine it with Transformer to capture long-range dependencies, dynamically construct a hypergraph model, and use spectral-spatial hybrid convolution. We also introduce early MLP branch optimization to finally generate sentiment prediction results.

Benefits of technology

It improves the model's robustness and generalization ability to emotional information, and significantly enhances the accuracy and stability of emotion recognition, especially in terms of performance on the CMU-MOSI and CH-SIMS datasets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120929912A_ABST
    Figure CN120929912A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-channel dynamic hypergraph sentiment analysis method and a multi-channel dynamic hypergraph sentiment analysis network fusing time sequence consistency, belongs to the field of artificial intelligence and multi-modal sentiment calculation, and aims to solve the problems existing in the existing sentiment analysis technology. The method comprises the following steps: S1, a multi-channel feature extraction step: extracting multi-channel features of a text mode and an audio mode through a heterogeneous pre-training model; s2, a local time sequence context fusion step based on a video number: fusing short-term emotional fluctuation based on a local context mechanism of the video number, and capturing long-range dependence across time dimensions through Transform; s3, a single-modal-multi-modal hypergraph collaborative prediction step: dynamically constructing a single-modal hypergraph and a multi-modal hypergraph in a training batch, and modeling a high-order relationship by adopting spectral domain-spatial domain hybrid convolution; and S4, a multi-level multi-branch supervision step: outputting a final emotion prediction result through joint optimization of an early MLP branch and a late hypergraph branch.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of artificial intelligence and multimodal emotion computing, specifically involving a deep learning network architecture based on multi-channel dynamic hypergraphs and local temporal consistency, used to identify and analyze human emotional states from multimodal data such as text, speech, and vision. This technology can be applied to fields such as intelligent human-computer interaction, mental health assessment, and social media emotion monitoring. Background Technology

[0002] Multimodal Sentiment Analysis (MSA) aims to identify and understand human emotions by fusing information from different modalities, including text, speech, and vision. Compared to unimodal methods, multimodal techniques can leverage the complementarity between different information sources, significantly improving the accuracy of emotion recognition. For example, when text content is neutral, tone of voice or facial expressions may reveal the true emotional state. In recent years, with the development of deep learning technology, neural network-based multimodal sentiment analysis methods have become mainstream research, attracting increasing attention due to their wide application in intelligent human-computer interaction, social media emotion monitoring, and mental health assessment. Since human emotional expression naturally involves multiple modalities such as language content, speech tone, and facial expressions, effectively fusing these multi-source signals has become a key issue in achieving robust emotion recognition.

[0003] Existing sentiment analysis techniques can be categorized into three types: dictionary-based methods, machine learning-based methods, and deep learning-based methods. With the continuous development of artificial intelligence, deep learning-based sentiment analysis methods have gradually become a research hotspot. However, existing deep learning-based MSA methods still face many challenges:

[0004] (1) Problem of feature simplification within modalities: Most current methods use a single pre-trained model (such as using only BERT to extract text features) for feature extraction, which leads to the loss of semantic information at different levels within a modality, making it difficult to cover semantic information at different levels within a modality and limiting the expressive power of the model. For example, traditional audio feature extraction only uses a single model such as COVAREP or LibROSA, which makes it difficult to capture both low-order acoustic features (such as fundamental frequency and energy) and high-order semantic features (such as speech sentiment).

[0005] (2) Insufficient temporal modeling capability: Existing methods have significant shortcomings in balancing short-term emotional dynamics and long-term semantic dependencies. Traditional RNN / LSTM structures are prone to the vanishing gradient problem, resulting in limited ability to capture long-term dependencies; while the standard Transformer architecture is difficult to effectively model local emotional fluctuations. For example, in movie review scenarios, brief facial expression changes (such as a momentary smile) may be diluted by the global attention mechanism. It is evident that balancing short-term emotional dynamics and long-term semantic dependencies remains a challenge, and traditional sequence modeling methods struggle to simultaneously capture local emotional fluctuations and global contextual relationships.

[0006] (3) Limitations of cross-modal interaction modeling: Existing cross-modal interaction methods mostly use simple splicing or attention mechanisms, which are difficult to capture complex high-order relationships. Existing methods have limited ability to model cross-modal high-order correlations.

[0007] (4) Information leakage risk of static graph structures: Methods that construct static graph structures based on global data (such as Graph-MFN and MTAG) usually construct static graph structures on global data, which may cause test set information to leak into the training process through graph connections. Furthermore, static graph structures are difficult to adapt to dynamic sentiment expression, thus limiting the generalization ability of the model. Summary of the Invention

[0008] To address the problems existing in current sentiment analysis techniques, this invention provides a multi-channel dynamic hypergraph sentiment analysis method and analysis network (MCDH-Net) that integrates temporal coherence. This method better models short-term sentiment changes and long-term semantic dependencies, enhances the model's adaptability to dynamic emotional expressions, and avoids information leakage from the test set caused by static graphs.

[0009] In a first aspect, the present invention provides a multi-channel dynamic hypergraph sentiment analysis method that integrates temporal consistency, comprising the following steps:

[0010] S1. Multi-channel feature extraction steps: Extract multi-channel features of text and audio modalities respectively through heterogeneous pre-trained models;

[0011] S2. Local Temporal Context Fusion Step Based on Video Number: Short-term sentiment fluctuations are fused using a local context mechanism based on video number, and long-range dependencies across time dimensions are captured through Transformer.

[0012] S3, Single-modal-multimodal hypergraph collaborative prediction steps: Dynamically construct single-modal and multimodal hypergraphs within training batches, and use spectral-spatial hybrid convolution to model higher-order relationships;

[0013] S4. Multi-level, multi-branch supervision steps: The final sentiment prediction result is output through joint optimization of the early MLP branch and the late hypergraph branch.

[0014] Preferably, the multi-channel feature extraction step S1 specifically includes:

[0015] Text modality feature extraction: Text features are extracted using a dual-channel approach combining the BERT and RoBERTa models;

[0016] Audio modality feature extraction: For the CMU-MOSI dataset, the COVAREP toolkit and Data2vec model were used to extract audio features in two channels. For the CH-SIMS dataset, the LibROSA model and HuberT model were used to extract audio features in two channels.

[0017] Preferably, the S2 local temporal context fusion step based on video ID specifically includes:

[0018] S21. Construct multi-channel context-enhanced features based on the local context mechanism of video number: For the current moment feature, fuse the features of the previous two moments of the video in which it is located, and fill in the missing features with the current features;

[0019] S22. Establish global dependencies based on Transformer encoder: First, fuse the context enhancement features of multiple channels in each modality, and then introduce the Transformer model to process the fused multi-channel features and generate multi-modal fused features.

[0020] Preferably, the S3 single-modal-multimodal hypergraph collaborative prediction step specifically includes:

[0021] S31. Single-modal dynamic hypergraph construction: A batch local strategy is adopted, and the hypergraph structure is constructed independently for each batch;

[0022] S32. Hypergraph Hybrid Convolution: Perform spectral domain convolution and spatial domain convolution on unimodal features and multimodal fusion features respectively to generate sentiment prediction results based on unimodal hypergraphs and sentiment prediction results based on multimodal hypergraphs.

[0023] Preferably, the outputs of spectral domain convolution and spatial domain convolution are connected using residual connections.

[0024] Preferably, the S4 multi-level multi-branch supervision step specifically includes:

[0025] S41. Utilize early MLP branches to directly perform regression prediction on Transformer pre-fusion features;

[0026] S42. Introduce a multi-branch loss fusion structure, which integrates the losses of the single-modal hypergraph branch, the multi-modal hypergraph branch, and the MLP branch in a weighted manner to construct the final training objective.

[0027] In a second aspect, the present invention provides the aforementioned multi-channel dynamic hypergraph sentiment analysis network with fused temporal consistency, comprising:

[0028] Multi-channel feature encoding module: used to extract multi-channel features of text and audio modalities respectively through heterogeneous pre-trained models;

[0029] Short-term dynamics and long-term dependency modeling module: used to fuse short-term sentiment fluctuations based on local context mechanism of video number, and capture long-term dependencies across time dimension through Transformer;

[0030] Single- and multimodal hypergraph collaborative prediction module: used to dynamically construct single-modal and multimodal hypergraphs within training batches, and to model higher-order relationships using spectral-spatial hybrid convolution;

[0031] Multi-level, multi-branch supervision module: used to jointly optimize through early MLP branches and late hypergraph branches, and output the final sentiment prediction result.

[0032] The beneficial effects of this invention are as follows: This invention utilizes heterogeneous pre-trained models for multi-channel feature extraction to capture diverse representational information within modalities; simultaneously, it introduces a local temporal context mechanism to model short-term sentiment changes, and inputs context-aware features into a Transformer structure to learn global dependencies across time dimensions; subsequently, it proposes unimodal and multimodal hypergraph convolutional modules to model high-order relational structures, and dynamically constructs the hypergraph in each training batch to prevent information leakage; finally, it introduces an auxiliary MLP regression branch based on early features to preserve fine-grained representational capabilities, and implements a multi-level optimization strategy through joint loss training. It has the following advantages:

[0033] 1. This invention utilizes multi-channel feature extraction to encode the same modality from multiple perspectives, thereby obtaining richer semantic and emotional information. Unlike relying solely on a single model, the multi-channel approach effectively integrates complementary features, enhancing the model's robustness and generalization ability.

[0034] 2. This invention employs a temporal context fusion method, which can capture the short-term temporal dependence of emotions, improve the model's ability to model fine-grained emotional dynamics, and also utilize self-attention mechanisms to capture the long-term and global dependencies of emotional states across channels.

[0035] 3. This invention employs dynamic single- and multi-modal hypergraph sentiment reasoning, which can simultaneously model higher-order associations within a modality and semantic collaboration across modalities, while the batch-based dynamic construction strategy can also strictly avoid information leakage;

[0036] 4. This invention improves the generalization and stability of the model by jointly optimizing the single-modal hypergraph loss, multimodal hypergraph loss and MLP loss.

[0037] This invention achieves significant performance improvements on the CMU-MOSI and CH-SIMS datasets:

[0038] Achieved 54.52% accuracy in the CMU-MOSI seven-class classification task, an improvement of 4.18% over the best baseline (MMML);

[0039] Achieved 50.44% accuracy in the CH-SIMS five-class classification task, an improvement of 1.06% over the best baseline (MMML);

[0040] The MAE index for the regression task decreased by 0.0322-0.4048;

[0041] Training efficiency improved by approximately 23%, while memory consumption decreased by 18.7%. Attached Figure Description

[0042] Figure 1 This is a schematic diagram of the overall architecture of the MCDH-Net multi-channel dynamic hypergraph sentiment analysis network that integrates local temporal consistency, as proposed in this invention.

[0043] Figure 2 This is a diagram of the pre-normalized Transformer structure.

[0044] Figure 3 The process for constructing a single-modal dynamic hypergraph.

[0045] Figure 4 This is a flowchart of the multimodal dynamic hypergraph hybrid convolution process. Detailed Implementation

[0046] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0047] It should be noted that, unless otherwise specified, the embodiments and features described in the present invention can be combined with each other.

[0048] The present invention will be further described below with reference to the accompanying drawings and specific embodiments, but this is not intended to limit the scope of the invention.

[0049] Specific Implementation Method 1: The following is combined with... Figures 1 to 4 This embodiment describes a multi-channel dynamic hypergraph sentiment analysis network that integrates temporal consistency. See [link to previous document]. Figure 1 It includes the following four modules:

[0050] Multi-channel feature encoding module: used to extract multi-channel features of text and audio modalities respectively through heterogeneous pre-trained models;

[0051] Short-term dynamics and long-term dependency modeling module: used to fuse short-term sentiment fluctuations based on local context mechanism of video number, and capture long-term dependencies across time dimension through Transformer;

[0052] Single- and multimodal hypergraph collaborative prediction module: used to dynamically construct single-modal and multimodal hypergraphs within training batches, and to model higher-order relationships using spectral-spatial hybrid convolution;

[0053] Multi-level, multi-branch supervision module: used to jointly optimize through early MLP branches and late hypergraph branches, and output the final sentiment prediction result.

[0054] based on Figure 1 The analysis method for multi-channel dynamic hypergraph sentiment analysis networks that incorporate temporal consistency includes the following steps:

[0055] S1. Multi-channel feature extraction steps: Extract multi-channel features of text and audio modalities respectively through heterogeneous pre-trained models;

[0056] S2. Local Temporal Context Fusion Step Based on Video Number: Short-term sentiment fluctuations are fused using a local context mechanism based on video number, and long-range dependencies across time dimensions are captured through Transformer.

[0057] S3, Single-modal-multimodal hypergraph collaborative prediction steps: Dynamically construct single-modal and multimodal hypergraphs within training batches, and use spectral-spatial hybrid convolution to model higher-order relationships;

[0058] S4. Multi-level, multi-branch supervision steps: The final sentiment prediction result is output through joint optimization of the early MLP branch and the late hypergraph branch.

[0059] In step S1, multiple heterogeneous models are introduced for feature extraction for each modality to form multi-channel modal features; the input is the original data of the two modalities.

[0060] The multi-channel feature extraction step specifically includes:

[0061] S11. Text Modal Feature Extraction: Text features are extracted using a dual-channel approach combining the BERT and RoBERTa models.

[0062] To enhance the semantic richness and robustness of the text representation, BERT and RoBERTa are first used to extract the initial text features F. t_1 With F t_2 Although both are based on the Transformer architecture, their different training strategies and corpus sizes make them complementary in semantic representation capabilities. The specific implementation process is as follows:

[0063] F t_1 =BERT(U t ;θ t_1 )

[0064] F t_2 =RoBERTa(U t ;θ t_2 )

[0065] U t Represents the original text data, θ t_1 and θ t_2 These are the hyperparameters of the pre-trained BERT model and the RoBERTa model, respectively.

[0066] S12. Audio Modal Feature Extraction: For the CMU-MOSI dataset, the COVAREP toolkit and Data2vec model are used to extract audio features in two channels. For the CH-SIMS dataset, the LibROSA model and HuberT model are used to extract audio features in two channels.

[0067] In terms of audio modal feature extraction, this invention uses the COVAREP toolkit and the Data2vec model to process the original audio for the CMU-MOSI dataset. For the CH-SIMS dataset, the LibROSA and HuberT models are used to extract audio features, respectively. This strategy enhances the model's generalization ability when faced with different speakers and diverse emotional expressions. The specific process for extracting audio features on the CMU-MOSI dataset is as follows:

[0068]

[0069] U a Represents the raw audio data. and These are the hyperparameters of the COVAREP model and the pre-trained Data2vec model, respectively. Similarly, the process for extracting audio features from the CH-SIMS dataset is as follows:

[0070]

[0071] in and These are the hyperparameters of the pre-trained LibROSA model and the HuberT model, respectively. This method combines the advantages of traditional handcrafted acoustic features with deep contextual representations, achieving multi-level modeling of audio information.

[0072] In step S2, a mechanism for constructing local context based on video IDs is proposed. Specifically, for each feature at a given moment, the two moments preceding it under its corresponding video ID are selected as context features. If no valid context exists, the current feature itself is used to complete the feature. Then, the current feature and the context features are weighted and fused column-wise to obtain new moment features. Finally, the new features from each channel are concatenated and input into the Transformer model for processing.

[0073] The local temporal context fusion steps based on video IDs specifically include:

[0074] S21. Construct multi-channel context-enhanced features based on the local context mechanism of video number: For the current moment feature, fuse the features of the previous two moments of the video in which it is located, and fill in the missing features with the current features;

[0075] In this step, to model the evolution of sentiment over time and improve the model's contextual modeling ability, a feature enhancement method based on local temporal consistency is proposed. The specific process is as follows:

[0076] (1) For any given time step, its features not only include the representation of the current moment, but also incorporate the features of the previous two time steps in the same video as contextual information. Assume... Let represent the initial feature vector of mode m, m∈{a,t} at time τ, where a corresponds to the audio mode and t corresponds to the text mode. This feature belongs to the video sequence v. The current feature... By weighting and concatenating the feature vectors from the previous two time steps along the column direction, we can obtain the local context-enhanced features.

[0077]

[0078] Where || represents the feature concatenation operation. and These represent the features of the two time steps that have the same video ID as the current sample, α and β. 0_m α 1_m and α 2_mThese are learnable weight coefficients used to control the contribution of features from different time steps in the fusion process. This design enables the model to learn temporal patterns of emotional states from historical emotional information, while supplementing any emotional ambiguity or missing information in the current time step, thereby generating more stable, coherent, and robust feature representations.

[0079] (2) If the context features are affected by crossing the video boundary or If missing, the current time feature is used. To maintain consistency of contextual information:

[0080]

[0081] Where k = 1, 2, and video(τ-k) represents the original video ID corresponding to time (τ-k). This strategy avoids introducing cross-video data interference, ensuring the consistency of feature semantics and the stability of contextual modeling. The above steps effectively integrate local temporal information and contextual features, enabling dynamic tracking of sentiment trends over time.

[0082] S22. Establish global dependencies based on Transformer encoder: First, fuse the context enhancement features of multiple channels in each modality, and then introduce the Transformer model to process the fused multi-channel features and generate multi-modal fused features.

[0083] Detailed calculation process Figure 2 Input F in m F', m∈{a,t} a The process of obtaining ′ and F t Consistent, output The acquisition process and Consistent. The specific process, taking the text modalities in the CMU-MOSI dataset as an example, is as follows:

[0084] (1) After obtaining the contextual fusion features of each time step, the information from each channel is further integrated to capture the global sentiment dependency across time steps:

[0085]

[0086] Where Mean(·) represents the mean operation, and F t ′∈R 1 ×(2304+2048) .

[0087] (2) The fused single-modal features are then input into the Transformer model for processing, as defined below:

[0088]

[0089] in This represents multi-channel context-enhanced features. The Transformer model can capture long-term dependencies across channels and time steps, thereby enabling the modeling of deep semantic relationships in the data.

[0090] In step S3, we construct hybrid convolutional networks for single-modal and multi-modal hypergraphs in the feature space encoded by the Transformer. Unlike existing global hypergraph construction methods, this invention adopts a batch dynamic construction strategy, dynamically generating hypergraphs based only on the current training batch data, strictly avoiding information leakage.

[0091] The specific steps of single-modal-multimodal hypergraph collaborative prediction include:

[0092] S31. Single-modal dynamic hypergraph construction: A batch local strategy is adopted, and the hypergraph structure is constructed independently for each batch;

[0093] S32. Hypergraph Hybrid Convolution: Perform spectral domain convolution and spatial domain convolution on unimodal features and multimodal fusion features respectively to generate sentiment prediction results based on unimodal hypergraphs and sentiment prediction results based on multimodal hypergraphs.

[0094] At the feature level, more complex non-Euclidean structural relationships still exist between different samples. We first construct a dynamic hypergraph based on single-modal features and introduce a spectral-spatial hypergraph hybrid convolution strategy to model single-modal features, thereby mining potential high-order structural dependencies within the modality and further improving the discriminative ability of sentiment features.

[0095] The process of obtaining the sentiment prediction results of a single-modal hypergraph includes:

[0096] (1) For the single-modal feature representation generated by Transformer We construct local hypergraphs in batches during training. For the feature vectors in each batch... Each sample is then considered as a central vertex. See also Figure 3 To distinguish between different batches, the feature vector of batch n is expressed as: Then, using a K-nearest neighbor strategy based on Euclidean distance, it is connected to the K-1 most similar samples in the feature space to construct a hyperedge for the central vertex. The similarity measurement rule is defined as follows:

[0097]

[0098] Where l represents the single-modal feature dimension. These are the single-modal features of the i-th and j-th samples, respectively. Then, the center node is selected. The K-1 most similar nodes are used to jointly construct the hyperedge.

[0099]

[0100] in The hypergraph G represents the mode m. m The j-th superedge in.

[0101] To distinguish between different batches, the batch n hyperedge is expressed as:

[0102] (2) Once all hyperedges are constructed, the connection relationships between vertices and hyperedges can be determined by the correlation matrix A. m To distinguish between different batches, the batch n correlation matrix is ​​represented as A. mn In matrix A m In this model, each column corresponds to a hyperedge. Non-zero elements in the column represent the vertex features contained within that hyperedge, and their numerical values ​​reflect the feature similarity between the vertex and the central vertex of the hyperedge. This invention uses a smooth inverse distance weighted method to calculate similarity, defined as follows:

[0103]

[0104] in Corresponding matrix A m The value of the element in the i-th row and j-th column.

[0105] Single-modal hypergraphs model the joint relationships between multi-channel features (such as the interaction between pitch, energy, and semantic cues in speech) through hyperedge structures, which helps to capture complex emotional patterns and improve emotion recognition performance. Figure 3 The process of constructing a single-modal dynamic hypergraph is demonstrated.

[0106] (3) After constructing the single-modal hypergraph, the corresponding adjacency matrix A is... m With multi-channel context-enhanced feature matrix A common input multi-layer hypergraph hybrid convolution module updates the vertex representation. Each layer simultaneously applies spectral and spatial convolutions to the single-modal features, and the results are summed as the input to the next layer. For the feature matrix... Its spectral domain hypergraph convolution can be represented as:

[0107]

[0108] Where F m_spec This indicates that the hypergraph convolution pairs are passed through the spectral domain. Enhanced features, D m_e and Dm_v Let A be the hyperedge and vertex normalization matrix of mode m, respectively. m and W m_e Let θ represent the incidence matrix and hyperedge weight matrix of mode m, respectively. m These are the learnable parameters of the convolutional layer. And the modality m... The corresponding spatial hypergraph convolution can be represented as:

[0109]

[0110] Where F m_spa This indicates that the spatial hypergraph convolution is performed from... The resulting enhanced features, θ m_e and θ m_v These are learnable convolutional layer parameters. This process enables each vertex to dynamically aggregate information from associated vertices in the original hypergraph, thereby capturing complex intramodal high-order correlations and promoting the fusion of local features.

[0111] (4) Finally, the spectral domain convolution output F m_spec Spatial convolution output F m_spa The summations are then fed into a linear transformation layer to generate sentiment prediction results based on a single-modal hypergraph.

[0112]

[0113] Where w m and b m These represent the weights and biases of the linear layer, respectively. This represents the single-modal sentiment prediction value for the nth batch of samples.

[0114] Furthermore, to fully explore the complementary information among the various modalities, we concatenate the single-modal features extracted by the Transformer to construct a corresponding multimodal hypergraph. Each hyperedge connects multiple semantically related samples, thus enabling the modeling of higher-order relationships between modalities. The process of obtaining the sentiment prediction results from the multimodal hypergraph includes:

[0115] (1) First, the individual modal features are concatenated to obtain the fused feature F, which is used to construct the multimodal hypergraph. The specific definition is as follows:

[0116]

[0117] (2) The construction method of multimodal hypergraphs is similar to that of unimodal hypergraphs, the only difference being that F is replaced by F. The remaining steps are the same as the input, and will not be repeated here. We only provide the adjacency matrix A of the multimodal hypergraph. f Definition:

[0118]

[0119] in Let be the element in the i-th row and j-th column of the adjacency matrix. Represents a multimodal hypergraph G f The j-th superedge in.

[0120] (3) After the multimodal hypergraph is constructed, the present invention performs hypergraph hybrid convolution processing on the fused feature F, in the following specific form:

[0121]

[0122] Where F f_spec With F f_spa D represents the enhanced feature of the fused feature F after convolution of the spectral and spatial hypergraphs, respectively; f_e and D f_v These are the hyperedge and vertex normalization matrices corresponding to the fused features, respectively; A f and W f_e These are the hypergraph's incidence matrix and hyperedge weight matrix, respectively; θ f θ is a learnable parameter for hypergraph convolution in the spectral domain. f_e and θ f_v These are the learnable parameters for spatial hypergraph convolution.

[0123] (4) Finally, the outputs of the spectral domain and spatial domain convolutions are weighted and fused, and the final multimodal sentiment prediction result is obtained through linear transformation.

[0124]

[0125] Where w f and b f These represent the weights and biases of the linear layer, respectively. This represents the multimodal prediction score of the nth sample.

[0126] This invention employs a single-modal-multimodal hypergraph modeling and inference strategy, which preserves the independence of each modality's features and achieves deep cross-modal information fusion through multimodal hypergraph convolution, significantly improving the model's representational ability and prediction performance.

[0127] In step S4, while utilizing the hypergraph to mine higher-order nonlinear relationships in the data, an MLP prediction path based on Transformer pre-features is introduced to retain fine-grained information from early fusion, forming a collaborative paradigm of "early fusion + late fusion" to achieve multi-level information complementarity. A multi-level, multi-branch supervision network is constructed to improve the robustness and generalization of the model.

[0128] The multi-level, multi-branch supervision process specifically includes:

[0129] S41. Utilize early MLP branches to directly perform regression prediction on Transformer pre-fusion features;

[0130] To fully explore the expressive power of multi-level fusion features, we will use context-aware features from different modalities and channels. The features are concatenated directly before being processed by the Transformer to obtain the fused feature F. MLP The fused features are then input into a multilayer perceptron (MLP) to generate corresponding multimodal sentiment prediction results.

[0131]

[0132] Where w1, w2, w3 and b1, b2, b3 are the weights and biases of the MLP network, respectively.

[0133] S42. Introduce a multi-branch loss fusion structure, which integrates the losses of the single-modal hypergraph branch, the multi-modal hypergraph branch, and the MLP branch in a weighted manner to construct the final training objective.

[0134] In summary, the MLP-based early fusion branch and the hypergraph-based late fusion branch are functionally complementary: the former focuses on capturing simple statistical correlations between modalities, such as the linear relationship between volume and emotion intensity; the latter is dedicated to modeling higher-order nonlinear interactions within and between modalities. This design allows the model to utilize both the fine-grained features obtained from early fusion and the complex dependency structures captured by the hypergraph module, ultimately achieving multi-level information complementarity.

[0135] The training method of the multi-channel dynamic hypergraph sentiment analysis network with temporal consistency described in this invention is as follows:

[0136] Step T1: For the original text data U t The initial text features F were obtained by processing them in batches using BERT and RoBERTa respectively. t_1 With F t_2 Among them, on the CMU-MOSI dataset, F t_1 ∈R 50×768 F t_2 ∈R 1×1024 On the CH-SIMS dataset, F t_1 ∈R 39×768 F t_2 ∈R 1×768 For the raw audio data U a On the CMU-MOSI dataset, the COVAREP toolkit and Data2vec model were used to process the original audio to obtain initial audio features. On the CH-SIMS dataset, audio features were extracted using the LibROSA and HuBERT models, respectively.

[0137] Step T2: To model the emotion dependency relationship within a local time window, the features at each moment are analyzed. Select the first two time-step samples belonging to the same video ID v and Used as contextual information; if the features from the previous two time steps are unavailable, then the features from the current time step are used. The features are then supplemented. Subsequently, the current time-to-time and context features are fused column-by-column, and then integrated using a weighted strategy to obtain a locally context-enhanced feature representation. Finally, the local context enhancement features of each channel are applied. and The concatenation operation is then fed into the Transformer module to obtain multi-channel context-enhanced features.

[0138] Step T3: For multi-channel contextual enhancement features During training, local hypergraphs are constructed in batches. For each batch of feature vectors... Constructing hyperedges based on the K-nearest neighbor algorithm This leads to the adjacency matrix A of the single-modal hypergraph. m In this invention, batch_size = 8 and K = 3. Then, A... m and Together, they serve as input to the spectral domain hypergraph convolutional layer C1 and the spatial domain hypergraph convolutional layer C2. Finally, the spectral domain convolution output F is... m_spec Spatial convolution output F m_spa The summations are then fed into a linear transformation layer to generate sentiment prediction results based on a single-modal hypergraph.

[0139] Step T4: Similarly, the single-modal features obtained in step T2 are... and The features are concatenated to obtain the fused feature F. For each batch of feature vectors... Constructing hyperedges based on the K-nearest neighbor algorithm This leads to the multimodal hypergraph adjacency matrix A. f Subsequently, A f Together with F, it serves as the input to the spectral domain hypergraph convolutional layer C1 and the spatial domain hypergraph convolutional layer C2. Finally, the spectral domain convolution output F is... f_spec Spatial convolution output F f_spa The sums are then fed into a linear transformation layer to generate sentiment prediction results based on a multimodal hypergraph.

[0140] Step T5: We will use context-aware features from different modalities and channels. The features are concatenated directly before being processed by the Transformer to obtain the fused feature F. MLP Subsequently, the fused features are passed through three linear layers to generate the corresponding multimodal sentiment prediction results.

[0141] Step T6: Finally, a multi-branch loss fusion structure is introduced to reduce the single-modal hypergraph branch loss. a and Loss t Multimodal hypergraph branch loss f and MLP branch loss MLP The data are then integrated in a weighted manner to construct the final training objective. The calculation formula is as follows:

[0142]

[0143] Where β1, β2, β3, and β4 represent the weighting coefficients of the loss terms for each branch. Furthermore, a smoothed L1 loss is used as the objective function during training. Taking the multimodal hypergraph loss as an example, it is specifically expressed as follows:

[0144]

[0145] in, Let F represent the true sentiment label of the nth input sample. Based on this, the training loss of the multimodal fusion feature F can be expressed as:

[0146]

[0147] Where N t This indicates the number of training samples.

[0148] To further verify the effectiveness of the proposed algorithm, this embodiment tested the model proposed in this application on the publicly available multimodal sentiment analysis datasets CMU-MOSI and CH-SIMS. The CMU-MOSI dataset is a benchmark English corpus in the field of multimodal sentiment analysis, consisting of 2199 speech-level video clips extracted from 93 YouTube movie review videos posted by 89 narrators. Each clip contains synchronized visual, audio, and textual modal information, and its sentiment intensity score is manually labeled, ranging from -3 (strongly negative) to +3 (strongly positive), comprehensively reflecting the polarity and intensity of sentiment. This dataset is divided into 1284 training samples, 229 validation samples, and 686 test samples.

[0149] The CH-SIMS dataset is a benchmark dataset in the field of Chinese multimodal sentiment analysis. It consists of 2281 video clips from various sources, including movies and TV series, and is manually annotated with sentiment information. This dataset contains rich real-world variability, such as spontaneous emoticons and diverse head poses, with sentiment labels ranging from -1 (strongly negative) to +1 (strongly positive). CH-SIMS includes text, audio, and visual modalities in the Chinese context and provides unimodal and multimodal sentiment labels for each speech, enabling researchers to conduct fine-grained sentiment analysis at both modality-specific and fusion levels.

[0150] The network model in this embodiment was built using the PyTorch framework. During training, the batch size was set to 8, and the initial learning rate was 5e-6. The model was optimized using the AdamW optimizer, with a smooth L1 loss function as the training objective. To ensure fairness in the comparison, each experiment was repeated five times, and the average result was reported.

[0151] The comparative models used in this patent application are of two types: non-graph models and graph models. The non-graph models include 21 models such as TFN, LMF, MFM, RAVEN, DialogueRNN, Multilogue-Net, MulT, TBJE, ALMT, SPECTRA, SPT, MAG-BERT, UniMSE, ICCN, MMLatch, MMIM, Self-MM, HMAI-BERT, Shapes-of-Emotion, MISA, and MMML. The graph models include Graph-MFN, MTAG, GraphCAGE, COGMEN, and MG.

[0152] The evaluation criteria used in this embodiment are divided into two forms: classification and regression. On the CMU-MOSI dataset, the classification tasks include binary classification and seven-class classification. For the binary classification task, the weighted F1 score (F1-Score) and binary classification accuracy (ACC2) are used as evaluation metrics. Specifically, when the test set does not contain neutral sentiment samples... At that time, the predicted sentiment score will be... Classified as positive Classified as negative; when the test set contains neutral sentiment samples, a sentiment score will be predicted. Classified as non-negative. Classified as negative. In the seven-class classification task, the seven-class classification accuracy (ACC7) is used as the evaluation metric. Specifically, the predicted sentiment score of the sample is... Rounding to the nearest integer in the range [-3, +3], the samples can be divided into 7 classes. Further, precision represents the proportion of correctly predicted samples in the test set. The F1 score, as the harmonic mean of precision and recall, ranges from 0 to 1. In sentiment regression tasks, mean absolute error (MAE) and Pearson correlation coefficient (Corr) are used as evaluation metrics. MAE measures the average absolute deviation between the model's predicted values ​​and the actual sentiment values; a smaller value indicates a better model fit. The Pearson correlation coefficient measures the linear correlation between predicted and actual values, defined as the ratio of the product of their covariance and their respective standard deviations, ranging from -1 to +1.

[0153] Furthermore, for the Chinese dataset CH-SIMS, we also conducted experiments on multiple classification metrics, covering binary, triangular, and quinary classification tasks. In the binary classification task, we used the weighted F1 score and binary classification accuracy as metrics, where samples with predicted sentiment scores falling within the interval [-1.0, 0.0] were considered to have negative emotions, while samples within the interval (0.0, 1.0) were considered to have positive emotions. In the triangular classification task, we used triangular classification accuracy (ACC3) as the evaluation metric. Specifically, when the predicted values ​​of different samples were mapped to the three intervals [-1.0, -0.1], (-0.1, 0.1], and (0.1, 1.0], they corresponded to negative emotions, neutral emotions, and positive emotions, respectively. Sentiment assessment. In the five-class classification task, the five-class classification accuracy (ACC5) was used as the evaluation metric. The predicted values ​​were divided into [-1.0, -0.7], (-0.7, -0.1], (-0.1, 0.1], (0.1, 0.7], and (0.7, 1.0], which correspond to strong negative sentiment, moderate negative sentiment, neutral sentiment, moderate positive sentiment, and strong positive sentiment, respectively. In the regression task, the same evaluation metrics as the CMU-MOSI dataset were used, namely mean absolute error (MAE) and Pearson correlation coefficient (Corr).

[0154] In summary, the multi-dimensional evaluation index system adopted in this application helps to comprehensively evaluate the performance of the model in multimodal sentiment analysis tasks.

[0155] Tables 1 and 2 show the multimodal sentiment classification performance of the baseline model and the proposed MCDH-Net on the CMU-MOSI and CH-SIMS datasets, respectively. Experimental results show that MCDH-Net exhibits superior classification performance on both datasets.

[0156] Specifically, on the CMU-MOSI dataset, MCDH-Net achieved the best results in all six classification metrics: 60.35% accuracy for five-class classification, 54.52% accuracy for seven-class classification, 87.76% accuracy for binary classification including neutral samples with a corresponding weighted F1 score of 87.71%, and 89.94% accuracy for binary classification after removing neutral samples with a corresponding weighted F1 score of 89.93%. Compared to 18 non-graph neural network baseline models, the improvements in each metric ranged from 3.94% to 17.67%, 4.18% to 19.12%, 0.25% to 6.57%, 0.26% to 7.61%, 0.25% to 8.24%, and 0.26% to 8.33%, respectively. Compared to 7 graph neural network baseline models, the improvements in each metric were 21.72%, 15.62% to 22.42%, 5.46% to 10.62%, 5.61% to 10.63%, 7.47% to 11.59%, and 7.24% to 11.58%, respectively. These results indicate that compared to existing methods, MCDH-Net shows the most significant performance improvement in five-class and seven-class classification tasks, followed by binary classification accuracy and F1 score after removing neutral samples; while the performance improvement in binary classification including neutral samples is relatively small. Furthermore, compared with existing baseline models based on graph neural networks, MCDH-Net still achieves significant performance gains, validating the effectiveness and advancement of the proposed method.

[0157] Table 1. Classification results of MCDH-Net and baseline models on CMU-MOSI

[0158]

[0159] On the CH-SIMS dataset, MCDH-Net achieves an accuracy of 81.58% in binary classification, slightly lower than MLF-DNN, MLMF, MTFN, MMML, and MDH by 0.70%, 0.74%, 0.87%, 1.35%, and 0.64%, respectively. However, it still represents a performance improvement of 0.39%–10.46% compared to other baseline models. Its corresponding F1 score is 82.03%, slightly lower than the aforementioned five models (with a difference of 0.36%–0.87%), but better than all other baseline methods, with an advantage ranging from 0.46% to 8.21%. In three-class and five-class classification tasks, MCDH-Net outperforms all baseline models, achieving accuracies of 71.71% and 50.44%, respectively, representing improvements of 2.34%–9.13% and 1.06%–16.52% over existing methods. As can be seen, MCDH-Net achieves the most significant improvement in the five-class and three-class classification tasks on the CH-SIMS dataset, while it is slightly inferior to some advanced models in terms of binary classification accuracy and F1 score.

[0160] Table 2. Classification results of MCDH-Net and baseline models on CH-SIMS

[0161]

[0162] Tables 3 and 4 present the regression results of MCDH-Net and various baseline models on the CMU-MOSI and CH-SIMS datasets, respectively. Specifically, on the CMU-MOSI dataset, MCDH-Net achieved the lowest MAE (0.5509) and the highest Corr (0.8871) among the 21 baseline models. Compared to 16 non-graph neural network methods, its MAE reduction ranged from 0.0322 to 0.3261, and its Corr improvement ranged from 0.0178 to 0.1871; compared to 5 graph neural network methods, MCDH-Net's MAE decreased by 0.3151 to 0.4048, and its Corr improvement ranged from 0.1651 to 0.2385. These results demonstrate that MCDH-Net's improvement in MAE is particularly significant; compared to graph-structured baseline models, dynamic hypergraph methods also exhibit significant advantages in sentiment regression. On the CH-SIMS dataset, MCDH-Net's MAE is 0.3513, slightly higher than MMML by 0.0193, but significantly lower than other baseline models, with a decrease between 0.0385 and 0.1196. Its Corr value is 0.7105, slightly lower than MMML (by 0.0221), but better than all other comparison models, with an improvement between 0.0358 and 0.1710. Although MCDH-Net's regression performance on this dataset is slightly inferior to MMML, its overall performance is still superior to all other baseline models, including MLF-DNN, MLMF, MTFN, and MDH.

[0163] In summary, MCDH-Net outperforms most existing baseline methods in multimodal sentiment classification tasks. On the CMU-MOSI dataset, it achieves state-of-the-art performance across all classification metrics; on the CH-SIMS dataset, it also achieves state-of-the-art performance in both 3-class and 5-class classification tasks. Although it slightly lags behind some state-of-the-art models (MLF-DNN, MLMF, MTFN, MMML, and MDH) in binary classification accuracy and F1 score, its overall performance significantly surpasses the other 11 baseline methods, fully validating the model's effectiveness and competitiveness in multimodal sentiment analysis. Furthermore, MCDH-Net also demonstrates superior performance in multimodal sentiment regression tasks. On the CMU-MOSI dataset, it achieves state-of-the-art results across all regression metrics; on the CH-SIMS dataset, although it slightly lags behind MMML in some metrics, it still significantly outperforms the other 14 baseline models.

[0164] Table 3. Regression results of MCDH-Net and baseline models on CMU-MOSI.

[0165]

[0166] Table 4. Regression results of MCDH-Net and baseline models on CH-SIMS.

[0167]

[0168] While the invention has been described with reference to specific embodiments, it should be understood that these embodiments are merely examples of the principles and applications of the invention. Therefore, it should be understood that many modifications can be made to the exemplary embodiments, and other arrangements can be designed without departing from the spirit and scope of the invention as defined by the appended claims. It should be understood that different dependent claims and features described in the invention can be combined in ways different from those described in the original claims. It is also understood that features described in conjunction with individual embodiments can be used in other described embodiments.

Claims

1. A multi-channel dynamic hypergraph sentiment analysis method integrating temporal consistency, characterized in that, Includes the following steps: S1. Multi-channel feature extraction steps: Extract multi-channel features of text and audio modalities respectively through heterogeneous pre-trained models; S2. Local Temporal Context Fusion Step Based on Video Number: Short-term sentiment fluctuations are fused using a local context mechanism based on video number, and long-range dependencies across time dimensions are captured through Transformer. S3. Single-modal-multimodal hypergraph collaborative prediction steps: Dynamically construct single-modal and multimodal hypergraphs within training batches, and use spectral-spatial hybrid convolution to model higher-order relationships; S4. Multi-level, multi-branch supervision steps: The final sentiment prediction result is output through joint optimization of the early MLP branch and the late hypergraph branch.

2. The multi-channel dynamic hypergraph sentiment analysis method with fused temporal consistency according to claim 1, characterized in that, The multi-channel feature extraction step described in S1 specifically includes: Text modality feature extraction: Text features are extracted using a dual-channel approach combining the BERT and RoBERTa models; Audio modality feature extraction: For the CMU-MOSI dataset, the COVAREP toolkit and Data2vec model were used to extract audio features in two channels. For the CH-SIMS dataset, the LibROSA model and HuberT model were used to extract audio features in two channels.

3. The multi-channel dynamic hypergraph sentiment analysis method with fused temporal consistency according to claim 1, characterized in that, The S2 local temporal context fusion steps based on video IDs specifically include: S21. Construct multi-channel context-enhanced features based on the local context mechanism of video number: For the current moment feature, fuse the features of the previous two moments of the video in which it is located, and fill in the missing features with the current features; S22. Establish global dependencies based on Transformer encoder: First, fuse the context enhancement features of multiple channels in each modality, and then introduce the Transformer model to process the fused multi-channel features and generate multi-modal fused features.

4. The multi-channel dynamic hypergraph sentiment analysis method with fused temporal consistency according to claim 3, characterized in that, The specific steps of S3 single-modal-multimodal hypergraph collaborative prediction include: S31. Single-modal dynamic hypergraph construction: A batch local strategy is adopted, and the hypergraph structure is constructed independently for each batch; S32. Hypergraph Hybrid Convolution: Perform spectral domain convolution and spatial domain convolution on unimodal features and multimodal fusion features respectively to generate sentiment prediction results based on unimodal hypergraphs and sentiment prediction results based on multimodal hypergraphs.

5. The multi-channel dynamic hypergraph sentiment analysis method with fused temporal consistency according to claim 4, characterized in that, The outputs of spectral domain convolution and spatial domain convolution are connected using residuals.

6. The multi-channel dynamic hypergraph sentiment analysis method with fused temporal consistency according to claim 1, characterized in that, The S4 multi-level, multi-branch supervision steps specifically include: S41. Utilize early MLP branches to directly perform regression prediction on Transformer pre-fusion features; S42. Introduce a multi-branch loss fusion structure, which integrates the losses of the single-modal hypergraph branch, the multi-modal hypergraph branch, and the MLP branch in a weighted manner to construct the final training objective.

7. A multi-channel dynamic hypergraph sentiment analysis network integrating temporal consistency, characterized in that... include Multi-channel feature encoding module: used to extract multi-channel features of text and audio modalities respectively through heterogeneous pre-trained models; Short-term dynamics and long-term dependency modeling module: used to fuse short-term sentiment fluctuations based on local context mechanism of video number, and capture long-term dependencies across time dimension through Transformer; Single- and multimodal hypergraph collaborative prediction module: used to dynamically construct single-modal and multimodal hypergraphs within training batches, and to model higher-order relationships using spectral-spatial hybrid convolution; Multi-level, multi-branch supervision module: used to jointly optimize through early MLP branches and late hypergraph branches to output the final sentiment prediction result.

Citation Information

Cited By

  • Postpartum comprehensive evaluation method based on multi-source heterogeneous data fusion

    CN122091215A

  • Visual emotion analysis method and system based on frequency domain enhancement and multi-attribute reasoning

    CN122416163A

  • A Visual Sentiment Analysis Method and System Based on Frequency Domain Enhancement and Multi-Attribute Inference

    CN122416163B