Multi-modal emotion calculation method and system based on adaptive reconstruction hypergraph
Through the adaptive reconstruction of the hypergraph method, the asynchronous data integration problem in multimodal sentiment analysis is solved, and efficient emotional information fusion between text, speech and visual modes is achieved, which improves the accuracy and interpretability of sentiment analysis.
Patent Information
- Application Number
- CN202510448175.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-10
- Publication Date
- 2025-07-25
AI Technical Summary
The prior art is difficult to effectively integrate asynchronous multimodal data, especially text, voice and visual modalities in videos, resulting in insufficient accuracy of sentiment analysis.
Adaptive reconstruction of hypergraph method is adopted to model the multi-dimensional emotional connection relationship of single-mode, dual-mode and three-mode through hypergraph structure, optimize the hypergraph module to adaptively adjust the connection relationship between modes, and use hypergraph convolution to fuse different modal features.
It improves the accuracy and interpretability of multimodal sentiment analysis, can effectively handle the timing asymmetry and asynchronousness of multimodal data, and reveals the interactive relationship between modes.
Smart Images

Figure CN120372354A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of natural language processing, and specifically, to a multi-modal sentiment computing method and system based on an adaptive reconstruction hypergraph. Background Art
[0002] The basic starting point of multi-modal sentiment computing is that different modalities in the same data are often complementary, providing additional semantic and sentiment recognition clues, making multi-modal systems more advantageous than single-modal systems in sentiment analysis. However, due to the asynchrony of multi-modal sequences, the behaviors expressing the same sentiment do not necessarily occur simultaneously. For example, a smile may be related to a positive word said in the past. Therefore, a key issue in the field of multi-modal sentiment analysis is how to effectively integrate heterogeneous modal data, extract and integrate the sentimentally meaningful interaction relationships from different modalities to obtain a unified representation.
[0003] To address the above asynchronous problem, some previous work often solved it by performing word-level alignment in the preprocessing stage. The specific approach was to force the lengths of the visual and speech feature sequences to be consistent with the text. This method not only increased additional computational resource consumption but also, since the interaction between modalities is not limited to a single word, this forced alignment was bound to cause information loss, thus affecting the final sentiment discrimination.
[0004] Subsequent methods often used recurrent neural networks, but they could only integrate single-modal information and could not model the interaction relationships between modalities. With the application of the attention mechanism in the field of multi-modal sentiment analysis, the interaction relationships between different modalities have been further explored. However, the current attention mechanism can only explore the relationships between two modalities. For example, a smile in the visual modality is related to a positive word in the text modality. However, for video sentiment analysis, which includes text, speech, and visual modalities, considering only the relationships between two modalities is insufficient. For example, a smile plus a positive word expresses a positive sentiment, while the speech has a sarcastic tone, resulting in a change in sentiment to a negative expression. Therefore, how to effectively model and analyze the sentiment relationships between different modalities remains a challenge.
[0005] Patent document CN118781524A (application number: 202411255941.X) discloses a multi-modal sentiment analysis method and system based on a two-stage heterogeneous hypergraph, belonging to the field of multi-modal sentiment analysis and prediction. The specific process is as follows: obtaining the original data of each modality in an audio-visual video clip, inputting the original data of each modality into a trained multi-modal sentiment analysis model, and outputting the probability of the sentiment polarity reflected in the audio-visual video clip; wherein, the original data of each modality is an audio clip, a video clip, and a dialogue text.
[0006] Based on the prior art, the present invention proposes a new multi-modal sentiment computing method and system based on adaptive reconstruction of hypergraphs. By modeling the multi-dimensional sentiment connection relationships of unimodal, bimodal, and trimodal through hypergraph structures, asynchronous data processing is achieved. At the same time, complementary sentiment information between different modalities is effectively extracted, and the interpretability of the model is improved through the edge connection relationships of the hypergraph. To ensure the effectiveness of the constructed hypergraph, the present invention designs an adaptive reconstruction hypergraph optimization hypergraph module, which adaptively changes the connection relationships of each time series point in the hypergraph according to the modal features, optimizes redundant connection relationships, and fuses the features of different modalities through hypergraph convolution. Summary of the Invention
[0007] Aiming at the deficiencies in the prior art, the object of the present invention is to provide a multi-modal sentiment computing method and system based on adaptive reconstruction of hypergraphs.
[0008] A multi-modal sentiment computing method based on adaptive reconstruction of hypergraphs provided by the present invention includes:
[0009] Step S1: Based on the video of the target object, obtain the text, speech, and visual sequences of the target object, and extract text features, speech features, and visual features based on the text, speech, and visual sequences of the target object;
[0010] Step S2: The text features, speech features, and visual features are respectively encoded through a feature encoder to obtain text, speech, and visual modalities with temporal relationships;
[0011] Step S3: Take all time series points of the text, speech, and visual modalities with temporal relationships as hypergraph nodes, and initialize the connection relationships between different time series points to obtain a hypergraph structure and hypergraph structure point features;
[0012] Step S4: For the hypergraph structure and hypergraph structure point features, use hypergraph convolution to aggregate the information between modalities to optimize the hypergraph structure point features and obtain optimized hypergraph structure point features;
[0013] Step S5: Obtain the point features of the last layer of the hypergraph structure, perform average pooling on the point features of the last layer, and aggregate the point features of the last layer into a final multi-modal sentiment vector representation; perform sentiment category judgment based on the final multi-modal sentiment vector representation.
[0014] Preferably, the step S1 includes:
[0015] Step S1.1: Based on the video of the target object, obtain the text, speech, and visual sequences of the target object;
[0016] Step S1.2: Based on the obtained text sequence of the target object, extract text features through the BERT model;
[0017] Step S1.3: Based on the obtained speech sequence of the target object, extract speech features through COVAREP. The speech features include: speech mel spectrogram features;
[0018] Step S1.4: Based on the obtained visual sequence of the target object, extract the video features of the speaker in the video frame through Facet. The video features include facial localization information.
[0019] Preferably, the step S2 includes: the text features, speech features, and visual features are respectively subjected to feature encoding through a bidirectional long short-term memory network to obtain text, speech, and visual modalities with temporal relationships; then, a fully connected network is used to map the different modality features to the same dimension;
[0020] X m = FC(BiLSTM(U m ))
[0021] where U m represents the single-modal extraction features, m ∈ {t, a, v} respectively represent the text, speech, and visual modalities; BiLSTM(·) and FC(·) respectively represent the bidirectional long short-term memory network and the fully connected network; represents the encoded single-modal features, T m represents the length of the temporal sequence, and d represents the unified mapping dimension.
[0022] Preferably, the step S3 includes:
[0023] Form the point set of the hypergraph by all the time series points of the text, speech, and visual modalities with temporal relationships;
[0024] N = [n t,1 ; n t,2 ; …; n a,1 ; n a,2 ; …; n v,1 ; n v,2 ;
[0025] where t, a, v respectively represent the text, speech, and visual modality representations;
[0026] The multi-modal hypergraph is represented as MHG = (N, E, W); where W represents the hyperedge weight; E represents the hyperedge feature;
[0027] Initialize the connection matrix through the self-attention mechanism. The hypergraph construction process for each layer is as follows:
[0028]
[0029] where W K and W Qrepresents the learnable weight matrix, and d is the dimension of the encoded feature.
[0030] Preferably, the step S4 includes:
[0031] Step S4.1: According to the constructed multi-modal hypergraph structure, aggregate the information of different modal time series points through hypergraph convolution:
[0032]
[0033] where l represents the number of layers, and θ m represents the learnable parameter matrix;
[0034] Step S4.2: Calculate the edge features of the hypergraph according to the updated point features and the connection matrix H (l) ;
[0035]
[0036] Step S4.3: Calculate the correlation between the point features and the edge features with each other;
[0037]
[0038] where W N and W E represent the learnable weights respectively;
[0039] Adopt a residual structure and introduce a self-loop matrix to ensure the availability of different point features:
[0040]
[0041] where W selfLoop is the identity matrix, and α and β are used to control the update ratio; after the connection matrix is updated, the above hypergraph convolution and hypergraph reconstruction are repeated for optimization.
[0042] Preferably, the step S5 includes:
[0043] Step S5.1: The output of the hypergraph of the last layer is N (L) , and the output N of the hypergraph of the last layer (L) is subjected to average pooling to obtain the multi-modal aggregation feature:
[0044] S = MeanPool(N (L) )
[0045] where MeanPool(·) represents average pooling,
[0046] Step S5.2: Input the multi-modal aggregation feature S into a multi-layer perceptron for final sentiment category judgment.
[0047] Preferably, the method further includes: calculating the difference between the true emotion and the predicted emotion, using the cross-entropy function as the task loss, and updating the parameters in the multi-modal emotion calculation method based on the adaptive reconstruction hypergraph through backpropagation.
[0048] A multi-modal emotion calculation system based on an adaptive reconstruction hypergraph according to the present invention includes:
[0049] Module M1: Obtain the text, speech, and visual sequences of the target object based on the video of the target object, and extract text features, speech features, and visual features based on the text, speech, and visual sequences of the target object;
[0050] Module M2: The text features, speech features, and visual features are respectively encoded through a feature encoder to obtain text, speech, and visual modalities with temporal relationships;
[0051] Module M3: Use all time series points of the text, speech, and visual modalities with temporal relationships as hypergraph nodes, and initialize the connection relationships between different time series points to obtain a hypergraph structure and hypergraph structure point features;
[0052] Module M4: For the hypergraph structure and hypergraph structure point features, use hypergraph convolution to aggregate information between modalities to optimize the hypergraph structure point features and obtain optimized hypergraph structure point features;
[0053] Module M5: Obtain the point features of the last layer of the hypergraph structure, perform average pooling on the point features of the last layer, and aggregate the point features of the last layer into a final multi-modal emotion vector representation; perform emotion category judgment based on the final multi-modal emotion vector representation.
[0054] Preferably, the module M1 includes:
[0055] Module M1.1: Obtain the text, speech, and visual sequences of the target object based on the video of the target object;
[0056] Module M1.2: Extract text features through a BERT model based on the obtained text sequence of the target object;
[0057] Module M1.3: Extract speech features through COVAREP based on the obtained speech sequence of the target object, and the speech features include: speech Mel spectrogram features;
[0058] Module M1.4: Extract video features of the speaker in the video frame through Facet based on the obtained visual sequence of the target object, and the video features include facial localization information;
[0059] The module M2 includes: text features, speech features, and visual features are respectively encoded through a bidirectional long short-term memory network to obtain text, speech, and visual modalities with temporal relationships; then, a fully connected network is used to map different modality features to the same dimension;
[0060] X m = FC(BiLSTM(U m ))
[0061] where U m represents single-modal extracted features, m ∈ {t, a, v} respectively represent text, speech, and visual modalities; BiLSTM(·) and FC(·) respectively represent a bidirectional long short-term memory network and a fully connected network; represents the encoded single-modal features, T m represents the length of the time series, and d represents the unified mapping dimension.
[0062] Preferably, the module M3 includes:
[0063] All time series points of the text, speech, and visual modalities with temporal relationships are used to form the point set of the hypergraph;
[0064] N = [n t,1 ; n t,2 ; …; n a,1 ; n a,2 ; …; n v,1 ; n v,2 ;
[0065] where t, a, v respectively represent text, speech, and visual modality representations;
[0066] The multi-modal hypergraph is represented as MHG = (N, E, W); where W represents the hyperedge weight; E represents the hyperedge feature;
[0067] The connection matrix is initialized through a self-attention mechanism, and the hypergraph construction process for each layer is as follows:
[0068]
[0069] where W K and W Q represent learnable weight matrices, and d is the dimension of the encoded features;
[0070] The module M4 includes:
[0071] Module M4.1: According to the constructed multi-modal hypergraph structure, the information of different modality time series points is aggregated through hypergraph convolution:
[0072]
[0073] Among them, l represents the number of layers, and θ m represents the learnable parameter matrix;
[0074] Module M4.2: Calculate the edge features of the hypergraph based on the updated point features and the connection matrix H (l) to obtain the edge features of the hypergraph;
[0075]
[0076] Module M4.3: Calculate the correlation between the point features and the edge features with each other;
[0077]
[0078] Among them, W N and W E respectively represent the learnable weights;
[0079] Adopt a residual structure and introduce a self-loop matrix to ensure the availability of different point features:
[0080]
[0081] Among them, W selfLoop is the identity matrix, and α and β are used to control the update ratio; After the connection matrix is updated, the above hypergraph convolution and hypergraph reconstruction are repeated for optimization;
[0082] The said module M5 includes:
[0083] Module M5.1: The hypergraph output of the last layer is N (L) , perform average pooling on the hypergraph output N of the last layer (L) to obtain the multi-modal aggregation feature:
[0084] S = MeanPool(N (L) )
[0085] Among them, MeanPool(·) represents average pooling,
[0086] Module M5.2: Input the multi-modal aggregation feature S into a multi-layer perceptron for final sentiment category judgment.
[0087] Compared with the prior art, the present invention has the following beneficial effects:
[0088] 1. In view of the problem that multi-modal sequences are asynchronous with each other, the present invention proposes a new multi-modal hypergraph method, aiming to effectively handle the temporal misalignment and asynchrony of multi-modal data. Specifically, all temporal points of different modalities such as text, speech, and vision are used as nodes of the hypergraph to establish a cross-modal and cross-time high-order relationship structure. Different from the traditional graph structure, the hypergraph can not only represent the one-to-one relationship between nodes, but also represent the high-order associations between multiple nodes, which is particularly suitable for modeling complex and asynchronous interaction relationships within and between modalities.
[0089] 2. The present invention models multi-modal emotion sequences through a hypergraph. The high-order connection relationship of the hypergraph helps to discover the common and complementary features between different modalities. However, the accuracy of its connection structure directly affects the subsequent extraction and fusion of modal emotion information, thus affecting the accuracy of emotion calculation. In order to obtain a more accurate emotion representation, the present invention designs a new adaptive hypergraph reconstruction module, which adaptively optimizes the connection relationship between nodes while updating the features of each temporal point, and generates the most suitable hypergraph structure for the task in an iterative manner.
[0090] 3. The present invention can improve the interpretability of multi-modal models. The connection relationship of the hypergraph not only intuitively provides the connection mode between each temporal point of different modalities in terms of structure, but also can effectively reveal the mutual influence and action mechanism of multi-modal sequences in the process of emotion expression. By constructing and analyzing the connection relationship between nodes and edges in the hypergraph, the model can capture the interaction patterns within the modality and between different modality features, and understand how they jointly act on the final result of emotion recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0091] Other features, objectives, and advantages of the present invention will become more apparent by reading the following detailed description of non-limiting embodiments with reference to the accompanying drawings:
[0092] Figure 1 It is a schematic diagram of a multi-modal emotion computing system based on an adaptive reconstruction hypergraph.
[0093] Figure 2 It is a flowchart of a multi-modal emotion computing method based on an adaptive reconstruction hypergraph. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0094] The present invention will be described in detail below with reference to specific embodiments. The following embodiments will help those skilled in the art to further understand the present invention, but do not limit the present invention in any form. It should be noted that those of ordinary skill in the art can make several changes and improvements without departing from the concept of the present invention. These all belong to the protection scope of the present invention.
[0095] Embodiment 1
[0096] A multi-modal sentiment calculation method based on an adaptive reconstructed hypergraph provided by the present invention is as follows Figure 1-2 shown, including:
[0097] Step 1: Obtain the text, speech, and visual sequences of the speaker, and respectively use BERT to extract text features, COVAREP to extract features such as speech mel spectrograms, and Facet to extract the facial localization information of the speaker in the video frames.
[0098] Step 2: Use bidirectional long short-term memory networks to respectively perform feature encoding on the text, speech, and visual modalities, and map the three-modal input sequences to the same dimension while maintaining the sequence timing information to reduce the heterogeneity between modal features.
[0099] Step 3: Use all the time series points of the text, speech, and visual modalities as hypergraph nodes, and use the self-attention mechanism to initialize the connection relationships between different time series points, where the initialized hypergraph structure and point features are the input for the adaptive hypergraph reconstruction module.
[0100] Step 4: For the modal point features and the initialized hypergraph structure, use hypergraph convolution to aggregate the information between modalities to obtain optimized point features. Since the hypergraph structure will directly affect the subsequent fusion performance, the present invention designs a hypergraph reconstruction module to reconstruct the connection relationships between the points of the hypergraph according to the new node features, and designs multiple layers of loops to gradually optimize the connection relationships of the time series points, where there are a total of L layers.
[0101] Step 5: The sentiment classification module takes the point features of the last layer of the hypergraph as input, first performs average pooling to aggregate all the point features into the final multi-modal sentiment vector representation, and then performs sentiment category judgment through a multi-layer perceptron.
[0102] The present invention simultaneously constructs a multi-modal hypergraph using different time series points of the three modalities, and can simultaneously model the sentiment relationships within and between the asynchronous sequences of the three modalities. In addition, by designing an adaptive hypergraph reconstruction module, it realizes the automatic synchronous capture of the relationships between the time series points of single-modal, dual-modal, and triple-modal, and removes the redundant connections between the time series points through multiple rounds of optimization, which helps the subsequent modal information fusion and improves the model performance. In addition, it can realize the mining of sentiment information while helping people understand the mechanism by which multi-modal information affects the final sentiment, and improves the interpretability of the model.
[0103] Further, in step 3, the text, speech, and visual modal temporal features jointly form the point set N of the hypergraph = [n t,1 ; n t,2 ; …; n a,1 ; n a,2 ; …; n v,1 ; n v,2, where \(t\), \(a\), and \(v\) represent the text, speech, and visual modalities respectively. The multi-modal hypergraph is represented as \(MHG=(N, E, W)\), where the hyper-edge weight \(W\) is usually set as the identity matrix. The connection matrix is initialized by the self-attention mechanism, and the hypergraph construction process of its first layer is as follows:
[0104]
[0105] where, \(W\) K and \(W\) Q represent learnable weight matrices, and \(d\) is the dimension of the encoded features.
[0106] Furthermore, in step 4, according to the constructed multi-modal hypergraph structure, the information of different modal time series points is aggregated through hypergraph convolution:
[0107]
[0108] where, \(l\) represents the number of layers, \(\theta\) m is a learnable parameter matrix. Subsequently, the updated point features and the connection matrix \(H\) (l) are input into the hypergraph reconstruction module. First, the edge features of the hypergraph are obtained from the point features:
[0109]
[0110] Subsequently, the correlation between the point features and the edge features is calculated.
[0111]
[0112] To ensure the continuity of the hypergraph sentiment information, the present invention adopts a residual structure and introduces a self-loop matrix to ensure the availability of different point features:
[0113]
[0114] where, \(W\) selfLoop is the identity matrix, and \(\alpha\) and \(\beta\) are used to control the update ratio. After the connection matrix is updated, the above hypergraph convolution and hypergraph reconstruction are repeated for optimization.
[0115] Furthermore, in step 5, the hypergraph output of the last layer is \(N\) (L) , first, average pooling is performed to obtain the multi-modal aggregation feature:
[0116] \(S = MeanPool(N\) (L) )
[0117] where, \(MeanPool(·)\) represents average pooling, Subsequently, it is input into a multi-layer perceptron for final sentiment category judgment.
[0118] Further, in order to calculate the difference between the true emotion and the predicted emotion, the present invention uses the cross-entropy function as the task loss and updates the parameters of the overall model through backpropagation.
[0119] The present invention also provides a multi-modal emotion calculation system based on an adaptive reconstructed hypergraph. The multi-modal emotion calculation system based on the adaptive reconstructed hypergraph can be implemented by executing the process steps of the multi-modal emotion calculation method based on the adaptive reconstructed hypergraph. That is, those skilled in the art can understand the multi-modal emotion calculation method based on the adaptive reconstructed hypergraph as a preferred embodiment of the multi-modal emotion calculation system based on the adaptive reconstructed hypergraph.
[0120] Embodiment 2
[0121] Embodiment 2 is a preferred example of Embodiment 1
[0122] According to a multi-modal emotion calculation system based on an adaptive reconstructed hypergraph provided by the present invention, it includes:
[0123] Feature encoder module: The input data of the three modalities come from different feature domains, and the features of their emotions have different expression forms. The quality of the constructed hypergraph directly affects the effect of subsequent modality fusion. To achieve the above process, first, a feature encoder is used to map the modality input information of different domains to the same feature dimension to maintain in the same feature space.
[0124] Hypergraph initialization module: For the encoded single-modal time-series features, the self-attention mechanism is used to construct the connection relationship between the time-series points of each modality to realize the initialization of the hypergraph structure.
[0125] Adaptive hypergraph reconstruction module: This module mainly includes a hypergraph convolution and a hypergraph reconstruction module. According to the initialized hypergraph structure and point features, the features of the nodes are optimized through hypergraph convolution. The hypergraph reconstruction module takes the new point-edge features and the original hypergraph structure as inputs for adaptive connection structure adjustment to optimize the connection relationship and remove redundant information. The point-edge features of the last layer are output to obtain the final multi-modal emotion expression.
[0126] Emotion classification module: For the point-edge features output by the last layer of the multi-modal hypergraph, a multi-layer perceptron is used as the emotion classifier, and the fused multi-modal emotion representation is used as the input for emotion discrimination, which can perform emotion category classification and emotion polarity classification.
[0127] Those skilled in the art know that, in addition to implementing the systems, devices and their respective modules provided by the present invention in the form of pure computer-readable program codes, it is entirely possible to make the systems, devices and their respective modules provided by the present invention be implemented in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, embedded microcontrollers, etc. by logically programming the method steps. Therefore, the systems, devices and their respective modules provided by the present invention can be regarded as a kind of hardware components, and the modules included therein for implementing various programs can also be regarded as the structures within the hardware components; the modules for implementing various functions can also be regarded as either software programs for implementing the methods or the structures within the hardware components.
[0128] The specific embodiments of the present invention have been described above. It should be understood that the present invention is not limited to the above specific embodiments, and those skilled in the art can make various changes or modifications within the scope of the claims, which does not affect the essence of the present invention. Without conflict, the embodiments of the present application and the features in the embodiments can be combined with each other arbitrarily.
Claims
1. A multimodal sentiment computing method based on an adaptive reconstructed hypergraph, characterized in that Including: Step S1: Based on the video of the target object, obtain the text, speech, and visual sequences of the target object, and extract text features, speech features, and visual features based on the text, speech, and visual sequences of the target object; Step S2: The text features, speech features, and visual features are respectively encoded through a feature encoder to obtain text, speech, and visual modalities with temporal relationships; Step S3: All time series points of the text, speech, and visual modalities with temporal relationships are used as hypergraph nodes, and the connection relationships between different temporal points are initialized to obtain a hypergraph structure and hypergraph structure point features; Step S4: For the hypergraph structure and hypergraph structure point features, use hypergraph convolution to aggregate information between modalities to optimize the hypergraph structure point features and obtain optimized hypergraph structure point features; Step S5: Obtain the point features of the last layer of the hypergraph structure, perform average pooling on the point features of the last layer, and aggregate the point features of the last layer into a final multi-modal sentiment vector representation; perform sentiment category judgment based on the final multi-modal sentiment vector representation.
2. The multimodal sentiment computing method based on an adaptive reconstructed hypergraph according to claim 1, wherein The said Step S1 includes: Step S1.1: Based on the video of the target object, obtain the text, speech, and visual sequences of the target object; Step S1.2: Based on the obtained text sequence of the target object, extract text features through the BERT model; Step S1.3: Based on the obtained speech sequence of the target object, extract speech features through COVAREP, and the speech features include: speech Mel spectrogram features; Step S1.4: Based on the obtained visual sequence of the target object, extract the video features of the speaker in the video frame through Facet, and the video features include face localization information.
3. The multimodal sentiment calculation method based on an adaptive reconstruction hypergraph according to claim 1, wherein, The said Step S2 includes: The text features, speech features, and visual features are respectively encoded through a bidirectional long short-term memory network to obtain text, speech, and visual modalities with temporal relationships; then use a fully connected network to map different modality features to the same dimension; X m = FC(BiLSTM(U m )) Among them, U m represents single-modal extraction features, where m ∈ {t, a, v} represent text, speech, and visual modalities respectively; BiLSTM(·) and FC(·) represent bidirectional long short-term memory network and fully connected network respectively; represents the encoded single-modal features, T m represents the length of the time series, and d represents the unified mapping dimension.
4. The multimodal sentiment computing method based on an adaptive reconstructed hypergraph according to claim 1, wherein The said Step S3 includes: All time series points of the text, speech, and visual modalities with temporal relationships form the point set of the hypergraph; N = [n t,1 ; n t,2 ; …; n a,1 ; n a,2 ; …; n v,1 ; n v,2 ; Among them, t, a, and v respectively represent text, speech, and visual modality representations; The multi-modal hypergraph is represented as MHG=(N, E, W); where, W represents the hyperedge weight; E represents the hyperedge feature; Initialize the connection matrix through the self-attention mechanism, and the hypergraph construction process of each layer is as follows: Among them, W K and W Q represent learnable weight matrices, and d is the dimension of the encoded features.
5. The multimodal sentiment computing method based on an adaptive reconstruction hypergraph according to claim 1, wherein The said Step S4 includes: Step S4.1: According to the constructed multi-modal hypergraph structure, aggregate information of different modality temporal points through hypergraph convolution: Among them, l represents the number of layers, and θ m represents the learnable parameter matrix; Step S4.2: Calculate the edge features of the hypergraph based on the updated point features and the connection matrix H (l) Step S4.3: Calculate the correlation between point features and edge features with each other; Among them, W N and W E represent learnable weights respectively; Adopt a residual structure and introduce a self-loop matrix to ensure the availability of different point features: where W selfLoop is the identity matrix, and α and β are used to control the update ratio; after the connection matrix is updated, the above hypergraph convolution and hypergraph reconstruction are repeated for optimization.
6. The multimodal sentiment computing method based on an adaptive reconstructed hypergraph according to claim 1, wherein The said Step S5 includes: Step S5.1: The hypergraph output of the last layer is N (L) , and the hypergraph output N (L) of the last layer is subjected to average pooling to obtain a multi-modal aggregation feature: S = MeanPool(N (L) ) Among them, MeanPool(·) represents average pooling, Step S5.2: Input the multi-modal aggregation feature S into a multi-layer perceptron for final sentiment category judgment.
7. The multimodal sentiment calculation method based on an adaptive reconstruction hypergraph according to claim 1, characterized in that The said method further includes: Calculate the difference between the true sentiment and the predicted sentiment, use the cross-entropy function as the task loss, and update the parameters in the multi-modal sentiment calculation method based on the adaptive reconstruction hypergraph through backpropagation.
8. A multimodal sentiment computing system based on an adaptive reconstructed hypergraph, characterized in that, Including: Module M1: Obtain the text, speech, and visual sequences of the target object based on the video of the target object, and extract text features, speech features, and visual features based on the text, speech, and visual sequences of the target object; Module M2: The text features, speech features, and visual features are respectively feature-encoded through a feature encoder to obtain text, speech, and visual modalities with temporal relationships; Module M3: All time series points of the text, speech, and visual modalities with temporal relationships are used as hypergraph nodes, and the connection relationships between different temporal points are initialized to obtain a hypergraph structure and hypergraph structure point features; Module M4: For the hypergraph structure and hypergraph structure point features, use hypergraph convolution to aggregate information between modalities to optimize the hypergraph structure point features and obtain optimized hypergraph structure point features; Module M5: Obtain the point features of the last layer of the hypergraph structure, perform average pooling on the point features of the last layer, and aggregate the point features of the last layer into a final multi-modal sentiment vector representation; perform sentiment category judgment based on the final multi-modal sentiment vector representation.
9. The multimodal sentiment computing system based on an adaptive reconstructed hypergraph according to claim 8, wherein The module M1 includes: Module M1.1: Obtain the text, speech, and visual sequences of the target object based on the video of the target object; Module M1.2: Extract text features through the BERT model based on the obtained text sequence of the target object; Module M1.3: Extract speech features through COVAREP based on the obtained speech sequence of the target object, and the speech features include: speech Mel spectrogram features; Module M1.4: Extract the video features of the speaker in the video frame through Facet based on the obtained visual sequence of the target object, and the video features include facial localization information; The module M2 includes: The text features, speech features, and visual features are respectively feature-encoded through a bidirectional long short-term memory network to obtain text, speech, and visual modalities with temporal relationships; then use a fully connected network to map different modality features to the same dimension; X m = FC(BiLSTM(U m )) Among them, U m represents single-modal extraction features, where m ∈ {t, a, v} represent text, speech, and visual modalities respectively; BiLSTM(·) and FC(·) represent bidirectional long short-term memory network and fully connected network respectively; represents the encoded single-modal features, T m represents the length of the time series, and d represents the unified mapping dimension.
10. The multimodal sentiment computing system based on an adaptive reconstructed hypergraph according to claim 8, characterized in that, The module M3 includes: Form the point set of the hypergraph with all time series points of the text, speech, and visual modalities with temporal relationships; N = [n t,1 ; n t,2 ; …; n a,1 ; n a,2 ; …; n v,1 ; n v,2 ; Among them, t, a, and v respectively represent the text, speech, and visual modality representations; The multi-modal hypergraph is represented as MHG=(N, E, W); where, W represents the hyperedge weight; E represents the hyperedge feature; Initialize the connection matrix through the self-attention mechanism, and the hypergraph construction process of each layer is as follows: Among them, W K and W Q represent learnable weight matrices, and d is the dimension of the encoded features; The module M4 includes: Module M4.1: According to the constructed multi-modal hypergraph structure, aggregate information of different modality time series points through hypergraph convolution: Among them, l represents the number of layers, and θ m represents the learnable parameter matrix; Module M4.2: According to the updated point features and the connection matrix H (l) calculate the edge features of the hypergraph; Module M4.3: Calculate the correlation between the point features and the edge features with each other; Among them, W N and W E represent learnable weights respectively; Adopt a residual structure and introduce a self-loop matrix to ensure the usability of different point features: Among them, W selfLoop is the identity matrix, and α and β are used to control the update ratio; after the connection matrix is updated, the above hypergraph convolution and hypergraph reconstruction are repeated for optimization; The module M5 includes: Module M5.1: The hypergraph output of the last layer is N (L) , the hypergraph output N of the last layer (L) is subjected to average pooling to obtain multi-modal aggregation features: S = MeanPool(N (L) ) Among them, MeanPool(·) represents average pooling, Module M5.2: Input the multi-modal aggregation feature S into a multi-layer perceptron for final sentiment category judgment.
Citation Information
Patent Citations
Multi-modal sentiment analysis method and system based on dual-stage heterogeneous hypergraph
CN118781524A
A multimodal sentiment analysis method and system based on two-stage heterogeneous hypergraph
CN118781524B