Multi-layer graph multi-mode sentiment analysis method and device based on meta-information
By introducing the method of meta-information alignment and multi-level feature fusion in multi-modal sentiment analysis, the problem of insufficient meta-information alignment in the prior art is solved, and more accurate multi-modal feature representation and higher sentiment analysis accuracy are achieved.
Patent Information
- Application Number
- CN202510013886.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-06
- Publication Date
- 2025-05-16
AI Technical Summary
The existing multimodal sentiment analysis methods ignore the alignment of meta information, resulting in the inaccurate representation of multimodal feature, which affects the prediction performance.
Using a multi-level graph multi-modal sentiment analysis method based on meta-information, high-level features of text, video and audio are extracted through BERT, OpenFace 2.0 and Librosa, and time-series encoding is performed using BiLSTM. Then, the modal invariant and specific features are captured by a shared and private multimodal encoder, a model diagram is constructed and the meta information within the modal is aligned by the GNN algorithm. Finally, the feature fusion is performed through the graph knowledge distillation module and the Transformer network to generate emotional values.
Effectively extract and fuse the timing, global and local information in the modal features, improve the accuracy of feature representation of multimodal sentiment analysis, alleviate the problem of inconsistent cross-modal distribution, and improve the overall effect and accuracy of sentiment analysis.
Smart Images

Figure CN120012008A_ABST
Abstract
Description
Technical Field
[0001] The present invention proposes a multi-level graph multimodal sentiment analysis method and device based on meta-information, which belongs to the field of deep learning, especially the research direction of natural language processing, and involves the application field of cross-modal sentiment analysis of vision, audio and language. Background Art
[0002] Multimodal sentiment analysis (MSA) aims to identify human emotions by integrating multimodal information. Compared with unimodal methods, multimodal sentiment analysis exploits cross-modal relations to provide additional clues for semantic and sentiment disambiguation, thereby improving prediction performance. In recent years, multimodal sentiment analysis methods can be roughly divided into two categories: one is representation learning methods, which aim to enhance modal semantics by effectively modeling inter-modal correlations. These methods emphasize capturing complex dependencies and interactions between modalities while optimizing single modal features. By leveraging this rich modal semantics, these methods provide a foundation for more detailed human sentiment analysis and improve the downstream efficiency of multimodal fusion in relational modeling. The other is methods centered on multimodal fusion, which mainly focus on developing advanced techniques to construct a unified joint representation of multimodal data. These methods usually rely on complex fusion mechanisms such as tensor-based fusion, cross-attention mechanisms, or hierarchical architectures to synergistically integrate heterogeneous modal features. By explicitly optimizing the integration of multimodal inputs, such methods aim to preserve the complementary information between modalities, thereby achieving superior performance in tasks that require a comprehensive understanding of multimodal relations.
[0003] Although these studies have made significant progress, many methods only focus on global feature alignment and often ignore the alignment of meta-information (i.e., a single word in text, a frame in video, and a clip in audio). This neglect may lead to inaccurate multimodal feature representations, because the meta-information may contain information that is irrelevant or conflicting to sentiment, thus affecting the prediction performance. Summary of the invention
[0004] In order to overcome the above-mentioned shortcomings of the prior art and solve the problem of sentiment analysis in multimodality, the present invention proposes a multi-level graph multimodal sentiment analysis method and device based on meta-information.
[0005] To solve the above problems, the technical solution provided by the present invention is:
[0006] A first aspect of the present invention provides a multi-level graph multimodal sentiment analysis method based on meta-information, comprising the following steps:
[0007] Step 1: BERT, OpenFace 2.0, and Librosa tools are used to extract features from the input text, video, and audio data to obtain high-level representations of each modality. BiLSTM is used to perform temporal encoding on the extracted video and audio features. Subsequently, MLP is used to map the features of different modalities to a unified dimensional space.
[0008] Step 2: To capture both the consistency and specificity of the modalities, a shared multimodal encoder and three private multimodal encoders are used to construct modality-invariant features and modality-specific features.
[0009] Step 3: For modal-specific features, the GNN algorithm is used to construct a modal graph. For each modal graph, the weight of the graph edge is calculated based on the grammatical dependency tree and Self-Attention to align the meta-information within the modality.
[0010] Step 4: Use Cross-Attention to align the video and audio features with the text features respectively, so as to construct a new multimodal graph and capture the correlation information between the modalities.
[0011] Step 5: Apply convolution operations, BiLSTM and GNN to the constructed new modal graph to extract local features, temporal features and global features. Based on the cross-modal attention relationship between the features, the edge weights of the graph are constructed to achieve multi-level modal feature fusion and generate multi-level feature graphs of each modality.
[0012] Step 6: Input the multi-level feature graphs and modality-invariant features of each modality into the graph knowledge distillation module respectively.
[0013] Step 7: Flatten the multi-level feature graphs of each modality and reduce the dimension through a private encoder to reduce the redundancy and dimension of the features. Weightedly fuse the modality-invariant features of each modality and further reduce the dimension through a shared encoder.
[0014] Step 8: Stack the reduced features and further process them through the Transformer network with a multi-head attention mechanism to finally obtain the fused sentiment value.
[0015] The second aspect of the present invention relates to a multi-level graph multimodal sentiment analysis device based on meta-information, comprising a memory and one or more processors, wherein the memory stores executable code, and when the one or more processors execute the executable code, they are used to implement a multi-level graph multimodal sentiment analysis method based on meta-information according to any one of claims 1 to 9.
[0016] In view of the heterogeneity between multimodal data and the resulting differences in distribution patterns, the present invention is committed to solving the challenge of multimodal feature fusion. By introducing a multi-level feature generation module and a graph distillation unit, the present invention can achieve intra-modal and inter-modal meta-information alignment, effectively improve the accuracy of feature representation in multimodal sentiment analysis, and alleviate the problem of inconsistent cross-modal distribution.
[0017] The advantages of the present invention are:
[0018] 1. Through multi-level feature graphs, the temporal, global and local information in modal features can be effectively extracted and integrated, thereby enriching the expressive power of modal features. The temporal features capture the temporal dependencies in each modal data through BiLSTM, revealing the dynamic changes of the data in time; the global features capture the overall structure and dependencies between modalities through the GNN model, helping to reveal the global pattern; the local features extract local information in the modality through convolution operations, highlighting important details and local interactions. These multi-level features work together to make the feature representation of each modality more comprehensive and refined.
[0019] 2. The performance of the model is further improved through the graph distillation unit. The graph distillation unit uses knowledge distillation to effectively transfer and optimize knowledge between modal features, making the fusion of modal features more coordinated and consistent. It can effectively reduce the heterogeneity between modal features and keep the features of different modalities consistent in the same representation space, thereby improving the overall effect and accuracy in multimodal learning tasks. In this way, the model can not only capture the characteristics of each modality, but also overcome the differences between modalities to achieve higher performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 It is a schematic diagram of the overall process of the method of the present invention.
[0021] Figure 2 It is the core architecture diagram of the multi-level feature graph of the present invention.
[0022] Figure 3 It is the core architecture diagram of the graph distillation unit of the present invention.
[0023] Figure 4 It is the core architecture diagram of the fusion module of the present invention.
[0024] Figure 5 It is a schematic diagram of the device of the present invention. DETAILED DESCRIPTION
[0025] The technical solution of the present invention is further described below in conjunction with the accompanying drawings.
[0026] Example 1
[0027] For most sentiment analysis, the features before input fusion are mostly global features. Using only global features will result in the loss of a lot of temporal information and local information, resulting in poor final sentiment classification results. The present invention obtains temporal, local and global information through multi-level feature graphs before fusion, enriches feature representation, and guides the model to correctly predict sentiment categories.
[0028] This embodiment relates to a multi-level graph multi-modal sentiment analysis method based on meta-information, such as Figure 1 , including the following steps:
[0029] Step 1: BERT, OpenFace 2.0, and Librosa tools are used to extract features from the input text, video, and audio data to obtain high-level representations of each modality. BiLSTM is used to perform temporal encoding on the extracted video and audio features. Subsequently, MLP is used to map the features of different modalities to a unified dimensional space.
[0030] Step 2: To capture both the consistency and specificity of the modalities, a shared multimodal encoder and three private multimodal encoders are used to construct modality-invariant features and modality-specific features.
[0031] Step 3: For modal-specific features, the GNN algorithm is used to construct a modal graph. For each modal graph, the weight of the graph edge is calculated based on the grammatical dependency tree and Self-Attention. The grammatical dependency tree is extracted from the text using the SpaCy toolkit.
[0032] Step 4: Use Cross-Attention to align the video and audio features with the text features respectively, so as to construct a new multimodal graph and capture the correlation information between the modalities.
[0033] Step 5: Apply convolution operations, BiLSTM and GNN to the constructed new modal graph to extract local features, temporal features and global features. Based on the cross-modal attention relationship between the features, the edge weights of the graph are constructed to achieve multi-level modal feature fusion and generate multi-level feature graphs of each modality, such as Figure 2 shown.
[0034] Step 6: Input the multi-level feature graphs and modality-invariant features of each modality into the graph knowledge distillation module respectively. The purpose is to effectively transfer and optimize knowledge between modality features so that the fusion of modality features is more coordinated. It can effectively reduce the heterogeneity between modality features and make the features of different modalities consistent in the same representation space, such as Figure 3 shown.
[0035] Step 7: Flatten the multi-level feature graphs of each modality and reduce the dimension through a private encoder to reduce the redundancy and dimension of the features. Weighted fusion of the modality-invariant features of each modality and further reduce the dimension through a shared encoder, such as Figure 4 shown.
[0036] Step 8: Stack the reduced features and further process them through the Transformer network with multi-head attention mechanism to finally obtain the fused sentiment value, such as Figure 4 shown.
[0037] In one embodiment of the present invention, step 1 specifically includes: using the BERT model to extract features from the input text, capturing the grammatical and semantic information of the text, converting it into a high-dimensional vector representation, and providing text features for multimodal fusion. At the same time, OpenFace 2.0 is used to extract facial expression features in the video, identify facial movements related to emotions, and provide video features for sentiment analysis. Audio data is extracted through Librosa, such as Mel-frequency cepstral coefficients (MFCC), to capture the emotional color of the audio and support sentiment analysis. Video and audio features are temporally encoded through BiLSTM, and then all modal features are mapped to a unified dimensional space through MLP, providing rich information for subsequent multimodal fusion and sentiment analysis.
[0038] X t =BERT(x t ) (1)
[0039] X a =BiLSTM(Librosa(x a )) (2)
[0040] X v =BiLSTM(OpenFace(x v )) (3)
[0041] u m =MLP m (X m ),m∈{t,v,a} (4)
[0042] In one embodiment of the present invention, step 2 specifically includes: in order to capture the consistency and specificity of the modalities at the same time, a shared multimodal encoder and three private multimodal encoders are used to construct a modality invariant feature c m and the modal-specific feature p m Private encoders are used to extract features of each modality separately, ensuring that the specificity of each modality is fully preserved. At the same time, shared encoders help capture the common rules between modalities, thereby obtaining modality-invariant feature representations.
[0043]
[0044] In one embodiment of the present invention, step 3 specifically includes: for the modal-specific features, a GNN algorithm is used to construct a modal graph, in which nodes represent feature units and edges represent the relationship between units. b and self-attention mechanism A m To calculate the graph edge weight, the former captures the grammatical relationship between words, and the latter adjusts the edge weight by similarity to enhance the relationship expression of important nodes. Through these two mechanisms, the edge weight of the modal graph is accurately constructed, the expressive power of the modal features is improved, and the complex structure and dependency within the modality are captured. Finally, rich feature representation is obtained through GNN propagation and learning.
[0045]
[0046] In one embodiment of the present invention, step 4 specifically includes: using a graph node alignment strategy to align the graph structure of the text modality with the graph structure of the video and audio modalities, thereby enhancing the information fusion between the modalities. Specifically, the graph node alignment strategy updates the nodes of the video and audio graphs through the Cross-Attention mechanism, aligns their representations with the text graph, enhances the information complementarity between the modalities, optimizes the feature representation of the video and audio graphs, and improves the semantic consistency across modalities.
[0047]
[0048] In one embodiment of the present invention, step 5 specifically includes: in order to enhance the richness of features, for the newly constructed modal graph, convolution operations are used to extract local features, BiLSTM captures timing features, and GNN extracts global features. The correlation between modalities is calculated through a cross-modal attention mechanism, and the weights of the graph edges are dynamically adjusted to achieve multi-level feature fusion and generate a comprehensive modal feature graph. First, the Readout module is used to aggregate node features through weighted pooling, and the attention mechanism is combined to highlight important nodes. LSTM features are weighted through a timing-dependent attention mechanism, and graph features focus on key node interactions through attention pooling. Subsequently, BiLSTM processes the timing data, and convolution operations extract local information. Finally, the LSTM, graph features, and convolution features are stacked for GNN to perform global information fusion, and Cross-Attention is used to calculate the relationship between features to obtain a richer feature representation.
[0049]
[0050] In one embodiment of the present invention, step 6 specifically includes: the multi-level feature graphs and modality-invariant features of each modality are then input into the graph knowledge distillation (GD) module. Specifically, graph distillation units (GD-Units) are used to semantically align features and enhance the semantic relationship between modality-invariant features. In order to effectively filter redundant information and improve the alignment effect, modality-specific GD units (Spec-GD) are used to process multi-level feature graphs, and for modality-invariant features, modality-invariant GD units (Inv-GD) are used to promote complementarity between multi-modal features. Specifically, for multi-level feature graphs, each modality graph is first complemented with the other two modality graphs through Cross-Attention, and then the aligned features are input into GD-Units; for modality-invariant features, they are directly input into GD-Units for further alignment.
[0051] In one embodiment of the present invention, step 7 specifically includes: after the multi-level feature graphs of each modality are flattened, the dimensionality is reduced through a private encoder to reduce feature redundancy and high-dimensional computational burden, effectively extract core information and compress unimportant features. For modality-invariant features, a weighted fusion strategy is adopted to combine the complementary information of each modality, and then the dimensionality is reduced and unified through a shared encoder to ensure that the representation scales between different modalities are consistent, significantly reduce computational complexity, retain key information, and improve the efficiency of multimodal feature fusion and the generalization ability of the model.
[0052] c=c t +c v +c a (17)
[0053]
[0054] In one embodiment of the present invention, step 8 specifically includes: the reduced-dimensional features are stacked to form a feature matrix The multi-head attention mechanism is used to process the Transformer network to enhance cross-modal and cross-subspace feature interactions. The multi-head attention mechanism enables each feature to effectively perceive the complementary information in other modal features, thereby improving the overall modeling ability of sentiment orientation and obtaining the feature matrix Finally, the processed feature vectors are concatenated into a joint representation and input into the fusion module G for final prediction. The fusion module further integrates and refines the features through linear transformation, nonlinear activation and regularization steps to generate the final sentiment prediction results.
[0055]
[0056] The present invention extracts the features of text, video and audio by using BERT, OpenFace 2.0 and Librosa, and performs temporal encoding through BiLSTM. Then, the features of different modalities are mapped to a unified dimensional space through MLP. Next, shared and private multimodal encoders are used to capture modality invariance and specific features, and a modal graph is constructed through GNN, and edge weights are calculated using grammatical dependency trees and self-attention. Video and audio features are aligned with text through Cross-Attention, a multimodal graph is constructed, and local, temporal and global features are extracted through convolution, BiLSTM and GNN. Further, features are fused through cross-modal attention relationships to generate multi-level feature graphs, and dimension reduction is performed through a graph knowledge distillation module and a private encoder. Finally, Transformer and multi-head attention mechanisms are used to fuse features to obtain sentiment values. Compared with previous methods, the present invention has a certain improvement in accuracy.
[0057] Example 2
[0058] like Figure 5 This embodiment relates to a multi-level graph multi-modal sentiment analysis device based on meta-information, including a memory and one or more processors, wherein the memory stores executable code, and when the one or more processors execute the executable code, they are used to implement the multi-level graph multi-modal sentiment analysis method based on meta-information described in Example 1.
[0059] The contents described in the embodiments of this specification are merely an enumeration of the implementation forms of the inventive concept. The protection scope of the present invention should not be regarded as limited to the specific forms described in the embodiments. The protection scope of the present invention also extends to equivalent technical means that can be conceived by those skilled in the art based on the inventive concept.
Claims
1. A multi-level graph multimodal sentiment analysis method based on meta-information, comprising the following steps: Step 1: BERT, OpenFace 2.0, and Librosa are used to extract features from the input text, video, and audio data, respectively, to obtain high-level representations of each modality. BiLSTM is used to perform temporal encoding on the extracted video and audio features. Subsequently, MLP is used to map the features of different modalities to a unified dimensional space. Step 2: To capture both the consistency and specificity of the modalities, a shared multimodal encoder and three private multimodal encoders are used to construct modality-invariant features and modality-specific features; Step 3: For modal-specific features, the GNN algorithm is used to construct a modal graph. For each modal graph, the weight of the graph edge is calculated based on the SyntacticDependency Tree and Self-Attention to align the meta-information within the modality. Step 4: Use Cross-Attention to align video and audio features with text features, thereby constructing a new multimodal graph, capturing the correlation information between the modalities, and aligning meta-information between the modalities. Step 5: Apply convolution operation, BiLSTM and GNN to the constructed new modal graph to extract local features, temporal features and global features; Based on the cross-modal attention relationship between features, multi-level modal feature fusion is achieved by constructing the edge weights of the graph, generating multi-level feature graphs of each modality; Step 6: Input the multi-level feature graphs and modality-invariant features of each modality into the graph knowledge distillation module respectively; Step 7: Flatten the multi-level feature graphs of each modality and reduce the dimension through a private encoder to reduce the redundancy and dimension of the features; perform weighted fusion of the modality-invariant features of each modality and further reduce the dimension through a shared encoder; Step 8: Stack the reduced features and further process them through the Transformer network with a multi-head attention mechanism to finally obtain the fused sentiment value.
2. According to the multi-level graph multimodal sentiment analysis method based on meta-information according to claim 1, the step 1 specifically comprises: Use the BERT model to extract features from the input text data; BERT converts input text into high-dimensional vector representation, captures the grammatical and semantic features of the text, and provides basic text feature representation for subsequent multimodal fusion; At the same time, OpenFace 2.0 is used to extract features from the input video data; key facial expressions and emotional information in the video are identified, thereby providing emotional features of the video data for multimodal sentiment analysis; In addition, the Librosa toolkit is used to extract features from the input audio data; Capture the emotional color in the audio and provide audio features for subsequent sentiment analysis; BiLSTM is used to perform temporal encoding on the extracted video and audio features to capture the temporal dependencies between these modalities. Subsequently, the features of different modalities are mapped to a unified dimensional space through MLP, laying the foundation for subsequent multimodal feature fusion; BERT, OpenFace 2.0, and Librosa tools are used to extract high-level feature representations of text, video, and audio, and temporal encoding and mapping of video and audio features are performed to provide rich modal information for subsequent multimodal fusion and sentiment analysis. X t =BERT(x t ) (1) X a =BiLSTM(Librosa(x a )) (2) X v =EiLSTM(OpenFace(x v )) (3) u m =MLP m (X m ),m∈{t,v,a} (4)。 3. According to the multi-level graph multimodal sentiment analysis method based on meta-information according to claim 1, the step 2 specifically comprises: To capture both the consistency and specificity of the modalities, a shared multimodal encoder and three private multimodal encoders are used to construct modality-invariant features and modality-specific features; Private encoders are used to extract features of each modality separately, ensuring that the specificity of each modality is fully preserved; at the same time, shared encoders help capture the common rules between modalities, thereby obtaining modality-invariant feature representations; 4. According to the multi-level graph multimodal sentiment analysis method based on meta-information according to claim 1, the step 3 specifically comprises: For modality-specific features, the GNN algorithm is used to construct the graph structure of each modality. Specifically, based on the feature data of each modality, the corresponding modality graph is first constructed, where the nodes of the graph represent the feature units of the modality, and the edges of the graph represent the relationship between the feature units. In order to enhance the expressiveness of the graph structure, the weights of the graph edges are further calculated in combination with Syntactic DependencyTree and Self-Attention. The grammatical dependency tree is used to capture the grammatical structure relationship between words in the text. By analyzing the dependency relationship between words, weights are assigned to the edges of the graph to reflect the semantic relevance of different words, denoted by A. b The self-attention mechanism calculates the similarity between feature units and dynamically adjusts the weight of the graph edge, thereby strengthening the relationship between important nodes, denoted as A m ; By combining these two mechanisms, it is possible to construct accurate edge weights for each modal graph, further improving the expressiveness of modal-specific features and helping to capture the complex structure and dependencies within the modality; finally, these modal graphs are propagated and learned using graph neural networks, thereby obtaining a richer and more accurate representation of modal features; 5. According to the multi-level graph multimodal sentiment analysis method based on meta-information according to claim 1, the step 4 specifically comprises: The graph node alignment strategy is used to align the graph structure of the text modality with the graph structure of the video and audio modalities, thereby enhancing the information fusion between the modalities. Specifically, the graph node alignment strategy updates the nodes of the video and audio graphs through the Cross-Attention mechanism to align their representations with the text graph, enhance the information complementarity between the modalities, optimize the feature representation of the video and audio graphs, and improve the semantic consistency across modalities.
6. According to the multi-level graph multimodal sentiment analysis method based on meta-information according to claim 1, the step 5 specifically comprises: In order to enhance the richness and comprehensiveness of features, for the newly constructed modal graph, convolution operations are used to extract local features, BiLSTM is used to capture temporal features, and GNN is used to extract global features; Subsequently, the correlation between the features of each modality is calculated through the cross-modal attention mechanism, and the edge weights of the graph are dynamically adjusted based on this, so as to achieve effective fusion of multi-level modal features and finally generate a modal feature graph containing multi-level information. In this multi-modal feature fusion process, the Readout module is first used. This module uses a weighted pooling method to aggregate node features. First, a dynamic attention score is calculated for each node, and these attention scores are learned based on the input features. The pooling process selectively focuses on the most relevant nodes through masking operations; for the features processed by LSTM, the attention mechanism can highlight the temporal dependencies between nodes; and for the features of the graph structure, attention pooling focuses on the most important node interactions; finally, weighted pooling is used to generate an aggregated representation of node features; The attention mechanism is applied to the time series features and graph structure features extracted by LSTM respectively, so as to highlight the key nodes and time series dependencies; then, the BiLSTM module is used to process the time series data and capture the dependencies between the previous and next sequences; then, the convolution operation is used to extract local features, and multiple convolution kernels are used to extract local information at different scales, and the nonlinear expression is increased by combining ReLU activation, which is recorded as Conv; finally, the LSTM and graph features are combined. The convolutional features are stacked for GNN to perform global information fusion, and the relationship between the features is calculated using Cross-Attention as the graph node edge, denoted as Ultimately providing a richer feature representation; 7. According to the multi-level graph multimodal sentiment analysis method based on meta-information according to claim 1, the step 6 specifically comprises: The multi-level feature maps and modality-invariant features of each modality are then input into the graph knowledge distillation GD module, which specifically includes: the graph distillation unit GD-Units is used to semantically align features and enhance the semantic relationship between modality-invariant features; in order to effectively filter redundant information and improve the alignment effect, the modality-specific GD unit Spec-GD is used to process the multi-level feature maps, and for the modality-invariant features, the modality-invariant GD unit Inv-GD is used to promote the complementarity between multi-modal features; for the multi-level feature maps, each modality map is first complemented with the other two modality maps through Cross-Attention, and then the aligned features are input into GD-Units; for the modality-invariant features, they are directly input into GD-Units for further alignment.
8. According to the multi-level graph multimodal sentiment analysis method based on meta-information according to claim 1, the step 7 specifically comprises: After flattening, the multi-level feature graphs of each modality are reduced in dimension through a private encoder to reduce feature redundancy and high-dimensional computational burden, effectively extract core information and compress unimportant features; For modality-invariant features, a weighted fusion strategy is used to combine the complementary information of each modality, and then the dimensionality reduction is unified through a shared encoder to ensure that the representation scales between different modalities are consistent, significantly reduce the computational complexity, retain key information, and improve the efficiency of multimodal feature fusion and the generalization ability of the model; c=c t +c v +c a (17) h com =E com (c) (19) 9. According to the multi-level graph multimodal sentiment analysis method based on meta-information according to claim 1, the step 8 specifically comprises: The reduced features are stacked to form a feature matrix The multi-head attention mechanism is used to process the Transformer network to enhance cross-modal and cross-subspace feature interactions. The multi-head attention mechanism enables each feature to effectively perceive the complementary information in other modal features, thereby improving the overall modeling ability of sentiment orientation and obtaining the feature matrix Finally, the processed feature vectors are concatenated into a joint representation and input into the fusion module G for final prediction; the fusion module further integrates and refines the features through linear transformation, nonlinear activation and regularization steps to generate the final sentiment prediction result; 10. A multi-level graph multi-modal sentiment analysis device based on meta-information, characterized in that: It includes a memory and one or more processors, wherein the memory stores executable code, and when the one or more processors execute the executable code, it is used to implement a multi-level graph multimodal sentiment analysis method based on meta-information as described in any one of claims 1-9.
Citation Information
Cited By
Audio and video dual-mode emotion recognition method and system based on adapter fusion
CN120411863A
Multi-modal feature splicing method and multi-modal data processing method based on rotation position coding technology
CN120724380A
Multi-modal interview automatic quality analysis and evaluation method and system based on large model
CN120849791A
Mongolian multi-modal sentiment analysis method based on dual-state space and multi-path transpose attention
CN121524732A
Adaptive dual-path evidence fusion method and system for multi-modal emotion recognition, and medium
CN122065269A