Dynamic graph multimodal sentiment analysis method based on learnable spline network

CN122778069APending Publication Date: 2026-09-18GUANGXI NORMAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610894617.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-22
Publication Date
2026-09-18

AI Technical Summary

Technical Problem

[0004]本发明根据现有技术中模态异质性、模态噪声污染和模态不平衡导致的情感分析精度不足、泛化能力差的问题,而提供一种基于可学习样条网络的动态图多模态情感分析方法

Benefits of technology

[0070] This technical solution relies on the radial basis function Kolmogorov-Arnold network encoder, replaces the traditional fixed activation function with a learnable B-spline activation function, and flexibly models the nonlinear mapping relationship of heterogeneous modes based on the Kolmogorov-Arnold representation theorem. Combined with self-supervised reconstruction loss to ensure the integrity of encoded information, it effectively improves cross-modal alignment capability and sentiment analysis accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122778069A_ABST
    Figure CN122778069A_ABST
Patent Text Reader

Abstract

The application discloses a kind of dynamic graph multimodal sentiment analysis methods based on learnable spline network, comprising the following steps: S1 multi-modal original feature extraction and radial basis Kolmogorov-Arnold network coding;S2 hierarchical multi-modal self-attention;S3 text-guided noise robust fusion;S4 local-global multi-modal fusion graph network;S5 cross-modal variational information bottleneck;S6 invariant learning regularization;S7 overall training of model;S8 sentiment prediction.This method realizes the unified projection of heterogeneous modal by learnable B-spline activation function, and improves the accuracy, robustness and cross-distribution generalization ability of multi-modal sentiment analysis by combining text-guided noise robust fusion and dynamic graph structured inference.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of artificial intelligence and multimodal data processing technology, specifically to a dynamic graph multimodal sentiment analysis method based on learnable spline networks. Background Technology

[0002] Multimodal sentiment analysis is a key technology in fields such as natural human-computer interaction and intelligent public opinion analysis. Currently, mainstream research approaches fall into four main categories: First, tensor fusion-based methods rely on tensor operations to achieve modality fusion, but depend on fixed activation functions, making it difficult to fit complex nonlinear relationships between heterogeneous modalities. Second, attention-based methods can align modal features, but are sensitive to acoustic and visual modal noise, which easily affects recognition accuracy as noise propagates with features. Third, graph neural network-based methods model modal interactions through graph structures, but use static graph topologies, failing to adaptively adjust the intensity of modal interactions between samples. Fourth, existing Kolmogorov-Arnold network methods leverage learnable spline activation functions to enhance nonlinear modeling capabilities, but lack modality denoising modules, resulting in incomplete dynamic interaction modeling and an inability to address modal imbalance issues. In summary, all existing approaches have significant technical shortcomings.

[0003] Further analysis reveals four common shortcomings in existing multimodal sentiment analysis technologies: First, significant modal heterogeneity, with textual, acoustic, and visual features belonging to different feature spaces and exhibiting large dimensional differences, resulting in poor cross-modal alignment using traditional linear mapping methods; Second, severe modal noise interference, with acoustic features susceptible to environmental noise and human voice interference, and visual features easily affected by lighting and facial pose, for which current technologies lack effective denoising methods; Third, modal information imbalance, with textual modal sentiment information accounting for the highest proportion, leading to an over-reliance on textual features in existing models and a weakening of the supplementary role of audiovisual modalities, resulting in a significant drop in model generalization ability when text becomes ineffective; Fourth, weak cross-distribution generalization ability, with models adapting to fixed training scenarios, leading to significant performance degradation when facing complex real-world environments involving multiple speakers and scenes. Summary of the Invention

[0004] This invention addresses the shortcomings of existing technologies, such as insufficient accuracy and poor generalization ability in sentiment analysis due to modal heterogeneity, modal noise contamination, and modal imbalance. It provides a dynamic graph multimodal sentiment analysis method based on learnable spline networks. This method achieves unified projection of heterogeneous modalities through a learnable B-spline activation function, and combines text-guided noise robust fusion and dynamic graph structured inference to improve the accuracy, robustness, and cross-distribution generalization ability of multimodal sentiment analysis.

[0005] The technical solution to achieve the objective of this invention is:

[0006] A dynamic graph multimodal sentiment analysis method based on learnable spline networks includes the following steps:

[0007] S1 Multimodal Original Feature Extraction and Radial Basis Kolmogorov-Arnold Network Encoding:

[0008] First, we acquire raw data from three modalities: text, acoustic, and visual, and then extract the corresponding raw features for each.

[0009] For each modality, a Radial Basis Function Kolmogorov-Arnold (RBKAM) network encoder is used to encode the features. The first layer of this encoder encodes the original features. The mapping is to 128-dimensional intermediate features, and the second layer outputs 32-dimensional uniform features. The third layer provides 32-dimensional unified features. Reconstruction is performed to obtain reconstruction features ;

[0010] S2-level multimodal self-attention:

[0011] Input uniform dimensional features ,Right now , , The 32-dimensional unified feature encoded by each modality is transformed into query, key, and value vectors through linear transformation and divided into multiple attention heads. First, intra-modal multi-head self-attention is performed on the query, key, and value vectors to mine semantic associations within a single modality. Then, cross-modal attention is performed to interactively compute the key and value vectors of the other two modalities using the query vector of the current modality.

[0012] Integrate intra-modal attention output and cross-modal attention output with unified dimensional features. Residual concatenation is performed, and feature refinement is completed by layer normalization, ultimately yielding optimized text, acoustic, and visual modal representations. ;

[0013] S3 text-guided noise-robust fusion:

[0014] S3.1. Sentiment-aware pooling: Introducing a learnable sentiment query vector and leveraging multi-head attention to optimize the text modality representation. Create an emotion-driven pooling mechanism to obtain... Cross-attention weighted output, and then with The mean pooling results along the temporal dimension are summed and fused, and then normalized to obtain text features that enhance sentiment semantics. ;

[0015] S3.2. Joint audiovisual modal denoising: Denoising of acoustic modal representations and visual modality representation First, a Fast Fourier Transform is performed, and a spectral threshold is set to filter out frequency domain noise. After inverse Fourier Transform and time-domain smoothing, the mixture is fused using an adaptive gating mechanism. Denoising features obtained by frequency domain denoising and time domain smoothing The gating coefficient is determined by a combination of characteristic statistics. Function generation:

[0016] in, Indicates the adaptive gating coefficient. This represents the Sigmoid activation function. This represents a multilayer perceptron. Denotes the mean function. This represents the audiovisual modal features obtained after frequency domain denoising and one-dimensional convolutional temporal smoothing;

[0017] Then output the denoised acoustic features. Visual features after noise reduction With noise reduction residual , ;

[0018] S3.3. Text-guided feature refinement: using optimized text features To query the acoustic features after denoising. and the visual features after noise reduction Perform cross-attention calculations to obtain text-acoustic fusion features. Text-visual fusion features Then, through gating network integration and Two sets of features are further aligned across modal semantics to output semantically aligned acoustic features. With visual features ;

[0019] S3.4. Uncertainty-Weighted Fusion: Input Text Features Noise reduction residual Text-acoustic fusion features Text-visual fusion features First, the features output in step S3.3 Text features The probability of sentiment prediction for each modality was calculated using linear sentiment head and Softmax respectively. Then calculate the cross-modal cosine similarity: And calculate the acoustic-text KL divergence and visual-text KL divergence to obtain and ,Will With noise reduction residual Cross-modal cosine similarity and acoustic text KL divergence Visual-Text KL Divergence The vectors are concatenated to form an uncertainty input vector, which is then used by a multilayer perceptron to estimate the uncertainty of each mode. Normalized fusion weights are calculated based on the inverse of uncertainty. : Complete modal weighted fusion:

[0020]

[0021] Add a feature-based MLP projection path and output Through gating network Adaptive combination and Two paths yield fused features:

[0022]

[0023] S3.5. Add hierarchical fusion auxiliary regression loss Constraint fusion features quality;

[0024] Then output the fused features. ;

[0025] S4 Local-Global Multimodal Fusion Graph Network:

[0026] Will respectively with , , The input features are obtained by weighting them at a ratio of 0.5:0.5. , , ,

[0027] These three types of fused input features are used as graph nodes to construct a two-layer graph structure;

[0028] Output global graph fusion features based on M-attention fusion and hypergraph convolution Modal-specific features , , With sentiment prediction logits;

[0029] S5 cross-modal variational information bottleneck:

[0030] right Variational coding is performed separately, and the mean and log-variance of the Gaussian distribution are output. Modal latent feature representations are obtained by sampling through reparameterization techniques. , , The joint latent representation is obtained by concatenating and encoding the latent features of the single modality. And combine sentiment tags to construct a prior distribution of label conditions;

[0031] By comprehensively incorporating multiple constraints, including KL divergence, modal variance constraints, cross-modal mutual information, and main task loss, a total cross-modal variational information bottleneck loss is constructed:

[0032]

[0033] This loss is used to filter out irrelevant noise information and retain effective features that are strongly correlated with the emotional task.

[0034] in, This represents the total loss due to the cross-modal variational information bottleneck. All of these represent hyperparameters. This represents the KL divergence between the unimodal posterior distribution and the standard Gaussian prior distribution. This represents the modal variance constraint loss. This represents the cross-modal mutual information loss. This represents the KL divergence between the joint posterior distribution and the standard Gaussian prior distribution. This represents the KL divergence between the unimodal posterior distribution and the label-conditional prior distribution. This indicates the loss in the main task of sentiment analysis;

[0035] S6 Invariant Learning Regularization:

[0036] The pseudo-environment is divided into three methods: label bucketing, sequence length bucketing, and fragment hashing, to simulate diverse application scenarios. The invariant risk minimization penalty term and the risk extrapolation penalty term are calculated separately: the invariant risk minimization penalty term constrains the model gradient consistency under different environments, and the risk extrapolation penalty term reduces the variance of task risk between different environments. The dual regularization improves the model's cross-scenario and cross-distribution generalization ability.

[0037] S7 Model Overall Training:

[0038] By integrating the main task regression loss, modality reconstruction loss, cross-modal variational information bottleneck loss, hierarchical fusion auxiliary loss, graph regularization loss, and IRM and REx regularization terms, the total loss function of the model is constructed:

[0039]

[0040] in, This represents the total loss function during the model training phase. This represents the mean squared error loss of the main task. This represents the total loss due to the cross-modal variational information bottleneck. This represents the auxiliary regression loss of the hierarchical fusion module. This represents the penalty term for minimizing constant risk. This represents the penalty term for minimizing the variance of environmental risk. All are weighted hyperparameters;

[0041] Training phase selection The optimizer distinguishes between setting the learning rate of the text encoder and other modules, configures weight decay, batch size, and number of iterations, and adopts a linear learning rate warm-up strategy.

[0042] S8 Sentiment Prediction:

[0043] Input the sentiment prediction logits from step S4, and map the logits to The interval is used to obtain the continuous emotion intensity prediction results. .

[0044] In step S1, raw data of three modalities—text, acoustic, and visual—are obtained, and corresponding raw features are extracted for each. The specific steps are as follows:

[0045] S1.1 Original Text Features Extraction was performed using a DeBERTav3 pre-trained model with a fixed dimension of L. 768, where L is the sequence length;

[0046] S1.2 Original Acoustic Features Extracted using tools such as COVAREP or OpenSMILE, with dimension L. 25 or L 74;

[0047] S1.3 Original Visual Features Extracted using OpenFace tool, with dimension L. 35. L 47 or L 177. All three types of features are in the form of equal-length sequences.

[0048] In step S1, the radial basis Kolmogorov-Arnold network encoder is divided into a three-layer structure:

[0049] The first layer of S1.4 is a Kolmogorov-Arnold network encoding layer, which encodes the original features. Mapped to 128-dimensional intermediate features;

[0050] The second layer in S1.5 is a linear projection layer, which sequentially performs layer normalization, ReLU activation, and random discarding on the intermediate features, outputting a 32-dimensional uniform feature. , ;

[0051] The third layer of S1.6 is an attention decoder, which reconstructs the 32-dimensional uniform features of the input to obtain the reconstructed features. ;

[0052] A self-supervised constraint mechanism is used to calculate the reconstruction loss to ensure the effectiveness of the encoding. The loss formula is as follows:

[0053] ,

[0054] in, Indicates the first The self-supervised reconstruction loss corresponding to the modality is This represents the square of the L2 norm.

[0055] In step S1, the encoder replaces the fixed scalar weights of the traditional network with learnable univariate spline functions, and the single-layer output expression is:

[0056]

[0057] In formula (1), This indicates the RadKAM encoder's first... The output value of each output neuron This represents the dimension of the input features of the RadKAM encoder. Indicates the RadKAM encoder's first... The input value of each input neuron. Indicates the connection of the first The input neuron and the first The learnable univariate spline function of each output neuron, composed of SiLU residual branches and... Step The branches are weighted and combined using spline basis functions. k represents the order of the B-spline basis function, which is a pre-set positive integer. The proportion of the two branches is adjusted by the learnable coefficients.

[0058] In step S4, the three types of fused input features are... The specific steps for constructing a two-layer graph structure using graph nodes are as follows:

[0059] S4.1 Feature projection will fuse the input features. After being reduced to 2D by mean pooling, it is then processed by linear projection, layer normalization, ReLU activation, and Dropout to be mapped to a unified form. Wei said Then stacked as 3D graph node tensor , Indicates the number of samples in the batch. A unified feature dimension representing graph nodes;

[0060] The S4.2 local graph module treats the three modal nodes of each sample as graph nodes, and constructs an adjacency matrix by calculating edge embeddings and connectivity scores from these graph nodes using MLP. After self-loop enhancement, the graph message passing is executed, the input is a bidirectional GRU for sequence modeling, and the output is an individual-level local embedding. ;

[0061] The S4.3 global graph module sequentially performs modality-specific graph attention extraction. Modal common graph convolution extracts shared features M-Attention Fusion Output Simultaneously, an adaptive graph learner is constructed, which incorporates the static adjacency matrix. With dynamic adjacency matrix Gated fusion is used to enhance effective connections by combining cosine similarity, where the static adjacency matrix... Generated from learnable 3rd order identity matrix parameters through Softmax normalization; dynamic adjacency matrix. The dynamic graph construction module uses graph node tensors Starting from this point, the system generates the following: cross-node attention scores are calculated using Q / K linear projection; the strongest connections are selected using Top-k sparsity filtering; and then Softmax normalization, matrix symmetry, and row normalization are applied.

[0062]

[0063]

[0064] For learnable gating coefficients, The learnable gating coefficients represent the fusion of static and dynamic adjacency matrices. The node cosine similarity matrix Element-wise multiplication is performed; a two-layer graph inference is executed via a GRU-style gated graph inference module, outputting a refined node representation. And update the modality-specific representation using residuals;

[0065] S4.4 Feature Aggregation: With three modality-specific representations , , Concatenate into a 5d dimensional aggregate vector: ,Will Input a three-layer MLP regression head and output sentiment prediction logits;

[0066] S4.5 Graph Structure Regularization Constraints: To constrain the rationality of graph structure, graph regularization loss is introduced, simultaneously achieving sparsity, uniform distribution, and structural symmetry constraints.

[0067]

[0068] in, This represents the total loss after regularization of the graph network structure. This represents the dynamically generated modal node adjacency matrix. , , All of these are manually set weight hyperparameters, which control the strength of sparsity constraints, uniform distribution constraints, and structural symmetry constraints, respectively. Let L1 norm represent the adjacency matrix A, used to constrain the sparsity of the graph. Representing the adjacency matrix The entropy is used to constrain the uniformity of node connectivity distribution. Representation matrix Rather than transpose The square of the Frobenius norm of the difference is used to constrain the symmetry of the graph structure.

[0069] Compared with the existing technology, the present technical solution has the following beneficial effects:

[0070] This technical solution relies on the radial basis function Kolmogorov-Arnold network encoder, replaces the traditional fixed activation function with a learnable B-spline activation function, and flexibly models the nonlinear mapping relationship of heterogeneous modes based on the Kolmogorov-Arnold representation theorem. Combined with self-supervised reconstruction loss to ensure the integrity of encoded information, it effectively improves cross-modal alignment capability and sentiment analysis accuracy.

[0071] This technical solution adds a text-guided spectral and temporal joint denoising mechanism, which combines frequency domain threshold filtering, temporal convolutional smoothing and adaptive gating fusion to effectively suppress audiovisual modal environmental interference noise, block noise propagation to back-end features, and improve the model's noise robustness.

[0072] This technical solution introduces a dynamic graph learning mechanism and an uncertainty-weighted fusion strategy, which can adaptively optimize the graph topology and modality fusion weights based on input samples. This overcomes the defects of traditional static modality interaction, avoids model dependence on a single modality, and achieves adaptive multimodal fusion. This invention integrates cross-modal variational information bottlenecks and a dual invariant learning regularization strategy. By mining task-related modality-invariant features and constraining the prediction risks of multiple pseudo-environments, it significantly improves the model's generalization ability across different data distribution scenarios. At the same time, this invention sets multiple auxiliary loss functions at all levels of encoding, fusion, graph modeling, and latent space representation, constructs hierarchical supervision signals, constrains the optimization direction of each module, and further improves the overall emotion recognition performance of the model. Attached Figure Description

[0073] Figure 1 This is a structural diagram of an embodiment. Detailed Implementation

[0074] The present invention will be further described below with reference to the accompanying drawings and embodiments, but this is not intended to limit the scope of the invention.

[0075] Example:

[0076] Reference Figure 1 A dynamic graph multimodal sentiment analysis method based on learnable spline networks includes the following steps:

[0077] S1 Multimodal Original Feature Extraction and Radial Basis Kolmogorov-Arnold Network Encoding:

[0078] First, we acquire raw data from three modalities: text, acoustic, and visual, and then extract the corresponding raw features for each.

[0079] S1.1 Original Text Features Extraction was performed using a DeBERTav3 pre-trained model with a fixed dimension of L. 768, where L is the sequence length;

[0080] S1.2 Original Acoustic Features Extracted using tools such as COVAREP or OpenSMILE, with dimension L. 25 or L 74;

[0081] S1.3 Original Visual Features Extracted using OpenFace tool, with dimension L. 35. L 47 or L 177, all three types of features are in the form of equal-length sequences;

[0082] For each modality, a Radial Basis Function Kolmogorov-Arnold (RBF-ANA) network encoder, also known as the RadKAM encoder, is used for feature encoding. This encoder replaces the fixed scalar weights of the traditional network with learnable univariate spline functions. The single-layer output expression is:

[0083]

[0084] In formula (1), This indicates the RadKAM encoder's first... The output value of each output neuron This represents the dimension of the input features of the RadKAM encoder. Indicates the RadKAM encoder's first... The input value of each input neuron. Indicates the connection of the first The input neuron and the first The learnable univariate spline function of each output neuron, composed of SiLU residual branches and... Step The branches are weighted and combined using spline basis functions, where k represents the order of the B-spline basis function and is a pre-defined positive integer. The proportion of the two branches is adjusted by learnable coefficients.

[0085] The radial-base Kolmogorov-Arnold network encoder is composed of a three-layer structure:

[0086] The first layer of S1.4 is a Kolmogorov-Arnold network encoding layer, which encodes the original features. Mapped to 128-dimensional intermediate features;

[0087] The second layer in S1.5 is a linear projection layer, which sequentially performs layer normalization, ReLU activation, and random discarding on the intermediate features, outputting a 32-dimensional uniform feature. , ;

[0088] The third layer of S1.6 is an attention decoder, which reconstructs the 32-dimensional uniform features of the input to obtain the reconstructed features. ;

[0089] A self-supervised constraint mechanism is used to calculate the reconstruction loss to ensure the effectiveness of the encoding. The loss formula is as follows:

[0090] ,

[0091] in, Indicates the first The self-supervised reconstruction loss corresponding to the modality is Represents the square of the L2 norm;

[0092] S2-level multimodal self-attention:

[0093] Input uniform dimensional features ,Right now , , The 32-dimensional unified feature encoded by each modality is transformed into query, key, and value vectors through linear transformation and divided into multiple attention heads. First, intra-modal multi-head self-attention is performed on the query, key, and value vectors to mine semantic associations within a single modality. Then, cross-modal attention is performed to interactively compute the key and value vectors of the other two modalities using the query vector of the current modality.

[0094] Integrate intra-modal attention output and cross-modal attention output with unified dimensional features. Residual concatenation is performed, and feature refinement is completed by layer normalization, ultimately yielding optimized text, acoustic, and visual modal representations. ;

[0095] S3 text-guided noise-robust fusion:

[0096] S3.1. Sentiment-aware pooling: Introducing a learnable sentiment query vector and leveraging multi-head attention to optimize the text modality representation. Create an emotion-driven pooling mechanism to obtain... Cross-attention weighted output, and then with The mean pooling results along the temporal dimension are summed and fused, and then normalized to obtain text features that enhance sentiment semantics. ;

[0097] S3.2. Joint audiovisual modal denoising: Denoising of acoustic modal representations and visual modality representation First, a Fast Fourier Transform is performed, and a spectral threshold is set to filter out frequency domain noise. After inverse Fourier Transform and time-domain smoothing, the mixture is fused using an adaptive gating mechanism. Denoising features obtained by frequency domain denoising and time domain smoothing The gating coefficient is determined by a combination of characteristic statistics. Function generation:

[0098] in, Indicates the adaptive gating coefficient. This represents the Sigmoid activation function. This represents a multilayer perceptron. Denotes the mean function. This represents the audiovisual modal features obtained after frequency domain denoising and one-dimensional convolutional temporal smoothing;

[0099] Then output the denoised acoustic features. Visual features after noise reduction With noise reduction residual , ;

[0100] S3.3. Text-guided feature refinement: using optimized text features To query the acoustic features after denoising. and the visual features after noise reduction Perform cross-attention calculations to obtain text-acoustic fusion features. Text-visual fusion features Then, through gating network integration and Two sets of features are further aligned across modal semantics to output semantically aligned acoustic features. With visual features ;

[0101] S3.4. Uncertainty-Weighted Fusion: Input Text Features Noise reduction residual Text-acoustic fusion features Text-visual fusion features First, the features output in step S3.3 Text features The probability of sentiment prediction for each modality was calculated using linear sentiment head and Softmax respectively. Then calculate the cross-modal cosine similarity: And calculate the acoustic-text KL divergence and visual-text KL divergence to obtain and ,Will With noise reduction residual Cross-modal cosine similarity and acoustic text KL divergence Visual-Text KL Divergence The vectors are concatenated to form an uncertainty input vector, which is then used by a multilayer perceptron to estimate the uncertainty of each mode. Normalized fusion weights are calculated based on the inverse of uncertainty. : Complete modal weighted fusion:

[0102]

[0103] Add a feature-based MLP projection path and output Through gating network Adaptive combination and Two paths yield fused features:

[0104]

[0105] S3.5. Add hierarchical fusion auxiliary regression loss Constraint fusion features quality;

[0106] Then output the fused features. ;

[0107] S4 Local-Global Multimodal Fusion Graph Network:

[0108] Will respectively with , , The input features are obtained by weighting them at a ratio of 0.5:0.5. , , ,

[0109] These three types of fused input features Used as graph nodes to construct a two-layer graph structure;

[0110] S4.1 Feature projection will fuse the input features. After being reduced to 2D by mean pooling, it is then processed by linear projection, layer normalization, ReLU activation, and Dropout to be mapped to a unified form. Wei said Then stacked as 3D graph node tensor , Indicates the number of samples in the batch. A unified feature dimension representing graph nodes;

[0111] The S4.2 local graph module treats the three modal nodes of each sample as graph nodes, and constructs an adjacency matrix by calculating edge embeddings and connectivity scores from these graph nodes using MLP. After self-loop enhancement, the graph message passing is executed, the input is a bidirectional GRU for sequence modeling, and the output is an individual-level local embedding. ;

[0112] The S4.3 global graph module sequentially performs modality-specific graph attention extraction. Modal common graph convolution extracts shared features M-Attention Fusion Output Simultaneously, an adaptive graph learner is constructed, which incorporates the static adjacency matrix. With dynamic adjacency matrix Gated fusion is used to enhance effective connections by combining cosine similarity, where the static adjacency matrix... Generated from learnable 3rd order identity matrix parameters through Softmax normalization; dynamic adjacency matrix. The dynamic graph construction module uses graph node tensors Starting from this point, the system generates the following: cross-node attention scores are calculated using Q / K linear projection; the strongest connections are selected using Top-k sparsity filtering; and then Softmax normalization, matrix symmetry, and row normalization are applied.

[0113]

[0114]

[0115] For learnable gating coefficients, The learnable gating coefficients represent the fusion of static and dynamic adjacency matrices. The node cosine similarity matrix Element-wise multiplication is performed; a two-layer graph inference is executed via a GRU-style gated graph inference module, outputting a refined node representation. And update the modality-specific representation using residuals;

[0116] S4.4 Feature Aggregation: With three modality-specific representations , , Concatenate into a 5d dimensional aggregate vector: ,Will Input a three-layer MLP regression head and output sentiment prediction logits;

[0117] S4.5 Graph Structure Regularization Constraints: To constrain the rationality of graph structure, graph regularization loss is introduced, simultaneously achieving sparsity, uniform distribution, and structural symmetry constraints.

[0118]

[0119] in, This represents the total loss after regularization of the graph network structure. This represents the dynamically generated modal node adjacency matrix. , , All of these are manually set weight hyperparameters, which control the strength of sparsity constraints, uniform distribution constraints, and structural symmetry constraints, respectively. Let L1 norm represent the adjacency matrix A, used to constrain the sparsity of the graph. Representing the adjacency matrix The entropy is used to constrain the uniformity of node connectivity distribution. Representation matrix Rather than transpose The square of the Frobenius norm of the difference is used to constrain the symmetry of the graph structure;

[0120] It outputs global graph fusion features based on M-attention fusion and hypergraph convolution. Modal-specific features , , With sentiment prediction logits;

[0121] S5 cross-modal variational information bottleneck:

[0122] right Variational coding is performed separately, and the mean and log-variance of the Gaussian distribution are output. Modal latent feature representations are obtained by sampling through reparameterization techniques. , , The joint latent representation is obtained by concatenating and encoding the latent features of the single modality. And combine sentiment tags to construct a prior distribution of label conditions;

[0123] By comprehensively incorporating multiple constraints, including KL divergence, modal variance constraints, cross-modal mutual information, and main task loss, a total cross-modal variational information bottleneck loss is constructed:

[0124]

[0125] This loss is used to filter out irrelevant noise information and retain effective features that are strongly correlated with the emotional task.

[0126] in, This represents the total loss due to the cross-modal variational information bottleneck. All of these represent hyperparameters. This represents the KL divergence between the unimodal posterior distribution and the standard Gaussian prior distribution. This represents the modal variance constraint loss. This represents the cross-modal mutual information loss. This represents the KL divergence between the joint posterior distribution and the standard Gaussian prior distribution. This represents the KL divergence between the unimodal posterior distribution and the label-conditional prior distribution. This indicates the loss in the main task of sentiment analysis;

[0127] S6 Invariant Learning Regularization:

[0128] The pseudo-environment is divided into three methods: label bucketing, sequence length bucketing, and fragment hashing, to simulate diverse application scenarios. The invariant risk minimization penalty term and the risk extrapolation penalty term are calculated separately: the invariant risk minimization penalty term constrains the model gradient consistency under different environments, and the risk extrapolation penalty term reduces the variance of task risk between different environments. The dual regularization improves the model's cross-scenario and cross-distribution generalization ability.

[0129] S7 Model Overall Training:

[0130] By integrating the main task regression loss, modality reconstruction loss, cross-modal variational information bottleneck loss, hierarchical fusion auxiliary loss, graph regularization loss, and IRM and REx regularization terms, the total loss function of the model is constructed:

[0131]

[0132] in, This represents the total loss function during the model training phase. This represents the mean squared error loss of the main task. This represents the total loss due to the cross-modal variational information bottleneck. This represents the auxiliary regression loss of the hierarchical fusion module. This represents the penalty term for minimizing constant risk. This represents the penalty term for minimizing the variance of environmental risk. All are weighted hyperparameters;

[0133] Training phase selection The optimizer distinguishes between setting the learning rate of the text encoder and other modules, configures weight decay, batch size, and number of iterations, and adopts a linear learning rate warm-up strategy.

[0134] S8 Sentiment Prediction:

[0135] Input the sentiment prediction logits from step S4, and map the logits to The interval is used to obtain the continuous emotion intensity prediction results. .

[0136] This example (KAG) is compared with existing state-of-the-art methods such as Graph-MFN, MFM, MMIM, HyCon, UniMSE, ConFEDE, MGCL, ULMD, MFON, ITHP, and KAN-MCP on two benchmark datasets, CMU-MOSI and CMU-MOSEI, as shown in Table 1.

[0137] On the CMU-MOSI dataset, this example achieves the following results: 7-class classification accuracy (Acc7) = 49.1%, 2-class classification accuracy (Acc2) = 88.9%, weighted F1 score = 88.9%, mean absolute error (MAE) = 0.605, and Pearson correlation coefficient (Corr) = 0.865. Both MAE and Corr are the best values ​​among all comparison methods. Compared to the strongest baseline, KAN-MCP, MAE is reduced by 0.010 (0.605 vs. 0.615), and Corr is improved by 0.008 (0.865 vs. 0.857). On the CMU-MOSEI dataset, this example achieves the following results: Acc7 = 54.8%, Acc2 = 86.9%, F1 score = 87.0%, MAE = 0.513, and Corr = 0.795. Among them, MAE and Corr were the best. Compared with KAN-MCP, MAE decreased by 0.009 (0.513 vs. 0.522), and Corr increased by 0.007 (0.795 vs. 0.788).

[0138] In summary, this example achieves or surpasses the performance levels of existing state-of-the-art methods in both MAE and Corr on the two datasets, validating the effectiveness of core modules such as the RadKAM encoder, text-guided noise robust fusion, and local-global dynamic graph network.

[0139] Table 1: Experimental results based on CMU-MOSI and CMU-MOSEI datasets in this example

[0140] .

[0141] The specific innovative points of this example are as follows:

[0142] 1. Radial-based Kolmogorov-Arnold network encoder

[0143] This example is the first to apply radial basis function Kolmogorov-Arnold (RBF-ANN) networks to the modality coding stage of multimodal sentiment analysis. Addressing the shortcomings of traditional multilayer perceptrons with fixed activation functions in terms of nonlinear fitting capabilities, the fixed scalar weights at the network edges are replaced with learnable B-spline univariate functions. The spline basis function parameterization is completed using a recursive formula. A dual-branch system with residual and spline branches is implemented, and a learnable scaling factor is introduced to adaptively adjust the weights of both branches, accurately fitting the complex nonlinear relationships between multiple heterogeneous modalities. Combined with a self-supervised reconstruction loss, encoding constraints are implemented, effectively ensuring the integrity of sentiment feature information during modality coding.

[0144] 2. Text-guided noise robust fusion module

[0145] This example proposes a hierarchical multimodal noise suppression and fusion mechanism. On the one hand, it combines frequency domain transformation, adaptive threshold filtering, and one-dimensional convolution to achieve joint spectral-temporal denoising, and uses an adaptive gating unit to complete the adaptive fusion of the original features and the denoised features. On the other hand, it uses the text modality with the lowest noise as a guide to complete cross-modal attention refinement of acoustic and visual modal features. At the same time, it introduces a modal uncertainty assessment mechanism to complete inverse uncertainty weighted fusion based on the reliability of the modality itself. From the three dimensions of feature denoising, semantic alignment, and weight allocation, it comprehensively suppresses the analysis interference caused by modal noise.

[0146] 3. Local-Global Two-Layer Multimodal Fusion Graph Network

[0147] This example constructs a local and global two-layer graph network architecture. The local graph network mines fine-grained modal interaction relationships, while the global graph network combines graph attention mechanisms, cross-modal feature projection, multi-head attention fusion, and hypergraph convolution to capture high-order modal association features. Simultaneously, a dynamic graph adaptive learning mechanism is designed to generate exclusive modal interaction edge weights in real time based on input samples. Effective fusion paths are selected through sparsity filtering, adaptively fusing dynamic and static graph structures, and strengthening modal associations based on node similarity. This achieves adaptive adjustment of modal interaction strength according to sample content, overcoming the limitations of traditional fixed modal interaction strategies.

[0148] 4. Dual-module joint invariant regularization strategy

[0149] This example integrates cross-modal variational information bottlenecks with a dual invariant learning regularization strategy. It leverages the cross-modal variational information bottleneck to mine task-related modal invariant features and constrains effective cross-modal information interaction. By constructing multi-dimensional pseudo-environments, it divides the training environment into different types and combines two types of invariant regularization constraints to unify the model gradient under different environments and reduce cross-environment prediction risk bias. This effectively addresses the technical pain point of weak cross-scene and cross-distribution generalization ability of existing multimodal sentiment analysis models.

Claims

1. A dynamic graph multimodal sentiment analysis method based on learnable spline networks, characterized in that, Includes the following steps: S1 Multimodal Original Feature Extraction and Radial Basis Kolmogorov-Arnold Network Encoding: First, we acquire raw data from three modalities: text, acoustic, and visual, and then extract the corresponding raw features for each. For each modality, a Radial Basis Function Kolmogorov-Arnold (RBKAM) network encoder is used to encode the features. The first layer of this encoder encodes the original features. The mapping is to 128-dimensional intermediate features, and the second layer outputs 32-dimensional uniform features. The third layer provides 32-dimensional unified features. Reconstruction is performed to obtain reconstruction features ; S2-level multimodal self-attention: Input uniform dimensional features ,Right now , , The 32-dimensional unified feature encoded by each modality is transformed into query, key, and value vectors through linear transformation and divided into multiple attention heads. First, intra-modal multi-head self-attention is performed on the query, key, and value vectors to mine semantic associations within a single modality. Then, cross-modal attention is performed to interactively compute the key and value vectors of the other two modalities using the query vector of the current modality. Integrate intra-modal attention output and cross-modal attention output with unified dimensional features. Residual concatenation is performed, and feature refinement is completed by layer normalization, ultimately yielding optimized text, acoustic, and visual modal representations. ; S3 text-guided noise-robust fusion: S3.

1. Sentiment-aware pooling: Introducing a learnable sentiment query vector and leveraging multi-head attention to optimize the text modality representation. Create an emotion-driven pooling mechanism to obtain... Cross-attention weighted output, and then with The mean pooling results along the temporal dimension are summed and fused, and then normalized to obtain text features that enhance sentiment semantics. ; S3.

2. Joint audiovisual modal denoising: Denoising of acoustic modal representations and visual modality representation First, a Fast Fourier Transform is performed, and a spectral threshold is set to filter out frequency domain noise. After inverse Fourier Transform and time-domain smoothing, the mixture is fused using an adaptive gating mechanism. Denoising features obtained by frequency domain denoising and time domain smoothing The gating coefficient is determined by a combination of characteristic statistics. Function generation: in, Indicates the adaptive gating coefficient. This represents the Sigmoid activation function. This represents a multilayer perceptron. Denotes the mean function. This represents the audiovisual modal features obtained after frequency domain denoising and one-dimensional convolutional temporal smoothing; Then output the denoised acoustic features. Visual features after noise reduction With noise reduction residual , ; S3.

3. Text-guided feature refinement: using optimized text features To query the acoustic features after denoising. and the visual features after noise reduction Perform cross-attention calculations to obtain text-acoustic fusion features. Text-visual fusion features Then, through gating network integration and Two sets of features are further aligned across modal semantics to output semantically aligned acoustic features. With visual features ; S3.

4. Uncertainty-Weighted Fusion: Input Text Features Noise reduction residual Text-acoustic fusion features Text-visual fusion features First, the features output in step S3.3 Text features The probability of sentiment prediction for each modality was calculated using linear sentiment head and Softmax respectively. Then calculate the cross-modal cosine similarity: And calculate the acoustic-text KL divergence and visual-text KL divergence to obtain and ,Will With noise reduction residual Cross-modal cosine similarity and acoustic text KL divergence Visual-Text KL Divergence The vectors are concatenated to form an uncertainty input vector, which is then used by a multilayer perceptron to estimate the uncertainty of each mode. Normalized fusion weights are calculated based on the inverse of uncertainty. : Complete modal weighted fusion: Add a feature-based MLP projection path and output Through gating network Adaptive combination and Two paths yield fused features: S3.

5. Add hierarchical fusion auxiliary regression loss Constraint fusion features quality; Then output the fused features. ; S4 Local-Global Multimodal Fusion Graph Network: Will respectively with , , The input features are obtained by weighting them at a ratio of 0.5:0.

5. , , , These three types of fused input features are used as graph nodes to construct a two-layer graph structure; Output global graph fusion features based on M-attention fusion and hypergraph convolution Modal-specific features , , With sentiment prediction logits; S5 cross-modal variational information bottleneck: right Variational coding is performed separately, and the mean and log-variance of the Gaussian distribution are output. Modal latent feature representations are obtained by sampling through reparameterization techniques. , , The joint latent representation is obtained by concatenating and encoding the latent features of the single modality. And combine sentiment tags to construct a prior distribution of label conditions; By comprehensively incorporating multiple constraints, including KL divergence, modal variance constraints, cross-modal mutual information, and main task loss, a total cross-modal variational information bottleneck loss is constructed: This loss is used to filter out irrelevant noise information and retain effective features that are strongly correlated with the emotional task. in, This represents the total loss due to the cross-modal variational information bottleneck. All of these represent hyperparameters. This represents the KL divergence between the unimodal posterior distribution and the standard Gaussian prior distribution. This represents the modal variance constraint loss. This represents the cross-modal mutual information loss. This represents the KL divergence between the joint posterior distribution and the standard Gaussian prior distribution. This represents the KL divergence between the unimodal posterior distribution and the label-conditional prior distribution. This indicates the loss in the main task of sentiment analysis; S6 Invariant Learning Regularization: The pseudo-environment is divided into three methods: label bucketing, sequence length bucketing, and fragment hashing, to simulate diverse application scenarios. The invariant risk minimization penalty term and the risk extrapolation penalty term are calculated separately: the invariant risk minimization penalty term constrains the model gradient consistency under different environments, and the risk extrapolation penalty term reduces the variance of task risk between different environments. The dual regularization improves the model's cross-scenario and cross-distribution generalization ability. S7 Model Overall Training: By integrating the main task regression loss, modality reconstruction loss, cross-modal variational information bottleneck loss, hierarchical fusion auxiliary loss, graph regularization loss, and IRM and REx regularization terms, the total loss function of the model is constructed: in, This represents the total loss function during the model training phase. This represents the mean squared error loss of the main task. This represents the total loss due to the cross-modal variational information bottleneck. This represents the auxiliary regression loss of the hierarchical fusion module. This represents the penalty term for minimizing constant risk. This represents the penalty term for minimizing the variance of environmental risk. All are weighted hyperparameters; Training phase selection The optimizer distinguishes between setting the learning rate of the text encoder and other modules, configures weight decay, batch size, and number of iterations, and adopts a linear learning rate warm-up strategy. S8 Sentiment Prediction: Input the sentiment prediction logits from step S4, and map the logits to The interval is used to obtain the continuous emotion intensity prediction results. .

2. The dynamic graph multimodal sentiment analysis method based on learnable spline networks according to claim 1, characterized in that, In step S1, raw data of three modalities—text, acoustic, and visual—are obtained, and corresponding raw features are extracted for each. The specific steps are as follows: S1.1 Original Text Features Extraction was performed using a DeBERTav3 pre-trained model with a fixed dimension of L. 768, where L is the sequence length; S1.2 Original Acoustic Features Extracted using tools such as COVAREP or OpenSMILE, with dimension L. 25 or L 74; S1.3 Original Visual Features Extracted using OpenFace tool, with dimension L.

35. L 47 or L 177. All three types of features are in the form of equal-length sequences.

3. The dynamic graph multimodal sentiment analysis method based on learnable spline networks according to claim 2, characterized in that, In step S1, the radial basis Kolmogorov-Arnold network encoder is divided into a three-layer structure: The first layer of S1.4 is a Kolmogorov-Arnold network encoding layer, which encodes the original features. Mapped to 128-dimensional intermediate features; The second layer in S1.5 is a linear projection layer, which sequentially performs layer normalization, ReLU activation, and random discarding on the intermediate features, outputting a 32-dimensional uniform feature. , ; The third layer of S1.6 is an attention decoder, which reconstructs the 32-dimensional uniform features of the input to obtain the reconstructed features. ; A self-supervised constraint mechanism is used to calculate the reconstruction loss to ensure the effectiveness of the encoding. The loss formula is as follows: , in, Indicates the first The self-supervised reconstruction loss corresponding to the modality is This represents the square of the L2 norm.

4. The dynamic graph multimodal sentiment analysis method based on learnable spline networks according to claim 1, characterized in that, In step S1, the encoder replaces the fixed scalar weights of the traditional network with learnable univariate spline functions, and the single-layer output expression is: In formula (1), This indicates the RadKAM encoder's first... The output value of each output neuron This represents the dimension of the input features of the RadKAM encoder. Indicates the RadKAM encoder's first... The input value of each input neuron. Indicates the connection of the first The input neuron and the first The learnable univariate spline function of each output neuron, composed of SiLU residual branches and... Step The branches are weighted and combined using spline basis functions. k represents the order of the B-spline basis function, which is a pre-set positive integer. The proportion of the two branches is adjusted by the learnable coefficients.

5. The dynamic graph multimodal sentiment analysis method based on learnable spline networks according to claim 1, characterized in that, In step S4, the three types of fused input features are... The specific steps for constructing a two-layer graph structure using graph nodes are as follows: S4.1 Feature projection will fuse the input features. After being reduced to 2D by mean pooling, it is then processed by linear projection, layer normalization, ReLU activation, and Dropout to be mapped to a unified form. Wei said Then stacked as 3D graph node tensor , Indicates the number of samples in the batch. A unified feature dimension representing graph nodes; The S4.2 local graph module treats the three modal nodes of each sample as graph nodes, and constructs an adjacency matrix by calculating edge embeddings and connectivity scores for these graph nodes using MLP. After self-loop enhancement, the graph message passing is executed, the input is a bidirectional GRU for sequence modeling, and the output is an individual-level local embedding. ; The S4.3 global graph module sequentially performs modality-specific graph attention extraction. Modal common graph convolution extracts shared features M-Attention Fusion Output Simultaneously, an adaptive graph learner is constructed, which incorporates the static adjacency matrix. With dynamic adjacency matrix Gated fusion is used to enhance effective connections by combining cosine similarity, where the static adjacency matrix... Generated from learnable 3rd order identity matrix parameters through Softmax normalization; dynamic adjacency matrix. The dynamic graph construction module uses graph node tensors Starting from this point, the system generates the following: cross-node attention scores are calculated using Q / K linear projection; the strongest connections are selected using Top-k sparsity filtering; and then Softmax normalization, matrix symmetry, and row normalization are applied. For learnable gating coefficients, The learnable gating coefficients represent the fusion of static and dynamic adjacency matrices. The node cosine similarity matrix This is element-wise multiplication; The GRU-style gated graph inference module performs two-layer graph inference and outputs a refined node representation. And update the modality-specific representation using residuals; S4.4 Feature Aggregation: With three modality-specific representations , , Concatenate into a 5d dimensional aggregate vector: ,Will Input a three-layer MLP regression head and output sentiment prediction logits; S4.5 Graph Structure Regularization Constraints: To constrain the rationality of graph structure, graph regularization loss is introduced, simultaneously achieving sparsity, uniform distribution, and structural symmetry constraints. in, This represents the total loss after regularization of the graph network structure. This represents the dynamically generated modal node adjacency matrix. , , All of these are manually set weight hyperparameters, which control the strength of sparsity constraints, uniform distribution constraints, and structural symmetry constraints, respectively. Let L1 norm represent the adjacency matrix A, used to constrain the sparsity of the graph. Representing the adjacency matrix The entropy is used to constrain the uniformity of node connectivity distribution. Representation matrix Rather than transpose The square of the Frobenius norm of the difference is used to constrain the symmetry of the graph structure.