Multi-modal sentiment analysis platform and method based on vertical federal learning
By adopting a multimodal sentiment analysis platform based on vertical federated learning in multimodal sentiment analysis, the problems of insufficient data privacy protection, lack of multimodal federated learning mechanisms and difficulty in feature alignment and fusion are solved, and efficient and accurate multimodal sentiment analysis is achieved.
Patent Information
- Application Number
- CN202510060391.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-15
- Publication Date
- 2025-05-30
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The prior art has problems such as insufficient data privacy protection, lack of multimodal federated learning mechanisms and difficulty in alignment and fusion of multimodal features in multimodal sentiment analysis.
Using a multimodal sentiment analysis platform based on vertical federated learning, multimodal features are divided into emotion-related features, modal independent features and interference features through feature decoupling modules, and these features are modeled and aligned with the distribution model, and combined with the joint training module to optimize model parameters.
It realizes multimodal sentiment analysis while protecting data privacy, improves the accuracy of feature representation, solves the problem of difficulty in multimodal data fusion, and improves model performance.
Smart Images

Figure CN120067568A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and specifically to a multi-modal sentiment analysis platform and method based on vertical federated learning Background Art
[0002] Multi-modal sentiment analysis aims to comprehensively process data from multiple modalities such as text, images, and audio to accurately identify the emotional state. However, traditional methods usually require all data to be centralized on one platform for unified processing. This centralized method has obvious deficiencies in terms of data privacy and security, especially in fields with high privacy requirements such as healthcare and finance. Therefore, how to achieve multi-modal sentiment analysis while protecting data privacy has become an urgent problem to be solved
[0003] Federated learning, as an emerging distributed machine learning method, allows each data holder to collaboratively train a model without sharing the original data. However, most existing federated learning methods are for single-modal data and lack an effective mechanism for processing multi-modal data. This results in the difficulty for existing federated learning methods to fully utilize the complementary information between modalities when facing multi-modal sentiment analysis tasks, and the model performance is limited
[0004] In addition, there is heterogeneity between multi-modal data, and data from different modalities have differences in feature space, distribution, etc. When traditional multi-modal learning methods align and fuse different modal features, they often need to share a large amount of intermediate data or model parameters, which may lead to the risk of privacy leakage in the framework of federated learning. Therefore, how to effectively align and fuse multi-modal features in federated learning while ensuring data privacy is still a huge challenge
[0005] In summary, the existing technologies mainly have the following deficiencies in multi-modal sentiment analysis: Insufficient data privacy protection: Traditional methods require centralized data processing, posing a risk of privacy leakage. Lack of multi-modal federated learning mechanism: Existing federated learning methods are mainly for single-modal and are difficult to handle multi-modal data. Difficulty in multi-modal feature alignment and fusion: The heterogeneity of different modal data increases the complexity of feature alignment and fusion and may cause privacy issues Summary of the Invention
[0006] Aiming at the deficiencies of the existing technologies, the present invention provides a multi-modal sentiment analysis platform and method based on vertical federated learning, which solves the problem that traditional methods require centralized data processing and pose a risk of privacy leakage
[0007] To achieve the above objectives, the present invention is realized through the following technical solutions: A multi-modal sentiment analysis platform based on vertical federated learning, including: The data holder module is configured to store multimodal data and extract sentiment-related features, modality-independent features, and interference features through a feature extraction tool; The federated learning coordination module is configured to coordinate feature alignment and model training among multiple data holders; The feature decoupling module is configured to divide multimodal features into sentiment-related features, modality-independent features, and interference features based on the subspace decomposition method; The distribution modeling module is configured to perform distribution modeling on the decoupled sentiment-related features and modality-independent features. The distribution modeling module fits the distribution of sentiment-related features based on the Gaussian mixture model and extracts the high-order statistical parameters of modality-independent features; The cross-modal alignment module is configured to align the distribution of sentiment-related features, and the modality-independent features achieve consistency by aligning high-order statistical characteristics; The joint training module is configured to update the model parameters according to the global optimization objective and distribute the updated model parameters to each data holder.
[0008] Preferably, the feature decoupling module includes: The feature decomposition unit is used to divide multimodal features into sentiment-related features, modality-independent features, and interference features based on the orthogonal projection method; The optimization unit is used to achieve orthogonal separation of the feature subspace by minimizing the decoupling loss function, and the decoupling loss function includes the inner product constraint of the subspace.
[0009] Preferably, the distribution modeling module includes: The distribution fitting unit is used to fit the distribution of sentiment-related features, and the fitting constructs the distribution based on the Gaussian mixture model; The parameter extraction unit is used to extract the mean, covariance, and skewness parameters from the modality-independent features as the distribution description of the modality-independent features.
[0010] Preferably, the cross-modal alignment module includes: The alignment calculation unit is used to calculate the distribution difference based on the distribution of sentiment-related features, and the distribution difference is calculated by the regularized Wasserstein distance; The mapping optimization unit is used to achieve distribution alignment through the optimization of the transport mapping matrix; The independent feature alignment unit is used to achieve consistency through the alignment of the mean, covariance, and skewness of the modality-independent features.
[0011] Preferably, the mapping optimization unit uses the Sinkhorn-Knopp algorithm to optimize the transport mapping matrix to reduce the computational complexity and achieve efficient alignment of sentiment-related features.
[0012] Preferably, the joint training module includes: An optimization objective generation unit for constructing a joint optimization objective by combining decoupling loss, alignment loss, and generalization loss; A parameter update unit for updating model parameters based on the joint optimization objective and generating a global model; A distribution unit for encrypting and distributing the optimized global model parameters to the data holders.
[0013] Preferably, the optimization objectives in the joint training module include: Decoupling loss for optimizing the orthogonality of the feature subspace; Alignment loss for minimizing the distribution difference between sentiment-related features and modality-independent features; Generalization loss for enhancing the adaptability of the model under distribution deviation scenarios.
[0014] Preferably, the data holder module is configured to interact with the federated learning coordination module after encrypting the distribution parameters through homomorphic encryption, and the federated learning coordination module calculates the feature distribution difference and optimization objective based on the encrypted data.
[0015] Preferably, the federated learning coordination module models the allowable set of distribution deviations based on a preset robust optimization mechanism, and realizes sentiment feature alignment under distribution deviation conditions by optimizing the generalization loss.
[0016] The present invention also provides a multi-modal sentiment analysis method based on vertical federated learning, including the following steps: Step 1, initialize the federated learning platform and configure the data holder module and the federated learning coordination module; Step 2, the data holder decomposes the multi-modal features into sentiment-related features, modality-independent features, and interference features through the feature decoupling module, and the feature decoupling is realized through the orthogonal projection algorithm; Step 3, the data holder performs Gaussian mixture distribution modeling on the sentiment-related features through the distribution modeling module, and extracts the mean, covariance, and skewness parameters of the modality-independent features; Step 4, the data holder sends the encrypted distribution parameters to the federated learning coordination module, and the federated learning coordination module calculates the distribution difference through the cross-modal alignment module and aligns the distribution of the sentiment-related features based on the regularized Wasserstein distance; Step 5, the federated learning coordination module realizes modality consistency by optimizing the high-order statistical parameters of the modality-independent features; Step 6, the federated learning coordination module optimizes the global model parameters through the joint training module, and the optimization objective combines decoupling loss, alignment loss, and generalization loss; Step 7: The federated learning coordination module distributes the updated global model parameters to the data holders, and the data holders use the global model to predict the local sentiment data; Step 8: Repeat Steps 2 to 7 to complete the next round of iterative optimization.
[0017] The present invention provides a multi-modal sentiment analysis platform and method based on vertical federated learning. It has the following beneficial effects: 1. By adopting a federated learning framework, the present invention coordinates the feature alignment and model training among multiple data holders, realizing multi-modal sentiment analysis while protecting data privacy. Compared with the existing technology that requires centralized data for training, it solves the problem of data privacy leakage.
[0018] 2. Through the feature decoupling module, the present invention uses the subspace decomposition method to divide multi-modal features into sentiment-related features, modality-independent features, and interference features, improving the accuracy of feature representation. Compared with the existing technology that fails to effectively distinguish different feature types, it overcomes the problem of model performance degradation caused by feature mixing.
[0019] 3. Through the cross-modal alignment module of the present invention, the distribution of sentiment-related features is aligned, and the modality-independent features are aligned by aligning the high-order statistical characteristics, realizing the consistency between different modalities. Compared with the existing technology lacking an effective alignment mechanism, it solves the problem of difficult multi-modal data fusion.
[0020] 4. Through the joint training module, the present invention updates the model parameters according to the global optimization objective and distributes the updated model parameters to each data holder, realizing the efficient training of the model. Compared with the existing technology of independent training, it overcomes the problem of poor model performance. Description of the Drawings
[0021] Figure 1 It is the main framework diagram of the present invention; Figure 2 It is the method flow chart of the present invention. Detailed Embodiments
[0022] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0023] Embodiment: Please refer to the attached Figure 1, the embodiment of the present invention provides a multi-modal sentiment analysis platform based on vertical federated learning, including: The data holder module is configured to store multi-modal data and extract sentiment-related features, modality-independent features, and interference features through a feature extraction tool; The federated learning coordination module is configured to coordinate feature alignment and model training among multiple data holders; The feature decoupling module is configured to divide multi-modal features into sentiment-related features, modality-independent features, and interference features based on the subspace decomposition method; The distribution modeling module is configured to perform distribution modeling on the decoupled sentiment-related features and modality-independent features. The distribution modeling module fits the distribution of sentiment-related features based on the Gaussian mixture model and extracts the high-order statistical parameters of modality-independent features; The cross-modal alignment module is configured to align the distribution of sentiment-related features, and the modality-independent features achieve consistency by aligning high-order statistical characteristics; The joint training module is configured to update the model parameters according to the global optimization objective and distribute the updated model parameters to each data holder.
[0024] The feature decoupling module includes: The feature decomposition unit is used to divide multi-modal features into sentiment-related features, modality-independent features, and interference features based on the orthogonal projection method; The optimization unit is used to achieve orthogonal separation of the feature subspace by minimizing the decoupling loss function, and the decoupling loss function includes the inner product constraint of the subspace.
[0025] The distribution modeling module includes: The distribution fitting unit is used to fit the distribution of sentiment-related features, and the fitting constructs a distribution based on the Gaussian mixture model; The parameter extraction unit is used to extract the mean, covariance, and skewness parameters from the modality-independent features as the distribution description of the modality-independent features.
[0026] The cross-modal alignment module includes: The alignment calculation unit is used to calculate the distribution difference based on the distribution of sentiment-related features, and the distribution difference is calculated by the regularized Wasserstein distance; The mapping optimization unit is used to achieve distribution alignment through the optimization of the transport mapping matrix; The independent feature alignment unit is used to achieve consistency through the alignment of the mean, covariance, and skewness of the modality-independent features.
[0027] The mapping optimization unit uses the Sinkhorn-Knopp algorithm to optimize the transport mapping matrix to reduce the computational complexity and achieve efficient alignment of sentiment-related features.
[0028] The joint training module includes: An optimization objective generation unit for constructing a joint optimization objective by combining decoupling loss, alignment loss, and generalization loss; A parameter update unit for updating model parameters based on the joint optimization objective and generating a global model; A distribution unit for encrypting and distributing the optimized global model parameters to data holders.
[0029] The optimization objectives in the joint training module include: Decoupling loss for optimizing the orthogonality of the feature subspace; Alignment loss for minimizing the distribution difference between sentiment-related features and modality-independent features; Generalization loss for enhancing the adaptability of the model under distribution deviation scenarios.
[0030] The data holder module is configured to interact with the federated learning coordination module after encrypting distribution parameters through homomorphic encryption. The federated learning coordination module calculates the feature distribution difference and optimization objective based on the encrypted data.
[0031] Based on a preset robust optimization mechanism, the federated learning coordination module models the allowable set of distribution deviations and achieves sentiment feature alignment under distribution deviation conditions by optimizing the generalization loss.
[0032] Please refer to the appendix Figure 2 The present invention also provides a multi-modal sentiment analysis method based on vertical federated learning, including the following steps: Step 1: Initialize the federated learning platform and configure the data holder module and the federated learning coordination module; Step 2: The data holder decomposes multi-modal features into sentiment-related features, modality-independent features, and interference features through a feature decoupling module. The feature decoupling is implemented by an orthogonal projection algorithm; Step 3: The data holder performs Gaussian mixture distribution modeling on the sentiment-related features through a distribution modeling module, and extracts mean, covariance, and skewness parameters for the modality-independent features; Step 4: The data holder sends the encrypted distribution parameters to the federated learning coordination module. The federated learning coordination module calculates the distribution difference through a cross-modal alignment module and aligns the distribution of sentiment-related features based on the regularized Wasserstein distance; Step 5: The federated learning coordination module achieves modality consistency by optimizing the high-order statistical parameters of the modality-independent features; Step 6: The federated learning coordination module optimizes the global model parameters through the joint training module. The optimization objective combines decoupling loss, alignment loss, and generalization loss; Step 7: The federated learning coordination module distributes the updated global model parameters to the data holders, and the data holders use the global model to predict the local sentiment data; Step 8: Repeat Steps 2 to 7 to complete the next round of iterative optimization.
[0033] In this embodiment, the data holder module is configured to store multimodal data and is responsible for initially extracting sentiment-related features, modality-independent features, and interference features.
[0034] Specifically, the data holder module includes the following technical content: First, the multimodal data is stored locally to maintain data integrity and privacy. For example, the voice data may be the collected user audio files, the text data may be the comment content generated by the user, and the image data may come from the facial expression images uploaded by the user. The data holder module does not need to share these raw data across parties, but completes the relevant feature extraction tasks through local processing.
[0035] In a possible implementation, the data holder module pre-configures feature extraction tools for performing specific feature extraction operations on different modality data. For example: For text data, a deep learning model (such as BERT) is usually used to extract context semantic features, and at the same time, the text is decomposed into sentiment-related word vector representations (such as sentiment word embeddings) and modality-independent syntactic structure information; For voice data, speech spectrum information is generally obtained through acoustic feature extraction tools (such as MFCC or RNN networks), where the pitch change of the speech may correspond to sentiment-related features, and the frequency stability can be used as a modality-independent feature; For image data, a convolutional neural network (such as ResNet) can be used to extract facial expression-related features, and at the same time, remove the background or other irrelevant regions, which are classified as interference features.
[0036] Generally, the results of feature extraction need to be further classified into three types of features, namely: Sentiment-related features : For example, the sentiment word embeddings of the text, the intonation change curve of the voice, and the facial expression key points of the image; Modality-independent features : For example, the syntactic structure of the text, the basic frequency stability of the voice, and the contour information of the image; Interference features : Such as background noise in the voice, meaningless stop words in the text, and non-face regions in the image.
[0037] In this embodiment, these three types of features are optimally partitioned through a decoupled loss function, and the specific optimization objectives are as follows:
[0038] Among them: represent emotion-related features; represent modality-independent features; represent interference features; Frobenius norm, used to measure the orthogonality between different feature subspaces As an option, the data holder module can also dynamically adjust the parameters of feature extraction to adapt to the quality of different modality data. For example, in some embodiments, if the sampling rate of voice data is low, the weight of low-frequency features can be enhanced to ensure that the extracted emotion-related features can more accurately reflect the user's emotional state.
[0039] In another possible implementation, the data holder module performs local distribution modeling on the extracted features. Specifically: For emotion-related features perform distribution fitting, usually using a Gaussian mixture model (GMM), and the fitting result is expressed as:
[0040] Among them: represents the mixture distribution of emotion-related features; is the number of Gaussian components; is the th component's weight; are the mean and covariance respectively.
[0041] For modality-independent features the data holder module calculates their higher-order statistical properties, including: mean ; covariance Cov skewness , where is the standard deviation.
[0042] As an implementation, the data holder module encrypts these modeling results and sends them to the federated learning coordination module. Generally, to protect data privacy, the encryption method can use homomorphic encryption or secure multi-party computation (MPC) technology. For example, the parameters of the feature distribution and They are encrypted and transmitted to avoid direct leakage of the original feature information.
[0043] In addition, the data holder module can also perform feature screening to further compress the amount of feature information to be transmitted. For example, in some embodiments, pruning can be performed on the distribution parameters fitted by GMM, and only the components with high weights are retained to reduce communication overhead.
[0044] The data holder module also has a certain degree of scalability. For example, in an extended scenario, the module can support other types of modal data, such as video data or physiological signal data. For video data, the module can extract emotion-related features frame by frame and use a temporal model (such as LSTM or Transformer) to capture the temporal correlation between features. For physiological signal data (such as heart rate or brain waves), the module can combine time-frequency analysis methods to extract features reflecting emotional fluctuations while removing physiological background noise.
[0045] In this embodiment, the federated learning coordination module first issues instructions for feature extraction and distribution modeling to each data holder according to the task requirements of the platform. Specifically, the coordination module will determine the required feature types and model parameters according to the preset learning objectives, and then send this information to each data holder. In some embodiments, the coordination module can also customize and issue different task instructions according to the computing power and data characteristics of each data holder to improve the overall efficiency.
[0046] After the data holders complete feature extraction and distribution modeling, the federated learning coordination module is responsible for collecting these feature parameters. To protect data privacy, all transmitted data is encrypted. Generally, the coordination module uses homomorphic encryption or secure multi-party computation (MPC) technology to ensure necessary calculations without decryption. As an option, the coordination module can also use differential privacy technology to further enhance data security.
[0047] In this embodiment, after receiving the encrypted feature parameters, the federated learning coordination module first calls the cross-modal alignment module for feature alignment. Specifically, the coordination module will calculate the feature distribution differences between each data holder and perform distribution alignment through an optimization algorithm (such as the Sinkhorn-Knopp algorithm). After feature alignment is completed, the coordination module is responsible for scheduling the joint training process. It will update the model parameters according to the global optimization objective, combining decoupling loss, alignment loss, and generalization loss. The optimization objective function can be expressed as:
[0048] where: Denotes the decoupling loss, which is used to optimize the orthogonality of the feature subspace; Denotes the alignment loss, which is used to minimize the distribution difference between modalities; Denotes the generalization loss, which is used to enhance the model's adaptability to distribution biases; 、 、 Are the corresponding weight parameters, which are used to balance each loss.
[0049] In a possible implementation, the coordination module uses the gradient descent algorithm to optimize the above loss function and update the global model parameters. The updated model parameters are encrypted and then distributed to each data holder. Each data holder uses these parameters to update the local model, thus completing one round of the federated learning process.
[0050] As an option, the federated learning coordination module can also have the following functions: Dynamic adjustment strategy: Dynamically adjust the task distribution and training scheduling strategies according to the real-time network conditions and data distribution to improve the learning efficiency.
[0051] Model evaluation and feedback: After each round of training, evaluate the model performance and send the feedback results to each data holder to guide the next round of training.
[0052] Anomaly detection: Monitor the behaviors of each data holder, detect and handle possible anomalies, such as data quality problems or malicious participants.
[0053] In this embodiment, the feature decoupling module first receives the multi-modal feature representations from the data holder module. These features may contain sentiment-related information, modality-specific information, and interference factors such as noise. To improve the accuracy of sentiment analysis, these features need to be decoupled.
[0054] Specifically, let the input feature be The feature decoupling module decomposes it into three parts:
[0055] Where: Denotes the sentiment-related feature; Denotes the modality-independent feature; Denotes the interference feature.
[0056] To ensure the effectiveness of the above decomposition, the feature decoupling module adopts the subspace decomposition method. Generally, it is required that these three feature subspaces are orthogonal to each other to reduce information redundancy.
[0057] In a possible implementation, the feature decoupling module uses the orthogonal projection algorithm to achieve feature decoupling. Specifically, by constructing a projection matrix, the input feature is projected onto each subspace.
[0058] Let: be the projection matrix for emotion-related features; be the projection matrix for modality-independent features; be the projection matrix for interference features.
[0059] Then there is:
[0060] To ensure the orthogonality of the subspaces, the projection matrix needs to satisfy the following conditions:
[0061] To optimize the above projection process, the feature decoupling module designs a decoupling loss function. This loss function aims to minimize the correlation between different feature subspaces.
[0062] The loss function is defined as:
[0063] Where: represents emotion-related features; represents modality-independent features; represents interference features; represents the Frobenius norm, which is used to measure the norm of a matrix. represents the correlation matrix between emotion-related features and modality-independent features.
[0064] By minimizing the independence between each feature subspace can be ensured, thus achieving effective feature decoupling.
[0065] As an option, the feature decoupling module can also combine other decoupling methods to improve the decoupling effect. For example, using an autoencoder structure to encode and decode the input features to learn a more refined feature representation.
[0066] In some embodiments, the feature decoupling module may adopt an adversarial training approach. By introducing a discriminator, it forces the decoupled features to have stronger discriminability in different subspaces.
[0067] In this embodiment, the distribution modeling module models the sentiment-related features output by the feature decoupling module Specifically, a Gaussian Mixture Model (GMM) is used to fit the distribution.
[0068] GMM is a commonly used probability model for representing distributions of arbitrary shapes. Its probability density function is expressed as:
[0069] where: represents the number of Gaussian components; is the weight of the -th Gaussian component, satisfying and represents a multivariate Gaussian distribution with mean and covariance , and its probability density function is:
[0070] where: represents the dimension of the feature ; represents the determinant of the covariance matrix ; represents minus the transposed vector of the mean .
[0071] In a possible implementation, the distribution modeling module uses the Expectation-Maximization (EM) algorithm to estimate the parameters of the GMM. The EM algorithm maximizes the log-likelihood function of the data under the model through iterative optimization to obtain the optimal parameter estimates.
[0072] For the mode-independent features output by the feature decoupling module the distribution modeling module extracts their high-order statistical parameters to capture the statistical characteristics of the data.
[0073] Specifically, the extracted high-order statistical parameters include: Mean vector: representing the central tendency of the feature , and the calculation formula is:
[0074] Wherein: represents the number of samples; represents the modal independent feature of the
[0075] Covariance matrix: represents the linear correlation of features The calculation formula is:
[0076] Skewness: represents the symmetry of the feature distribution, and the calculation formula is:
[0077] Wherein: represents the standard deviation of the feature; Kurtosis: represents the sharpness of the feature distribution, and the calculation formula is:
[0078] After completing the distribution modeling, the distribution modeling module transmits the model parameters to the federated learning coordination module. Generally, to protect data privacy, the parameters are encrypted before transmission. As an option, homomorphic encryption technology can be used to ensure that necessary calculations can still be performed in the encrypted state.
[0079] The distribution modeling module can also select other suitable probability distribution models for fitting according to the characteristics of the data. For example, for emotion-related features with multimodal characteristics, a mixture model can be used for modeling. For modal independent features with a long-tailed distribution, a stable distribution can be used for fitting.
[0080] In this embodiment, the cross-modal alignment module first receives the emotion-related feature distribution parameters from the distribution modeling module. Specifically, for each modality , the distribution of the emotion-related feature is modeled by a Gaussian mixture model (GMM) to obtain the mean vector and the covariance matrix .
[0081] To achieve alignment between different modalities, the cross-modal alignment module uses the Optimal Transport method to calculate the optimal transport mapping between the distributions of modalities. Specifically, the goal is to find a transport matrix such that from modality to modality Minimize the transmission cost.
[0082] The optimal transmission problem can be formulated as the following optimization problem:
[0083] Where: represents the transmission volume from the -th sample of modality to the -th sample of modality ; represents the transmission cost between sample and sample , usually expressed as the square of the Euclidean distance, i.e., .
[0084] By solving the above optimization problem, the optimal transmission matrix is obtained, and then the distribution alignment between modalities is achieved.
[0085] For modality-specific features , the cross-modal alignment module achieves consistency by aligning the high-order statistical features. Specifically, the mean vector , covariance matrix , skewness and kurtosis of each modality are calculated. Then, using the method of statistical matching, the high-order statistical features of different modalities are aligned. For example, by linear transformation, the mean and covariance of modality are adjusted to be consistent with those of modality . Specifically, the linear transformation can be expressed as:
[0086] Where: is the transformation matrix for adjusting the covariance; is the offset vector for adjusting the mean.
[0087] By selecting appropriate and , the transformed features are matched with the high-order statistical features of the target modality, thus achieving the alignment of modality-specific features.
[0088] As an alternative, the cross-modal alignment module can also use the method of adversarial learning for alignment. Specifically, a generative adversarial network (GAN) is introduced. The discriminator is used to evaluate the alignment degree of different modality features, and the generator is responsible for generating the aligned feature representations.
[0089] In addition, the cross-modal alignment module can also select other suitable alignment methods according to specific application requirements, such as kernel-based alignment methods or deep learning models, to improve the alignment effect.
[0090] In this embodiment, the joint training module first initializes the global model parameters , and distributes them to each data holder. Specifically, assume there are data holders, and each holder has a local dataset . The joint training module distributes the initial model parameters to each holder to ensure the synchronization of the training process.
[0091] After receiving the global model parameters, each data holder uses its private data for model training locally. Specifically, holder uses its local dataset and the received model parameters for training, and the updated local model parameters are denoted as . The goal of local training is to minimize the following loss function:
[0092] Where: represents the size of the dataset of holder ; represents the data sample and the corresponding label; represents the loss function, such as cross-entropy loss; represents the predicted output of the model.
[0093] After each holder completes local training, the joint training module collects the model updates of all holders and performs aggregation of the global model. Generally, a weighted average method is used for aggregation to update the global model parameters:
[0094] This method ensures that holders with a larger amount of data have a greater impact on the update of the global model.
[0095] In a possible implementation, the joint training module needs to handle the case where different holders use heterogeneous models. For this purpose, a model distillation method can be adopted, using the model outputs of each holder as soft labels to train a unified student model. Specifically, the model output of holder is , the joint training module performs weighted averaging on the outputs of all holders to obtain soft labels:
[0096] Where: is the weight of the holder satisfying represents the model output of the holder .
[0097] Then, use the soft label to train the student model and minimize the following loss function:
[0098] In this way, the integration and collaborative training of heterogeneous models are achieved.
[0099] The joint training module repeats the above processes of local training and global aggregation until the performance of the model on the validation set reaches the expectation or meets the preset convergence conditions. Generally, the maximum number of iterations or the threshold of performance improvement is set as the stopping condition.
[0100] As an option, the joint training module can also introduce a differential privacy mechanism to ensure the protection of the data privacy of the holders during the model update process. Specifically, noise is added when uploading the local model update to prevent the leakage of sensitive information.
[0101] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A multimodal sentiment analysis platform based on vertical federated learning, characterized by: include: A data holder module configured to store multimodal data and extract emotion-related features, modality-independent features, and interference features through a feature extraction tool; A federated learning coordination module configured to coordinate feature alignment and model training among multiple data holders; a feature decoupling module configured to divide the multimodal features into emotion-related features, modality-independent features, and interference features based on a subspace decomposition method; A distribution modeling module, configured to perform distribution modeling on the decoupled emotion-related features and modality-independent features, wherein the distribution modeling module fits the distribution of the emotion-related features based on a Gaussian mixture model and extracts high-order statistical parameters of the modality-independent features; A cross-modal alignment module, configured to align the distribution of sentiment-related features and modality-independent features to achieve consistency by aligning high-order statistical properties; The joint training module is configured to update the model parameters according to the global optimization goal and distribute the updated model parameters to each data holder.
2. The multimodal sentiment analysis platform based on vertical federated learning according to claim 1, characterized in that: The feature decoupling module comprises: A feature decomposition unit, used for dividing the multimodal features into emotion-related features, modality-independent features and interference features based on an orthogonal projection method; An optimization unit is used to achieve orthogonal separation of feature subspaces by minimizing a decoupling loss function, wherein the decoupling loss function includes a subspace inner product constraint.
3. The multimodal sentiment analysis platform based on vertical federated learning according to claim 1, characterized in that: The distribution modeling module includes: A distribution fitting unit, used for fitting the distribution of emotion-related features, wherein the fitting constructs the distribution based on a Gaussian mixture model; The parameter extraction unit is used to extract the mean, covariance and skewness parameters from the modal independent features as the distribution description of the modal independent features.
4. The multimodal sentiment analysis platform based on vertical federated learning according to claim 1, characterized in that: The cross-modal alignment module includes: an alignment calculation unit, configured to calculate a distribution difference based on the distribution of the emotion-related features, wherein the distribution difference is calculated by a regularized Wasserstein distance; A mapping optimization unit, for achieving distribution alignment through transmission mapping matrix optimization; Independent feature alignment unit, used to achieve consistency by aligning the mean, covariance, and skewness of modality-independent features.
5. The multimodal sentiment analysis platform based on vertical federated learning according to claim 4 is characterized in that: The mapping optimization unit optimizes the transmission mapping matrix using the Sinkhorn-Knopp algorithm to reduce computational complexity and achieve efficient alignment of emotion-related features.
6. The multimodal sentiment analysis platform based on vertical federated learning according to claim 1, characterized in that: The joint training module includes: The optimization target generation unit is used to construct a joint optimization target by combining the decoupling loss, alignment loss and generalization loss; A parameter updating unit, used to update the model parameters based on the joint optimization objective and generate a global model; The distribution unit is used to encrypt the optimized global model parameters and distribute them to the data holder.
7. The multimodal sentiment analysis platform based on vertical federated learning according to claim 1, characterized in that: The optimization objectives in the joint training module include: Decoupling loss, used to optimize the orthogonality of feature subspaces; Alignment loss, used to minimize the distribution difference between sentiment-related features and modality-independent features; Generalization loss is used to improve the model's adaptability in the case of distribution deviation.
8. The multimodal sentiment analysis platform based on vertical federated learning according to claim 1, characterized in that: The data holder module is configured to interact with the federated learning coordination module after encrypting the distribution parameters through homomorphic encryption. The federated learning coordination module calculates the feature distribution difference and optimization target based on the encrypted data.
9. The multimodal sentiment analysis platform based on vertical federated learning according to claim 1, characterized in that: The federated learning coordination module models the allowed distribution deviation set based on a preset robust optimization mechanism, and achieves sentiment feature alignment under distribution deviation conditions by optimizing the generalization loss.
10. A multimodal sentiment analysis method based on vertical federated learning, according to the multimodal sentiment analysis platform based on vertical federated learning according to any one of claims 1 to 9, characterized in that: The following steps are involved: Step 1: Initialize the federated learning platform and configure the data holder module and the federated learning coordination module; Step 2: The data holder decomposes the multimodal features into emotion-related features, modality-independent features, and interference features through the feature decoupling module. Feature decoupling is achieved through the orthogonal projection algorithm. Step 3: The data holder uses the distribution modeling module to perform Gaussian mixture distribution modeling on the emotion-related features, and extracts the mean, covariance, and skewness parameters of the modal independent features; Step 4: The data holder sends the encrypted distribution parameters to the federated learning coordination module. The federated learning coordination module calculates the distribution difference through the cross-modal alignment module and aligns the distribution of emotion-related features based on the regularized Wasserstein distance. Step 5: The federated learning coordination module achieves modal consistency by optimizing the high-order statistical parameters of modal independent features; Step 6: The federated learning coordination module optimizes the global model parameters through the joint training module, and the optimization objective combines the decoupling loss, the alignment loss and the generalization loss; Step 7: The federated learning coordination module distributes the updated global model parameters to the data holders, and the data holders use the global model to predict the local sentiment data; Step 8: Repeat steps 2 to 7 to complete the next round of iterative optimization.
Citation Information
Cited By
Feature de-entanglement emotion decoding system combining electroencephalogram and electrocardio
CN120296687A
Intelligent terminal multi-mode sentiment analysis method, device and server
CN120951171A
Intelligent terminal multi-modal sentiment analysis method, device and server
CN120951171B