Multi-modal big data-oriented interpretable safe longitudinal federal representation learning method and device

The multimodal representation learning method based on cross-modal transfer and information bottleneck loss function optimization solves the problems of scarcity and poor interpretability of single modality data, improves the performance and interpretability of target modality representation, and is suitable for secure longitudinal federated representation learning of multimodal big data.

CN120688585APending Publication Date: 2025-09-23GUANGXI POWER GRID CORP
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510785291.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-12
Publication Date
2025-09-23

AI Technical Summary

Technical Problem

Existing multimodal representation learning methods are prone to training overfitting when single-modal data is scarce, and the extracted features contain a lot of redundant information, resulting in poor interpretability, which limits their application, especially in high-risk areas.

Method used

An interpretable and secure longitudinal federated representation learning method for multimodal big data is adopted. Through cross-modal migration, information from other modalities is transferred to the target modality to generate cross-modal multimodal representations. The model is optimized through the information bottleneck loss function and variational approximation method to improve the interpretability and generalization ability of the representation.

Benefits of technology

While maintaining generalization capabilities, it protects the inherent characteristics of the target domain, improves the performance and interpretability of the target modality representation, solves the problem of scarcity of single modality data, and enhances the interpretability of representation transfer learning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120688585A_ABST
    Figure CN120688585A_ABST
Patent Text Reader

Abstract

The invention provides an interpretable safe longitudinal federal representation learning method and device for multi-modal big data. According to the method, for the problem of target domain modal data scarcity, in the target domain modal representation learning process, an attention mechanism is used for supplementing information of a source domain modal into a target domain modal, and an algorithm framework is established under a longitudinal federated learning framework, so that on one hand, it is guaranteed that data of a source domain and data of a target domain are not locally output, and the algorithm framework is established under a longitudinal federated learning framework; therefore, the data security is improved; on the other hand, the problem of insufficient modal data of the target domain is solved, and the performance of downstream tasks is improved; besides, in the process of constructing the loss function, an information bottleneck theory is introduced, redundant information between input and intermediate representation is minimized, and related information between representation and a target task is maximized, so that efficient information compression and feature extraction are realized, and the interpretability of extracted representation for downstream tasks is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the intersection of data security and multimodal learning, and specifically relates to an explainable and secure longitudinal federated representation learning method and device for multimodal big data. Background Art

[0002] In recent years, as data-driven machine learning models and deep learning models have penetrated various fields and achieved remarkable results, machine learning algorithms have also placed increasingly higher demands on data representation extraction. Rationally and systematically utilizing data and obtaining good representations to improve the performance of downstream tasks has become an inevitable trend. Existing representation learning methods mainly extract representations from data of the same modality, ignoring the auxiliary information gain that other modalities can bring to the target modality data. This has led to limitations in the application of representation learning in many practical scenarios. Especially in real-world scenarios where target domain data is scarce, existing unimodal representation learning methods may lead to training overfitting, while some existing multimodal representation learning algorithms focus on finding common representations for each modality, while ignoring the extraction of personalized representations of the target modality that can be improved through cross-modal transfer. For such problems, cross-modal knowledge transfer is a good solution.

[0003] In actual scenarios, cross-modal transfer requires joint training with data from all participants, but due to user privacy restrictions, users cannot directly share original data. In addition, in actual scenarios, data from different participants may be heterogeneous in feature space and label space. These problems may cause the performance of representations extracted by existing representation learning methods to be greatly reduced in downstream tasks. Summary of the Invention

[0004] Existing representation learning methods for multimodal representation learning face two challenges. First, faced with the scarcity of single-modality data, existing representation learning approaches for multimodal data focus on extracting common representations across all modalities, while neglecting the inherent characteristics of each modality. This proposed approach focuses on transferring information from other modalities to the target modality to facilitate the extraction of target modality representations. This approach effectively preserves the inherent characteristics of the target domain while maintaining generalization capabilities, a problem that existing representation learning methods fail to address. Second, the features extracted by existing representation learning methods often contain redundant information from the original data, while the effective information is not sufficiently condensed, resulting in poor interpretability. In high-risk domains such as healthcare, interpretability is crucial for the applicability of the extracted representations to downstream tasks. Therefore, improving the interpretability of representation learning is a pressing issue. While addressing data scarcity through modality transfer, the interpretability of representation transfer learning should be maximized to overcome the feasibility limitations of its underlying algorithms.

[0005] In order to solve the problem of scarcity of data in the same modality during representation learning and poor interpretability of existing representation learning methods, this paper proposes an interpretable and secure longitudinal federated representation learning method and device for multimodal big data. This method solves the problem of scarcity of modal data in the target domain.

[0006] In order to achieve the above object, the present invention adopts the following technical solutions:

[0007] A method for learning explainable and secure longitudinal federated representations for multimodal big data, including the following steps:

[0008] Identify the federation participants, designate one of them as the coordination server, the remaining federation participants as clients, and set the federated learning hyperparameters.

[0009] Each client uses a local feature extractor to process the raw data from several source domain modalities and the raw data from the target domain modalities to generate a unimodal representation of each modality;

[0010] Interactively modeling unimodal representations through a preset fusion method to generate cross-modal multimodal representations;

[0011] Design an information bottleneck loss function based on multimodal representation and optimize the information bottleneck objective function through variational approximation method to update the federated model;

[0012] When the preset termination condition is reached, the federated representation model is output, and the final multimodal fusion network and interpretability report are published through the coordination server, while each client retains the local unimodal representation.

[0013] As a preferred solution, the federation participants are determined, one of the federation participants is used as a coordination server, the remaining federation participants are used as clients, and the federation learning hyperparameters are set, specifically including:

[0014] Initialize and pre-process the privacy data of federated participants to determine their definitions and parameter configurations; each participant holds a private multimodal dataset;

[0015] Randomly designate one of the participants as the coordination server and the rest of the federated participants as clients;

[0016] Set federated learning hyperparameters, including the Bloom filter's hash function set, Bloom filter capacity limit, false positive rate threshold, and homomorphic encryption algorithm parameters.

[0017] As a preferred solution, after setting the hyperparameters of federated learning, the method further includes:

[0018] The coordination server generates public and private keys that satisfy additive homomorphism and broadcasts the public key to all clients, so that each client uses the public key to encrypt and pre-process local private data.

[0019] As a preferred solution, each client uses a feature extractor locally to process the raw data from several source domain modalities and the raw data from the target domain modalities to generate a unimodal representation of each modality, specifically including:

[0020] Each client uses a Transformer-based feature extractor locally to process the raw data from several source domain modalities and the raw data from the target domain modality to generate a unimodal representation of each modality;

[0021] According to the unimodal representation of each modality, a multimodal unimodal learning network is constructed; the target domain hidden layer weight parameters are dynamically recorded during the training process and transferred to the source domain modal feature extraction network through the cross-modal transfer learning mechanism.

[0022] As a preferred solution, after generating the unimodal representation of each modality, the method further includes:

[0023] The target domain client additionally records the hidden layer weights and migrates the hidden layer weights to the source domain client through a secure migration protocol.

[0024] As a preferred solution, interactive modeling of unimodal representations by a preset fusion method to generate cross-modal multimodal representations specifically includes:

[0025] The client maps the unimodal representation to a Bloom filter, whereby the client uploads the encrypted Bloom filter to the coordination server;

[0026] The coordination server decrypts the received Bloom filter and dynamically selects a preset fusion method based on task requirements, thereby interactively modeling the unimodal representation and generating a cross-modal multimodal representation.

[0027] As a preferred solution, the information bottleneck loss function is designed based on the multimodal representation, and the information bottleneck objective function is optimized by the variational approximation method to update the federated model, which specifically includes:

[0028] Based on the cross-modal multimodal representation, we adopt variational approximation optimization, sample the multimodal representation through reparameterization techniques, and jointly optimize the cross entropy loss and KL divergence constraints to design the information bottleneck loss function;

[0029] The coordination server calculates the gradient update of the global fusion network parameters and the unimodal network parameters, masks the gradients, and distributes the encrypted gradients to each client through a secure aggregation protocol, so that the client can decrypt and update the local federated model.

[0030] As a preferred solution, when the preset termination condition is reached, the federated representation model is output, and the final multimodal fusion network and interpretability report are released through the coordination server. At the same time, each client retains the local unimodal representation, specifically including:

[0031] Based on the information bottleneck strength coefficient and the gradient contribution of the fusion path, the importance score of each modality is calculated and an interpretable report is generated;

[0032] When the global loss function converges or reaches the maximum number of iterations, the federated training is terminated, and the federated representation model is output, so that the coordination server publishes the final multimodal fusion network and interpretability report, and each client retains the local unimodal network.

[0033] Accordingly, the present invention also provides an interpretable and secure longitudinal federated representation learning device for multimodal big data, comprising:

[0034] The determination module is used to determine the federation participants, set one of the federation participants as the coordination server, the remaining federation participants as clients, and set the hyperparameters of federated learning;

[0035] The extraction module is used by each client to use a local feature extractor to process the raw data from several source domain modalities and the raw data from the target domain modality to generate a unimodal representation of each modality;

[0036] A modeling module, which is used to interactively model unimodal representations through a preset fusion method to generate cross-modal multimodal representations;

[0037] The update module is used to design the information bottleneck loss function based on the multimodal representation and optimize the information bottleneck objective function through the variational approximation method to update the federated model;

[0038] The output module is used to output the federated representation model when the preset termination condition is reached, and publish the final multimodal fusion network and interpretability report through the coordination server, while each client retains the local unimodal representation.

[0039] Accordingly, the present invention also provides a computer-readable storage medium, which includes a stored computer program, wherein when the computer program is running, the device where the computer-readable storage medium is located is controlled to execute the interpretable and secure longitudinal federated representation learning method for multimodal big data as described in any one of the above items.

[0040] Compared with the prior art, the present invention has at least the following beneficial technical effects:

[0041] This paper proposes an interpretable and secure longitudinal federated representation learning method for multimodal big data. By transferring relevant knowledge from other modalities to the target modality, this method assists in representation extraction and improves the representation performance of the target modality. Unlike multimodal representation learning, which focuses on extracting common representations across all modalities, this method focuses on transferring information from other modalities to the target modality to assist in training the target modality. The extracted representation of the target modality is both class-discriminative and domain-invariant, maintaining generalization capabilities while preserving the inherent characteristics of the target domain without sacrificing information inherent in the task's target domain. This addresses the data scarcity issue faced by existing representation learning algorithms for single-modality datasets. Furthermore, this paper combines the representation transfer algorithm with information bottleneck theory and core techniques in decoupled representation learning and interpretability. This deeply integrates knowledge from information theory and statistics with representation learning, enhancing the interpretability of representation transfer learning and providing theoretical support for an interpretable and secure longitudinal federated representation learning method for multimodal big data. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the specific embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0043] Figure 1 The figure is a flow chart of the steps of the method of the present invention.

[0044] Figure 2 The figure shows the implementation results of the method of the present invention on an actual data set.

[0045] Figure 3 This is a structural block diagram of an explainable and secure longitudinal federated representation learning device for multimodal big data. DETAILED DESCRIPTION

[0046] Hereinafter, only certain exemplary embodiments are briefly described. As will be appreciated by those skilled in the art, the described embodiments may be modified in various ways without departing from the spirit or scope of the present invention. Therefore, the drawings and description are to be considered as illustrative in nature and not restrictive.

[0047] It will be understood that when used in this specification and the appended claims, the terms “comprises” and “comprising” indicate the presence of described features, integers, steps, operations, elements and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof.

[0048] It should also be understood that the terms used in the present specification are only for the purpose of describing particular embodiments and are not intended to limit the present invention. As used in the present specification and the appended claims, the singular forms "a", "an", and "the" are intended to include the plural forms unless the context clearly indicates otherwise.

[0049] It should be further understood that the term "and / or" used in the present description and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.

[0050] The accompanying drawings illustrate various schematic diagrams of structures according to embodiments disclosed herein. These figures are not drawn to scale; for clarity, some details are exaggerated and some details may be omitted. The shapes of the various regions and layers shown in the figures, as well as their relative sizes and positional relationships, are merely exemplary and may deviate in practice due to manufacturing tolerances or technical limitations. Those skilled in the art may design regions / layers with different shapes, sizes, and relative positions as needed.

[0051] The embodiments of the present invention are described in detail below with reference to the accompanying drawings.

[0052] Example 1

[0053] See also Figure 1 , which is an explainable and secure longitudinal federated representation learning method for multimodal big data, including the following steps S1-S5:

[0054] S1: Determine the federation participants, use one of them as the coordination server, and the rest as clients, and set the federated learning hyperparameters.

[0055] As a preferred solution, the federation participants are determined, one of the federation participants is used as a coordination server, the remaining federation participants are used as clients, and the federation learning hyperparameters are set, specifically including:

[0056] Initialize and pre-process the privacy data of federated participants to determine their definitions and parameter configurations; each participant holds a private multimodal dataset;

[0057] Randomly designate one of the participants as the coordination server and the rest of the federated participants as clients;

[0058] Set federated learning hyperparameters, including the Bloom filter's hash function set, Bloom filter capacity limit, false positive rate threshold, and homomorphic encryption algorithm parameters.

[0059] As a preferred solution, after setting the hyperparameters of federated learning, the method further includes:

[0060] The coordination server generates public and private keys that satisfy additive homomorphism and broadcasts the public key to all clients, so that each client uses the public key to encrypt and pre-process local private data.

[0061] In this embodiment, federated participants are initialized and privacy data is preprocessed: participant definition and parameter configuration are performed, multiple parties are defined, each party holds a private multimodal data set (such as text, image, audio), and one of the parties is designated as the coordination server (Server), and the rest are clients (Clients).

[0062] In this embodiment, the federated learning hyperparameters are set, including but not limited to the hash function set H = {h1, h2, ..., h k}, the upper limit m of the Bloom filter capacity, the false positive rate threshold ∈, and the homomorphic encryption algorithm parameters (such as the key length λ of Paillier encryption, where Paillier encryption is an asymmetric encryption algorithm).

[0063] In this embodiment, homomorphic key generation and distribution, the coordination server generates a public key p that satisfies additive homomorphism. k and private key s k , and p k Broadcast to all clients. Each client uses p k Encryption pre-processes local private data to ensure that the data is irreversible and cannot be inferred during transmission and calculation.

[0064] S2: Each client uses a feature extractor locally to process the raw data from several source domain modalities and the raw data from the target domain modalities to generate a unimodal representation of each modality.

[0065] As a preferred solution, each client uses a feature extractor locally to process the raw data from several source domain modalities and the raw data from the target domain modalities to generate a unimodal representation of each modality, specifically including:

[0066] Each client uses a Transformer-based feature extractor locally to process the raw data from several source domain modalities and the raw data from the target domain modality to generate a unimodal representation of each modality;

[0067] According to the unimodal representation of each modality, a multimodal unimodal learning network is constructed; the target domain hidden layer weight parameters are dynamically recorded during the training process and transferred to the source domain modal feature extraction network through the cross-modal transfer learning mechanism.

[0068] As a preferred solution, after generating the unimodal representation of each modality, the method further includes:

[0069] The target domain client additionally records the hidden layer weights and migrates the hidden layer weights to the source domain client through a secure migration protocol.

[0070] In this embodiment, in the unimodal learning network, the feature extractors of the source domain and the target domain both adopt the Transformer architecture, where the text modality uses the BERT model (pre-trained language model), the output dimension is mapped to the shared feature space through the temporal convolution layer, and the weight transfer mechanism is used to inject the target domain hidden layer parameters into the source domain network.

[0071] Specifically, privacy-preserving unimodal feature extraction and unimodal network localization training are performed. Each client locally processes the raw data using a Transformer-based feature extractor (e.g., BERT for text modality and ViT for image modality) to generate unimodal representations: where θ i are local model parameters.

[0072] S3: Interactively model the unimodal representations through a preset fusion method to generate cross-modal multimodal representations.

[0073] As a preferred solution, interactive modeling of unimodal representations by a preset fusion method to generate cross-modal multimodal representations specifically includes:

[0074] The client maps the unimodal representation to a Bloom filter, whereby the client uploads the encrypted Bloom filter to the coordination server;

[0075] The coordination server decrypts the received Bloom filter and dynamically selects a preset fusion method based on task requirements, thereby interactively modeling the unimodal representation and generating a cross-modal multimodal representation.

[0076] In this embodiment, the target domain client additionally records the hidden layer weight W t , and W is transferred through a secure migration protocol (such as differential privacy noise injection) t Migrate to the source domain client to achieve cross-modal parameter alignment.

[0077] In this embodiment, Bloom filter encoding and encryption transmission: the client represents the single peak Z i Mapping to a Bloom filter:

[0078]

[0079] The client will be encrypted Bloom filter CBF i Upload to the coordination server.

[0080] Secure multimodal fusion and information bottleneck regularization and multimodal fusion method selection, coordinate the server to decrypt the received {CBF i}, dynamically select the fusion strategy according to the task requirements: Additive fusion: Tensor outer product fusion: Graph fusion: Construct an interaction graph G = (V, E), with node v i Represents a unimodal feature, edge weight e ij Calculated through the attention mechanism, Z is finally generated through graph convolution fuse .

[0081] In this embodiment, additive fusion generates multimodal representations by element-by-element addition; multiplicative fusion realizes feature interaction by element-by-element multiplication; splicing fusion maps the spliced ​​high-dimensional features to the target dimension through a fully connected layer; tensor outer product fusion generates a high-order interaction tensor by calculating the outer product of the unimodal representation; the graph fusion network models multimodal interactions as graph nodes and captures unimodal, bimodal and multimodal dynamic associations through a message passing mechanism.

[0082] S4: Design the information bottleneck loss function based on the multimodal representation, and optimize the information bottleneck objective function through the variational approximation method to perform federated model update.

[0083] As a preferred solution, the information bottleneck loss function is designed based on the multimodal representation, and the information bottleneck objective function is optimized by the variational approximation method to update the federated model, which specifically includes:

[0084] Based on the cross-modal multimodal representation, we adopt variational approximation optimization, sample the multimodal representation through reparameterization techniques, and jointly optimize the cross entropy loss and KL divergence constraints to design the information bottleneck loss function;

[0085] The coordination server calculates the gradient update of the global fusion network parameters and the unimodal network parameters, masks the gradients, and distributes the encrypted gradients to each client through a secure aggregation protocol, so that the client can decrypt and update the local federated model.

[0086] In this embodiment, a joint distribution decomposition hypothesis is defined to constrain the multimodal representation to depend only on the input data and be conditionally independent of the label. The mutual information boundary is approximated by the variational distribution, and maximizing the label mutual information is transformed into minimizing the cross entropy loss. Information compression is achieved by enforcing the KL divergence constraint between the multimodal representation and the standard Gaussian distribution.

[0087] In this embodiment, a vertical federated learning framework is used to achieve cross-domain privacy protection. Each participant only shares the multimodal representation regularized by the information bottleneck, and the original data and unimodal network parameters are locally protected through differential privacy or homomorphic encryption technology.

[0088] In this embodiment, in the variational approximation optimization, it is assumed that the target distribution obeys the Gaussian distribution, and the gradient back propagation is realized by the reparameterization technique; combined with the Monte Carlo sampling to estimate the KL divergence term, a random optimization algorithm is used to jointly train the unimodal and fusion networks.

[0089] Specifically, the information bottleneck-driven joint optimization designs the information bottleneck loss function:

[0090]

[0091] Using variational approximation optimization: sampling Z via reparameterization techniques fuse ~N(μ,σ 2 ), and jointly optimize the cross entropy loss and KL divergence constraint.

[0092] Then the federated model is updated and the interpretability analysis is performed, and the security parameters are aggregated. The coordination server calculates the global fusion network parameter θ g and the unimodal network parameters {θ i The gradient update Δθ of} is masked using Paillier homomorphic encryption and distributed to each client via the secure aggregation protocol. The client decrypts and updates the local model:

[0093] S5: When the preset termination condition is reached, the federated representation model is output, and the final multimodal fusion network and interpretability report are released through the coordination server, while each client retains the local unimodal representation.

[0094] As a preferred solution, when the preset termination condition is reached, the federated representation model is output, and the final multimodal fusion network and interpretability report are released through the coordination server. At the same time, each client retains the local unimodal representation, specifically including:

[0095] Based on the information bottleneck strength coefficient and the gradient contribution of the fusion path, the importance score of each modality is calculated and an interpretable report is generated;

[0096] When the global loss function converges or reaches the maximum number of iterations, the federated training is terminated, and the federated representation model is output, so that the coordination server publishes the final multimodal fusion network and interpretability report, and each client retains the local unimodal network.

[0097] In this embodiment, the method of the present invention supports multimodal interpretability analysis, and by calculating the information bottleneck intensity coefficient of each fusion path, quantifying the contribution of different modalities to the final decision, generates an explanatory report based on information flow.

[0098] In this embodiment, interpretability is quantified and output. Based on the information bottleneck strength coefficient β and the gradient contribution of the fusion path, the importance score of each modality is calculated. Generate explainable reports and visualize the rationale for multimodal decisions.

[0099] Termination conditions and result output, federated training terminates. When the global loss function converges Or reach the maximum number of iterations T max Terminate training.

[0100] Output the federated representation model and coordinate the server to publish the final multimodal fusion network and interpretability reports, each client retains a local unimodal network

[0101] Example 2

[0102] like Figure 2 As shown, the present invention provides an interpretable and secure longitudinal federated representation learning method for multimodal big data: MTIB. Its overall architecture, at a macro level, consists of multiple unimodal learning networks, a multimodal fusion network, and an encoder-decoder. The optimization of these networks is driven by the information bottleneck principle. In terms of the model's specific flow, MTIB includes the following steps S101-S106:

[0103] S101: Constructing Unimodal Learning Networks: F i m 、F t

[0104] First, each source domain mode M i (i=1,2,...,n) data X i (i=1,2,...,n) and target domain modality M t The original data Y are input into the corresponding feature extractor F i m and F tDuring training, the weight parameters of the target domain hidden layer are recorded for transfer learning. The original data in the target domain, such as text modality, is represented unimodally using the BERT-Transformer model as its feature extractor. Specifically, the steps of the BERT unimodal learning network are as follows:

[0105]

[0106]

[0107] Among them, Conv1D represents the time convolution operation, K i is the kernel size, used to map the output dimension of BERT to the shared dimension. t It is the weight parameter of the hidden layer in the target domain, which is used to transfer the weight parameters extracted from the target domain to the source domain, thereby achieving the goal of cross-modal representation transfer learning.

[0108] For feature extractors of other general modal raw data, we generally use the Transformer model

[66] as its feature extractor to obtain a unimodal representation, which is basically the same as above:

[0109]

[0110] It is worth noting that in the choice of feature extractor, we need to use the Transformer model to obtain unimodal representation.

[0111] S102: Establish multimodal fusion network: F f

[0112] Our algorithm is independent of the specific fusion mechanism. We can inject various fusion methods into our multimodal fusion structure to provide higher expression capabilities. This paper mainly studies five fusion methods to verify the effectiveness of the algorithm. The fusion methods are described as follows:

[0113] 1. Direct addition: The equation is expressed as follows:

[0114] U j =U1+U2+…+U n +U t ,

[0115] Among them U j ∈R d It is a multi-peak representation.

[0116] This fusion method is not learnable. However, in our experiments, we show that even with such a simple fusion method, our algorithm can still achieve very competitive performance.

[0117] 2. Multiplication: The equation is expressed as follows:

[0118] U j =U1·U2……U n ·U t ,

[0119] Multiplication is another non-parametric fusion method.

[0120] 3. Splicing: The equation is expressed as follows:

[0121]

[0122] Then, we use a fully connected network to map U j to the feature dimension of d. Together with direct addition and multiplication, they serve as the baseline fusion method throughout multimodal learning research.

[0123] 4. Tensor Fusion: Tensor fusion is a widely used feature fusion algorithm that has recently attracted extensive research attention. By applying outer products to explore the interactions between unimodal representations, the resulting multimodal representation has the highest expressive power, but is also high-dimensional and complex. The equation for tensor fusion is shown below:

[0124] U′ i =[U i ,1],i∈(1,2,...,n),

[0125]

[0126] in represents the outer product of a set of tensors. We then use a fully connected network to map U j To the feature dimension of d. 6, each unimodal embedding is padded with 1 to preserve the interaction of any subset of modalities.

[0127] 5. Graph Fusion Network: Graph fusion treats each interaction as a node and implements message passing between nodes to model unimodal, bimodal, and trimodal dynamics. The final graph representation is obtained by averaging the node embeddings. For more details, please refer to the graph fusion network in [1]. In our MTIB, the fused multimodal representation is regularized using the IB principle, making it sufficient for predicting labels while incorporating as little noise information from all three modalities as possible. MTIB regularization will be demonstrated in the next section.

[0128] S103: Establishing loss function based on information bottleneck theory

[0129] In this section, we will introduce the derivation of the MTIB loss function in detail:

[0130] The objective function of MTIB is defined as:

[0131] L MTIB =I(y;z)-βI(y,U j ),

[0132] To optimize the objective function of the MTIB, we employ the solution introduced in the Variable Information Bottleneck (VIB). First, we describe how to optimize the first term of the MTIB, I(y; z). Based on the formulas derived from the information bottleneck theory, we obtain:

[0133]

[0134] Then:

[0135]

[0136] However, this goal is difficult to achieve. Therefore, we define q(z|y) as a variational approximation of p(z|y), usually assuming a Gaussian distribution. Applying the property that the KL-divergence of the two distributions is greater than or equal to zero, we can obtain I(y;z):

[0137]

[0138] Combining the above mentioned equations, we can get:

[0139] I(y;z)≥∫dydzp(y,z)logq(z|y)-∫dzp(z)logp(z)

[0140] =∫dU j dydzp(y|U j )p(z|U j )p(U j )logq(z|y),

[0141] The entropy of the target label, H(z) = -∫dzp(z)logp(z), is irrelevant to parameter optimization and can be ignored. In this way, we can instead optimize I(y;z) by maximizing the lower bound of the objective function.

[0142] According to the equation, we will now optimize the second term of MTIB (i.e., the minimum information constraint). We can write the minimum information constraint as:

[0143]

[0144] We assume that q(z) is a variational approximation of the marginal distribution p(z), which is usually fixed to a standard normal Gaussian distribution. Similarly, applying the property that the KL-divergence of two distributions is greater than or equal to zero, we can get I(U j ,y), as shown in the following formula:

[0145]

[0146] Combining the above conditions, we can get:

[0147]

[0148] Combining the above two constraints and applying the equations mentioned above, we can obtain the lower bound of the objective loss function of MTIB:

[0149]

[0150] In the above formula, J MTIB It's L MTIB A lower bound for . By maximizing J MTIB , L MTIB The lower limit is improved.

[0151] S104: Encoder-Decoder Architecture

[0152] In the central server, the intermediate representation containing redundant information generated by the multimodal fusion network is input into the encoder to obtain the hidden representation, and then the hidden representation is output through the decoder and corresponds to the label. Through the network architecture combined with the multimodal information bottleneck loss function, the correspondence between the hidden representation and the label is maximized, and the correspondence between the hidden representation and the intermediate representation is minimized. The hidden representation y j It is the optimal multimodal representation adapted to downstream tasks.

[0153]

[0154]

[0155] Table 1

[0156] Table 1 shows the experimental results of the MTIB model and comparison methods on the CMU-MOSI dataset. As can be seen from the table, our MTIB model significantly outperforms the baseline method and other comparison methods across all five evaluation metrics. In terms of optimal performance, the MTIB architecture combined with the tensor fusion network achieves the best performance across several metrics, which is largely due to the algorithmic advantages of the tensor fusion network itself.

[0157] Specifically, when comparing the MTIB architecture combined with the tensor fusion network with a baseline method, the accuracy of the seven-category emotion classification (Acc-7) metric improved by 40.4% compared to the baseline method; the accuracy of the two-category emotion classification (Acc-2) metric improved by 12.8% compared to the baseline method; the F1-score metric for bidirectional sentiment classification improved by 14.6% compared to the baseline method; the mean absolute error (MAE) metric reduced the absolute error by 22.5% compared to the baseline method; and the Pearson correlation coefficient metric improved the correlation by 24% compared to the baseline method. Overall, the average classification accuracy improved by 22.6% compared to the baseline method.

[0158] Comparing the MTIB architecture with a tensor fusion network against the state-of-the-art MISA method, the accuracy of the seven-category emotion classification Acc-7 metric improved by 15.2% over the baseline method; the accuracy of the two-category emotion classification Acc-2 metric improved by 4.4% over the baseline method; the F1-score metric for bidirectional sentiment classification improved by 5.1% over the baseline method; the mean absolute error (MAE) metric reduced the absolute error by 9% over the baseline method; and the Pearson correlation coefficient metric improved the correlation by 2.6% over the baseline method. Overall, the MTIB architecture combined with the tensor fusion network architecture achieved an average performance improvement of 8.23% over the state-of-the-art MISA method in terms of classification accuracy.

[0159] Example 3

[0160] See also Figure 3 , which is an explainable and secure longitudinal federated representation learning device for multimodal big data provided by the present invention, comprising:

[0161] Determination module 201, used to determine federation participants, and set one of the federation participants as a coordination server, the remaining federation participants as clients, and set federated learning hyperparameters;

[0162] Extraction module 202, configured for each client to locally use a feature extractor to process the raw data from the multiple source domain modalities and the raw data from the target domain modalities to generate a unimodal representation of each modality;

[0163] A modeling module 203 is configured to interactively model the unimodal representation using a preset fusion method to generate a cross-modal multimodal representation;

[0164] An updating module 204 is configured to design an information bottleneck loss function based on the multimodal representation and optimize the information bottleneck objective function using a variational approximation method, thereby updating the federated model.

[0165] The output module 205 is used to output the federated representation model when the preset termination condition is reached, and publish the final multimodal fusion network and interpretability report through the coordination server, while each client retains the local unimodal representation.

[0166] As a preferred solution, the federation participants are determined, one of the federation participants is used as a coordination server, the remaining federation participants are used as clients, and the federation learning hyperparameters are set, specifically including:

[0167] Initialize and pre-process the privacy data of federated participants to determine their definitions and parameter configurations; each participant holds a private multimodal dataset;

[0168] Randomly designate one of the participants as the coordination server and the rest of the federated participants as clients;

[0169] Set federated learning hyperparameters, including the Bloom filter's hash function set, Bloom filter capacity limit, false positive rate threshold, and homomorphic encryption algorithm parameters.

[0170] As a preferred solution, after setting the hyperparameters of federated learning, the method further includes:

[0171] The coordination server generates public and private keys that satisfy additive homomorphism and broadcasts the public key to all clients, so that each client uses the public key to encrypt and pre-process local private data.

[0172] As a preferred solution, each client uses a feature extractor locally to process the raw data from several source domain modalities and the raw data from the target domain modalities to generate a unimodal representation of each modality, specifically including:

[0173] Each client uses a Transformer-based feature extractor locally to process the raw data from several source domain modalities and the raw data from the target domain modality to generate a unimodal representation of each modality;

[0174] According to the unimodal representation of each modality, a multimodal unimodal learning network is constructed; the target domain hidden layer weight parameters are dynamically recorded during the training process and transferred to the source domain modal feature extraction network through the cross-modal transfer learning mechanism.

[0175] As a preferred solution, after generating the unimodal representation of each modality, the method further includes:

[0176] The target domain client additionally records the hidden layer weights and migrates the hidden layer weights to the source domain client through a secure migration protocol.

[0177] As a preferred solution, interactive modeling of unimodal representations by a preset fusion method to generate cross-modal multimodal representations specifically includes:

[0178] The client maps the unimodal representation to a Bloom filter, whereby the client uploads the encrypted Bloom filter to the coordination server;

[0179] The coordination server decrypts the received Bloom filter and dynamically selects a preset fusion method based on task requirements, thereby interactively modeling the unimodal representation and generating a cross-modal multimodal representation.

[0180] As a preferred solution, the information bottleneck loss function is designed based on the multimodal representation, and the information bottleneck objective function is optimized by the variational approximation method to update the federated model, which specifically includes:

[0181] Based on the cross-modal multimodal representation, we adopt variational approximation optimization, sample the multimodal representation through reparameterization techniques, and jointly optimize the cross entropy loss and KL divergence constraints to design the information bottleneck loss function;

[0182] The coordination server calculates the gradient update of the global fusion network parameters and the unimodal network parameters, masks the gradients, and distributes the encrypted gradients to each client through a secure aggregation protocol, so that the client can decrypt and update the local federated model.

[0183] As a preferred solution, when the preset termination condition is reached, the federated representation model is output, and the final multimodal fusion network and interpretability report are released through the coordination server. At the same time, each client retains the local unimodal representation, specifically including:

[0184] Based on the information bottleneck strength coefficient and the gradient contribution of the fusion path, the importance score of each modality is calculated and an interpretable report is generated;

[0185] When the global loss function converges or reaches the maximum number of iterations, the federated training is terminated, and the federated representation model is output, so that the coordination server publishes the final multimodal fusion network and interpretability report, and each client retains the local unimodal network.

[0186] The specific embodiments described above further illustrate the objectives, technical solutions, and beneficial effects of the present invention. It should be understood that the above descriptions are merely specific embodiments of the present invention and are not intended to limit the scope of protection of the present invention. In particular, it should be noted that any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included within the scope of protection of the present invention for those skilled in the art.

Claims

1. A method for learning explainable and secure longitudinal federated representations for multimodal big data, characterized by: The following steps are involved: Identify the federation participants, designate one of them as the coordination server, the remaining federation participants as clients, and set the federated learning hyperparameters. Each client uses a local feature extractor to process the raw data from several source domain modalities and the raw data from the target domain modalities to generate a unimodal representation of each modality; Interactively modeling unimodal representations through a preset fusion method to generate cross-modal multimodal representations; Design an information bottleneck loss function based on multimodal representation and optimize the information bottleneck objective function through variational approximation method to update the federated model; When the preset termination condition is reached, the federated representation model is output, and the final multimodal fusion network and interpretability report are published through the coordination server, while each client retains the local unimodal representation.

2. The method for learning explainable and secure longitudinal federated representations for multimodal big data according to claim 1, wherein: The federation participants are determined, one of them is used as the coordination server, the rest of the federation participants are used as clients, and the hyperparameters of federated learning are set, including: Initialize and pre-process the privacy data of federated participants to determine their definitions and parameter configurations; each participant holds a private multimodal dataset; Randomly designate one of the participants as the coordination server and the rest of the federated participants as clients; Set federated learning hyperparameters, including the Bloom filter's hash function set, Bloom filter capacity limit, false positive rate threshold, and homomorphic encryption algorithm parameters.

3. The method for learning explainable and secure longitudinal federated representations for multimodal big data according to claim 2, wherein: After setting the hyperparameters for federated learning, the following steps are also included: The coordination server generates public and private keys that satisfy additive homomorphism and broadcasts the public key to all clients, so that each client uses the public key to encrypt and pre-process local private data.

4. The method for learning explainable and secure longitudinal federated representations for multimodal big data according to claim 1, wherein: Each client uses a feature extractor locally to process the raw data from several source domain modalities and the raw data from the target domain modalities to generate a unimodal representation of each modality, specifically including: Each client uses a Transformer-based feature extractor locally to process the raw data from several source domain modalities and the raw data from the target domain modality to generate a unimodal representation of each modality; According to the unimodal representation of each modality, a multimodal unimodal learning network is constructed; the target domain hidden layer weight parameters are dynamically recorded during the training process and transferred to the source domain modal feature extraction network through the cross-modal transfer learning mechanism.

5. The method for learning explainable and secure longitudinal federated representations for multimodal big data according to claim 4, wherein: After generating the unimodal representation of each modality, the method further includes: The target domain client additionally records the hidden layer weights and migrates the hidden layer weights to the source domain client through a secure migration protocol.

6. The method for learning explainable and secure longitudinal federated representations for multimodal big data according to claim 1, wherein: The interactive modeling of the unimodal representation by a preset fusion method to generate a cross-modal multimodal representation specifically includes: The client maps the unimodal representation to a Bloom filter, whereby the client uploads the encrypted Bloom filter to the coordination server; The coordination server decrypts the received Bloom filter and dynamically selects a preset fusion method based on task requirements, thereby interactively modeling the unimodal representation and generating a cross-modal multimodal representation.

7. The method for learning explainable and secure longitudinal federated representations for multimodal big data according to claim 6, wherein: The information bottleneck loss function is designed based on the multimodal representation, and the information bottleneck objective function is optimized by the variational approximation method to update the federated model, specifically including: Based on the cross-modal multimodal representation, we adopt variational approximation optimization, sample the multimodal representation through reparameterization techniques, and jointly optimize the cross entropy loss and KL divergence constraints to design the information bottleneck loss function; The coordination server calculates the gradient update of the global fusion network parameters and the unimodal network parameters, masks the gradients, and distributes the encrypted gradients to each client through a secure aggregation protocol, so that the client can decrypt and update the local federated model.

8. The method for learning explainable and secure longitudinal federated representations for multimodal big data according to claim 7, wherein: When the preset termination condition is reached, the federated representation model is output, and the final multimodal fusion network and interpretability report are published through the coordination server. At the same time, each client retains the local unimodal representation, specifically including: Based on the information bottleneck strength coefficient and the gradient contribution of the fusion path, the importance score of each modality is calculated and an interpretable report is generated; When the global loss function converges or reaches the maximum number of iterations, the federated training is terminated, and the federated representation model is output, so that the coordination server publishes the final multimodal fusion network and interpretability report, and each client retains the local unimodal network.

9. An interpretable and secure longitudinal federated representation learning device for multimodal big data, characterized by: include: The determination module is used to determine the federation participants, set one of the federation participants as the coordination server, the remaining federation participants as clients, and set the hyperparameters of federated learning; The extraction module is used by each client to use a local feature extractor to process the raw data from several source domain modalities and the raw data from the target domain modality to generate a unimodal representation of each modality; A modeling module, which is used to interactively model unimodal representations through a preset fusion method to generate cross-modal multimodal representations; The update module is used to design the information bottleneck loss function based on the multimodal representation and optimize the information bottleneck objective function through the variational approximation method to update the federated model; The output module is used to output the federated representation model when the preset termination condition is reached, and publish the final multimodal fusion network and interpretability report through the coordination server, while each client retains the local unimodal representation.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium includes a stored computer program, wherein, when the computer program is running, the device where the computer-readable storage medium is located is controlled to execute the steps of the explainable and secure longitudinal federated representation learning method for multimodal big data as described in any one of claims 1 to 8.

Citation Information

Cited By

  • Mobile crowdsourcing trust evaluation method and system

    CN121563322A