Multi-modal data privacy protection distributed machine learning method based on federal learning
Through the federated learning method of multimodal feature alignment and lightweight privacy protection, the problems of multimodal data heterogeneity and distributed training imbalance are solved, efficient and secure multimodal data processing is achieved, and the stability and resource utilization of the model are improved.
Patent Information
- Application Number
- CN202510357826.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-07-11
AI Technical Summary
When the existing federated learning methods process multimodal data, there are problems such as heterogeneity of multimodal data, large privacy protection computing overhead and unbalanced distributed training, resulting in inefficient and poor practicality of model training.
By introducing multimodal feature alignment mechanism, lightweight privacy protection strategy and adaptive aggregation algorithm, the feature fusion and standardized representation of multimodal data are realized, and the distributed training process is optimized by protecting model parameters through differential privacy noise.
It improves the stability and consistency of model training, reduces the computing overhead of privacy protection, optimizes the resource utilization rate of distributed systems, and improves the efficiency and practicality of model training.
Smart Images

Figure CN120297448A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a distributed machine learning method for multi-modal data privacy protection based on federated learning, belonging to the technical field of federated learning. Background Art
[0002] With the rapid development of big data and artificial intelligence technologies, the processing and analysis of multi-modal data have wide applications in fields such as healthcare, finance, and intelligent transportation. However, multi-modal data is usually distributed across different devices or institutions and involves sensitive information (such as personal privacy data). Traditional centralized machine learning methods require uploading all data to a central server for processing, which not only increases the risk of data leakage but also leads to waste of computing resources due to data transmission and storage requirements.
[0003] Federated Learning, as a distributed machine learning framework, allows collaborative learning to be achieved by training models locally and only exchanging model parameters or gradients without sharing the original data. However, existing federated learning methods still face the following technical problems when dealing with multi-modal data:
[0004] (1) Heterogeneity of multi-modal data: The feature representation dimensions and distributions of different modal data (such as text and images) vary significantly, resulting in low model training efficiency;
[0005] (2) Computational overhead of privacy protection: Existing privacy protection mechanisms (such as differential privacy or homomorphic encryption) have high computational complexity and communication costs in multi-modal scenarios, limiting the practicality of the methods;
[0006] (3) Imbalance of distributed training: The significant differences in computing power, data volume, and communication latency among participating parties affect the convergence and stability of the global model.
[0007] In response to the above problems, there are already some improvement solutions in the prior art. Some literature has proposed a federated learning method based on feature fusion, but it requires additional feature preprocessing steps, increasing the computational burden; other literature has introduced encryption technology to protect parameter updates but has not solved the problem of multi-modal data heterogeneity. Therefore, there is an urgent need for a comprehensive method that can simultaneously solve the heterogeneity of multi-modal data, privacy protection overhead, and imbalance of distributed training. Summary of the Invention
[0008] In view of this, the purpose of the present invention is to provide a distributed machine learning method for multi-modal data privacy protection based on federated learning, which solves the problems of multi-modal data heterogeneity, large privacy protection computational overhead, and imbalance of distributed training by introducing a multi-modal feature alignment mechanism, a lightweight privacy protection strategy, and an adaptive aggregation algorithm.
[0009] The object of the present invention is achieved by the following technical solutions:
[0010] A distributed machine learning method for multi-modal data privacy protection based on federated learning, comprising the following steps:
[0011] S1. Data preprocessing and feature extraction: On multiple distributed nodes, preprocess the locally stored multi-modal data, extract the feature representations of each modality respectively, and perform feature matching;
[0012] S2. Multi-modal feature fusion and alignment: Using a feature alignment model, iteratively update through global and local parameters, map the heterogeneous features of each node to a unified feature space, and form a standardized feature representation;
[0013] S3. Local model training: On each distributed node, use the standardized feature representation to train the local machine learning model, and generate local model parameter updates;
[0014] S4. Privacy protection processing: Use privacy protection technology to process the model, perform lightweight privacy protection processing on the local model parameter updates, and generate encrypted parameter updates;
[0015] S5. Parameter aggregation: On the central server, receive the encrypted parameter updates uploaded by each distributed node, and perform weighted aggregation on the parameter updates through an adaptive aggregation algorithm to generate global model parameters;
[0016] S6. Global model distribution: Distribute the global model parameters to each distributed node, update the local model, and repeat steps S3 to S6 until the model converges or reaches a preset number of training rounds.
[0017] Further, in step S1, on multiple distributed nodes, preprocess the locally stored multi-modal data, extract the feature representations of each modality, calculate the feature similarity, determine the matching feature pairs, and ensure the efficiency of subsequent feature alignment and model training. Among them, the multi-modal data includes but is not limited to text, image and audio data.
[0018] Further, in step S2, based on a preset feature alignment model, map the different modality features extracted by each distributed node to a unified feature space, eliminate the influence of heterogeneity, and form a standardized feature representation, where the feature alignment is achieved through the following formula:
[0019] F i =W i ×X i +b i
[0020]
[0021] Among them, F i is the normalized feature representation of the i-th node;
[0022] X i is the original feature of the i-th node;
[0023] W i and b i are the weight matrix and bias vector of the alignment model respectively;
[0024] represents the square of the L2 norm, and the goal is to minimize the distance between the feature of different nodes.
[0025] Furthermore, the feature alignment is realized by iteratively updating the global and local parameters, and the specific implementation of the alignment process includes:
[0026] a) The central server initializes the global feature alignment model parameters W g and b g ;
[0027] b) Distribute W g and b g to each node;
[0028] c) Each distributed node optimizes the local alignment parameters W i and b i based on the local data X i , and calculates the feature representation through the formula F i ;
[0029] d) Upload the local parameter update to the central server, and update the global parameters through weighted average:
[0030]
[0031] e) Iteratively execute steps b) to d) until the feature alignment model converges.
[0032] Furthermore, in step S4, the lightweight privacy protection process adds differential privacy noise through the following formula:
[0033]
[0034] Among them, α i is the encrypted parameter update of the i-th node, β i is the original parameter update, Lap(·) represents Laplace distribution noise, Δf is the sensitivity of the parameter update, and ε i is the privacy budget of the i-th node, and the range is usually [0.1, 1.0].
[0035] Furthermore, privacy protection supports dynamic εi Configuration and parameter compression, and the lightweight privacy protection process further includes:
[0036] a) Calculate the sensitivity of parameter updates where δ i is the parameter update on the neighboring dataset, usually limited to 1 by gradient clipping;
[0037] b) Add noise according to the formula where ε i is the privacy budget, dynamically selected by each node within the interval [0.1, 1.0] according to privacy requirements, with a default of 0.5;
[0038] c) Quantize and compress the parameters after adding noise to reduce communication overhead.
[0039] Furthermore, in step S5, the specific steps of parameter aggregation include:
[0040] a) On the central server, receive the encrypted parameter updates uploaded by each distributed node, and preliminarily check data integrity and format;
[0041] b) Standardize the parameter data and handle outliers and missing values;
[0042] c) Select an aggregation algorithm according to the parameter type and objective;
[0043] d) Perform weighted aggregation on the parameter updates through the adaptive aggregation algorithm, perform summary calculations on the distributed data, and verify the aggregation result.
[0044] If the verification is qualified, output the aggregation result for subsequent use;
[0045] If the verification is unqualified, through the anomaly detection and elimination mechanism, iteratively execute steps a) to d) to generate global model parameters.
[0046] Furthermore, the adaptive aggregation algorithm is calculated by the following formula:
[0047]
[0048] In the formula, θ g is the global model parameter, N is the total number of distributed nodes, ω i is the aggregation weight of the i-th node, n i is the data volume of the i-th node in the local sample number, γ i is the computing power factor of the i-th node; defined as where ∈ i is the computing resource of the node;
[0049] To ensure aggregation stability, abnormal parameter updates are performed through an anomaly detection and rejection mechanism, specifically including:
[0050] a) Calculate the computing power factor of each node where ∈ i is the computing resource of the i-th node;
[0051] b) Allocate aggregation weights according to the formula ω i ;
[0052] c) Detect abnormal parameter updates. If ( is the parameter mean and τ is the threshold), then reject this update.
[0053] Furthermore, in step S6, the specific steps for model convergence or training are:
[0054] a) After each round of training, evaluate the global model performance using the local validation set. If the accuracy drops by more than 5%, then adjust ε i or ω i ;
[0055] b) After the model converges, perform multi-modal task verification to ensure the stability and accuracy of parameter distribution.
[0056] Furthermore, this method is implemented through a module set on the central server. The module includes:
[0057] A data preprocessing module for preprocessing local multi-modal data and feature extraction on multiple distributed nodes;
[0058] A feature alignment module for mapping different modal features to a unified feature space through the formula F i = W i × X i + b i ;
[0059] A local training module for training the local model based on the standardized feature representation and generating parameter updates;
[0060] A privacy protection module for protecting the privacy of parameter updates through the formula ;
[0061] A parameter aggregation module for weighted aggregation of encrypted parameters through the formula ;
[0062] A model distribution module for distributing global model parameters to each distributed node;
[0063] It also includes an anomaly detection module for identifying and rejecting abnormal parameter updates
[0064] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0065] (1) High precision. Through formulaic feature alignment and parameter aggregation, the stability and consistency of model training are improved.
[0066] (2) Good security. The accurate calculation of differential privacy noise effectively balances privacy protection and model performance, and the computational overhead of privacy protection is small.
[0067] (3) High model training efficiency. Adaptive weight allocation optimizes the resource utilization rate of the distributed system.
[0068] It effectively solves the problems in the existing federated learning methods when dealing with multi-modal data, such as uneven distributed training, significant differences in the computing power, data volume, and communication latency of each participant, poor convergence and stability of the global model, high computational complexity and communication cost of the privacy protection mechanism in the multi-modal scenario, and large differences in the feature representation dimensions and distributions of different modal data, resulting in low model training efficiency. It has strong practicability and high popularization potential.
[0069] Other advantages, objectives, and features of the present invention will be described to some extent in the subsequent specification, and to some extent, will be obvious to those skilled in the art based on the study of the following text, or can be taught from the practice of the present invention. The objectives and other advantages of the present invention can be achieved and obtained through the following specification. Brief Description of the Drawings
[0070] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings, where:
[0071] Figure 1 is the flowchart of the multi-modal data privacy protection distributed machine learning method based on federated learning provided by the present invention;
[0072] Figure 2 is the flowchart of multi-modal feature fusion and alignment;
[0073] Figure 3 is the flowchart of privacy protection processing;
[0074] Figure 4 is the flowchart of parameter aggregation;
[0075] Figure 5 is the system architecture diagram of the multi-modal data privacy protection distributed machine learning method based on federated learning;
[0076] Figure 6 is the architecture diagram of the feature alignment module. Detailed Embodiments
[0077] Next, in combination with the accompanying drawings in the embodiments of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described. Obviously, the described embodiments are only partial embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without making creative efforts belong to the scope of protection of the present invention.
[0078] The object of the present invention is to provide a multi-modal data privacy protection distributed machine learning method based on federated learning. By introducing a multi-modal feature alignment mechanism, a lightweight privacy protection strategy, and an adaptive aggregation algorithm, the problems of multi-modal data heterogeneity, large privacy protection computing overhead, and unbalanced distributed training are solved.
[0079] As Figures 1-6 shown, the first embodiment of the present invention provides a multi-modal data privacy protection distributed machine learning method based on federated learning, which includes the following steps:
[0080] S1. Data preprocessing and feature extraction. On multiple distributed nodes, preprocess the multi-modal data stored locally, extract the feature representations of each modality respectively, and perform feature matching;
[0081] S2. Multi-modal feature fusion and alignment. Using the feature alignment model, iteratively update through global and local parameters, map the heterogeneous features of each node to a unified feature space, and form a standardized feature representation;
[0082] S3. Local model training. On each distributed node, use the standardized feature representation to train the local machine learning model, and generate local model parameter updates;
[0083] S4. Privacy protection processing. Use privacy protection technology to process the model, perform lightweight privacy protection processing on the local model parameter updates, and generate encrypted parameter updates;
[0084] S5. Parameter aggregation. On the central server, receive the encrypted parameter updates uploaded by each distributed node, and perform weighted aggregation on the parameter updates through the adaptive aggregation algorithm to generate global model parameters;
[0085] S6. Global model distribution. Distribute the global model parameters to each distributed node, update the local model, and repeat steps S3 to S6 until the model converges or reaches the preset number of training rounds.
[0086] In step S1, on multiple distributed nodes, the locally stored multimodal data is preprocessed, the feature representations of each modality are extracted, the feature similarity is calculated, and the matching feature pairs are determined to ensure the efficiency of subsequent feature alignment and model training. Among them, the multimodal data includes but is not limited to text, image, and audio data, and the feature extraction includes the adaptive processing of text, image, and audio data. The feature extraction methods and parameter designs for each modality are as follows:
[0087] (1) For text data, a pre-trained model based on Transformer (such as BERT) is used to extract word embedding vectors;
[0088] (2) For image data, a convolutional neural network (such as ResNet-50) is used to extract deep features;
[0089] (3) For audio data, short-time Fourier transform (STFT) or mel-frequency cepstral coefficients (MFCC) are used to extract time-frequency features.
[0090] To reduce the impact of data heterogeneity, each node adaptively adjusts the feature extraction parameters according to the statistical distribution of local data (such as mean, variance), for example, adjusting the convolutional kernel size or word embedding dimension to optimize the quality of feature representation.
[0091] In step S2, based on a preset feature alignment model, the different modality features extracted by each distributed node are mapped to a unified feature space to eliminate the impact of heterogeneity and form a standardized feature representation, where the feature alignment is achieved through the following formula:
[0092] F i =W i ×X i +b i
[0093]
[0094] where, F i is the standardized feature representation of the i-th node;
[0095] X i is the original feature of the i-th node;
[0096] W i and b i are the weight matrix and bias vector of the alignment model respectively;
[0097] represents the square of the L2 norm, and the goal is to minimize the distance between the features of different nodes.
[0098] The feature alignment is achieved through the iterative update of global and local parameters. The specific implementation of the alignment process includes:
[0099] a) The central server initializes the global feature alignment model parameters W g and b g ;
[0100] b) Distribute W g and b g to each node;
[0101] c) Each distributed node optimizes the local alignment parameters W i and b i based on the local data X i , and calculates the feature representation through the formula F i ;
[0102] d) Upload the local parameter updates to the central server, and update the global parameters through weighted average:
[0103]
[0104]
[0105] e) Iteratively execute steps b) to d) until the feature alignment model converges.
[0106] In step S3, on each distributed node, the local machine learning model is trained using the normalized feature representation to generate local model parameter updates. Among them, each node uses the normalized feature F i to train the local machine learning model (such as a neural network), updates the model parameters through the gradient descent method, and generates the parameter update β i . The training objective can be a classification, regression, or generation task, and the specific loss function is determined according to the application scenario.
[0107] In step S4, collect the parameter updates on the local model, protect the privacy of the data by adding Laplace noise (Lap(Δf / ε i )) while trying to maintain the availability of the data, and generate the encrypted parameter updates. Among them, the lightweight privacy protection process adds differential privacy noise through the following formula:
[0108]
[0109] where α i is the encrypted parameter update of the i-th node, β i is the original parameter update, Lap(·) represents the Laplace distribution noise, Δf is the sensitivity of the parameter update, and ε iis the privacy budget for the i-th node, usually in the range of [0.1, 1.0], and is dynamically configured by the node according to privacy requirements. To further reduce communication overhead, the parameters after adding noise can be compressed through quantization coding (such as 8-bit integer quantization).
[0110] Privacy protection supports dynamic ε i Configuration and parameter compression, the lightweight privacy protection process further includes:
[0111] a) Calculate the sensitivity of parameter updates where δ i is the parameter update on the neighboring dataset, usually limited to 1 by gradient clipping;
[0112] b) Add noise according to the formula where ε i is the privacy budget, dynamically selected by each node within the interval [0.1, 1.0] according to privacy requirements, with a default of 0.5;
[0113] c) Quantize and compress the parameters after adding noise, for example, map the parameter values to 8-bit integers to reduce communication overhead.
[0114] In step S5, the specific steps of parameter aggregation include:
[0115] a) On the central server, receive the encrypted parameter updates uploaded by each distributed node, and preliminarily check data integrity and format;
[0116] b) Standardize the parameter data, and handle outliers and missing values;
[0117] c) Select an aggregation algorithm according to the parameter type and target;
[0118] d) Perform weighted aggregation on the parameter updates through an adaptive aggregation algorithm, perform summary calculations on distributed data, and perform aggregation result verification.
[0119] If the verification is qualified, output the aggregation result for subsequent use;
[0120] If the verification is unqualified, through the outlier detection and elimination mechanism, iteratively execute steps a) to d) to generate global model parameters.
[0121] Among them, the central server receives the encrypted parameter updates α i from each node, and generates the global model parameter θg through an adaptive aggregation algorithm. The adaptive aggregation algorithm is calculated by the following formula:
[0122]
[0123] In the formula, θ gis the global model parameter, N is the total number of distributed nodes, ω i is the aggregation weight of the i-th node, n i is the data volume of the i-th node in the local sample number, γ i is the computing power factor of the i-th node; defined as where ∈ i is the computing resource of the node (such as CPU frequency).
[0124] To ensure the aggregation stability, abnormal parameter updates (such as outliers caused by malicious attacks) are detected through ( is the parameter mean, τ is the threshold) and are removed. The specific steps include:
[0125] a) Calculate the computing power factor of each node where ∈ i is the computing resource of the i-th node (such as CPU frequency);
[0126] b) Allocate the aggregation weight according to the formula ω i ;
[0127] c) Detect the abnormal parameter update, if ( is the parameter mean, τ is the threshold), then remove this update.
[0128] In step S6, the specific steps for model convergence or training are:
[0129] a) After each round of training, evaluate the global model performance using the local validation set. If the accuracy drops by more than 5%, then adjust ε i or ω i ;
[0130] b) After the model converges, perform multi-modal task verification, such as image classification or text sentiment analysis, to ensure the stability and accuracy of parameter distribution.
[0131] Based on the above multi-modal data privacy protection distributed machine learning method for federated learning, it includes six core steps. The technical details of each step are elaborated below, especially the specific parameters of feature extraction, the mathematical derivation of noise calculation, and the optimization method of aggregation weight.
[0132] 1. Data preprocessing and feature extraction
[0133] Technical objective: On multiple distributed nodes, preprocess the locally stored multi-modal data, extract the feature representations of each modality, and ensure the efficiency of subsequent feature alignment and model training.
[0134] Technical implementation:
[0135] Multimodal data includes, but is not limited to, text, images, and audio. The feature extraction methods and parameter designs for each modality are as follows:
[0136] 1.1 Feature Extraction of Text Data
[0137] Method: Use a pre-trained model based on Transformer (such as BERT) to extract word embedding vectors.
[0138] Specific parameters:
[0139] Input: The original text sequence, with the maximum length set to 128 words (truncate if exceeding, pad if insufficient);
[0140] Model: BERT-Base (12-layer Transformer, 768-dimensional hidden layer, 12 attention heads);
[0141] Output: Take the 768-dimensional vector corresponding to the [CLS] token as the sentence-level feature representation;
[0142] Hyperparameters: Learning rate 2×10 5 , batch size 16, optimizer Adam (β1 = 0.9, β2 = 0.999);
[0143] Adaptive adjustment: Dynamically adjust the maximum sequence length according to the word frequency distribution of the text data. For example, if 90% of the text lengths are less than 100, adjust it to 100 to reduce the padding overhead.
[0144] 1.2 Feature Extraction of Image Data
[0145] Method: Use ResNet-50 (Residual Network) to extract deep features.
[0146] Specific parameters:
[0147] Input: The image resolution is uniformly adjusted to 224×224 pixels, with three RGB channels;
[0148] Model: ResNet-50 (50-layer convolutional network, pre-trained on ImageNet);
[0149] Output: A 2048-dimensional feature vector after the global average pooling layer;
[0150] Hyperparameters: Convolution kernel size 3×3, stride 1, padding 1; batch size 32, learning rate 1×10 -4 ;
[0151] Adaptive adjustment: If the resolution distribution of the node image data is concentrated at a relatively low value (such as 128×128), adjust the input resolution and retrain the last layer to adapt to local characteristics.
[0152] 1.3. Audio data feature extraction
[0153] Method: Mel Frequency Cepstral Coefficients (MFCC) are used to extract time-frequency features.
[0154] Specific parameters:
[0155] Input: Audio sampling rate is 16 kHz, frame length is 25 ms, and frame shift is 10 ms;
[0156] Preprocessing: Pre-emphasis coefficient is 0.97, and Hamming window processing is performed;
[0157] MFCC: 40 Mel filters, 13-dimensional MFCC coefficients (including 0th order) + first and second order differences, a total of 39-dimensional features;
[0158] Output: Take the average of each audio segment to generate a fixed-length feature vector;
[0159] Adaptive adjustment: Dynamically adjust the number of filters according to the audio signal-to-noise ratio (SNR) (for example, increase the number of filters to 48 when SNR < 20 dB to capture more details).
[0160] When preprocessing multi-modal data, perform heterogeneity processing and optimize the computational complexity,
[0161] Heterogeneity processing: Each node adjusts parameters according to the local data statistical characteristics (such as mean, variance). For example, if the variance of the image data is large, increase the number of convolution kernels to enhance the feature expression ability.
[0162] Computational complexity: The complexity of text feature extraction is O(L H 2 )(L is the sequence length, H is the hidden layer dimension); for images, it is O(C H W) (C is the number of channels, H, W are the height and width); for audio, it is O(T F) (T is the number of frames, F is the number of filters). Specifically, optimizing the complexity includes:
[0163] Reduce dimensions: Such as reducing the hidden layer dimension H, the number of channels C, or the number of filters F;
[0164] Reduce the sequence length: Such as truncating the text or downsampling the audio;
[0165] Use efficient models: Such as using Depthwise Separable Convolution or Sparse Attention mechanism.
[0166] 2. Multi-modal feature alignment
[0167] Technical objective: Map the different modal features extracted by each node to a unified feature space to eliminate the influence of heterogeneity.
[0168] Technical implementation: Feature alignment is achieved through linear transformation and federated optimization. The core formula is:
[0169] F i = W i ·X i + b i ,
[0170]
[0171] 2.1 Parameter design
[0172] Input: X i is the original feature vector of the i-th node (dimension d i , such as 768, 2048, or 39);
[0173] Transformation matrix: d is the dimension of the unified feature space (default 256); b i ∈ R d ;
[0174] Initialization: W g and b g are randomly initialized on the central server and follow the normal distribution N(0, 0.01).
[0175] 2.2 Optimization process
[0176] Local optimization: Each node optimizes W i and b i based on the local data X i , and the loss function is:
[0177]
[0178] where is the feature predicted by the global model, and n i is the number of local samples.
[0179] Gradient descent update:
[0180] η = 0.01 is the learning rate;
[0181] Global update: The central server receives W i and b i , and performs weighted averaging:
[0182]
[0183] Convergence condition: Iterate until the change is less than 10^-4.
[0184] 2.3 Mathematical Derivation
[0185] Optimization Objective is equivalent to minimizing the variance between features. Assume F i is an independent and identically distributed random variable, then
[0186]
[0187] Approximate the global optimal solution through gradient descent to ensure feature consistency.
[0188] 2.4 Detail Optimization
[0189] Dimension Selection: d = 256 balances the expressive power and computational cost, and can be adjusted according to the task (such as 512 or 128);
[0190] Complexity: The computational complexity of a single node is O(n i ·d·d i ) and the communication complexity is O(N d d i ).
[0191] 3. Local Model Training
[0192] Technical Objective: Use the standardized feature F i to train the local model and generate the local parameter update β i .
[0193] Technical Implementation:
[0194] Model: By default, use a multi-layer perceptron (MLP) with a 3-layer structure (256-128-64) and the ReLU activation function.
[0195] Loss Function: Use cross entropy for classification tasks and mean squared error for regression tasks. For example, the classification loss:
[0196]
[0197] Optimization: Use the SGD optimizer with a learning rate of 0.01, a momentum of 0.9, and a batch size of 32.
[0198] Parameter Update:
[0199] Adaptive Adjustment: If the local data volume n i <1000, increase the number of training rounds to 20 to improve convergence.
[0200] 4. Privacy Protection Processing
[0201] Technical Objective: Add noise to β i to ensure privacy protection. Core formula:
[0202]
[0203] 4.1 Parameter Design
[0204] Sensitivity Δf: Usually limited to 1 by gradient clipping;
[0205] Privacy budget ε i : Range [0.1, 1.0], default 0.5;
[0206] Noise scale:
[0207] 4.2 Mathematical Derivation
[0208] Differential privacy definition: For any adjacent datasets D and D ' (Differing by only one sample), the output probability satisfies:
[0209]
[0210] Laplace noise: The probability density function is:
[0211] The noise variance is
[0212] Sensitivity calculation: Assume the gradient clipping threshold C = 1, then d θ is the parameter dimension.
[0213] 4.3 Detail Optimization
[0214] Dynamic ε i : If the node privacy requirement is high, then reduce ε i to 0.2;
[0215] Compression: Quantize to 8 bits, reducing the communication volume by about 75%;
[0216] Complexity: The noise generation complexity is O(d θ );
[0217] 5. Parameter Aggregation
[0218] Technical objective: Generate the global parameter θ through adaptive aggregation g , formula:
[0219]
[0220] 5.1 Parameter Design
[0221] Data volume n i : The number of local samples;
[0222] Computing power factor γ i: ∈ j is the CPU frequency (GHz).
[0223] 5.2 Optimization Method
[0224] Weight calculation: ω i Integrate the data volume and computing power to ensure fairness and efficiency. Optimization goal:
[0225] θ * is the ideal global parameter, approximated by weighting.
[0226] Anomaly detection: If (δ is the standard deviation), reject α i .
[0227] 5.3 Mathematical Deduction
[0228] Consistency: Assume that α i is an unbiased estimate of θ * , then E[θ g = θ * .
[0229] Variance: Smooth the noise effect by increasing ω i .
[0230] 5.4 Detail Optimization
[0231] Dynamic adjustment: If the communication delay increases, reduce the weight of μ i nodes.
[0232] Complexity: The aggregation complexity is O(N·d θ ).
[0233] 6. Global Model Distribution
[0234] Technical implementation: Broadcast θ g to each node, update the local model, and repeat steps 3 - 6. The communication uses a compressed format (such as Protobuf), and the bandwidth requirement is about d θ 4 bytes (32 - bit floating point).
[0235] For the above - mentioned distributed machine learning method for multi - modal data privacy protection in federated learning, this method is implemented through a module set on the central server. The module includes:
[0236] A data pre - processing module for pre - processing and feature extraction of local multi - modal data on multiple distributed nodes;
[0237] A feature alignment module for aligning features through the formula F i = W i ×Xi +b i Map different modality features to a unified feature space;
[0238] A local training module for training a local model based on the standardized feature representation and generating parameter updates;
[0239] A privacy protection module for protecting the privacy of parameter updates through the formula Protect the parameter updates;
[0240] A parameter aggregation module for weighted aggregation of encrypted parameters through the formula Aggregate the encrypted parameters;
[0241] A model distribution module for distributing the global model parameters to each distributed node.
[0242] The feature alignment module realizes feature alignment through iterative updates of global and local parameters, and supports dynamic adjustment of W i and b i .
[0243] The privacy protection module supports dynamic configuration of ε i and integrates a parameter compression function.
[0244] The system further includes an anomaly detection module for identifying and removing abnormal parameter updates.
[0245] Specific applications are as in Embodiment 1 and Embodiment 2
[0246] Embodiment 1: Medical scenario
[0247] Background: Hospital A has medical record text data (X1, dimension 300), and Hospital B has CT image data (X2, dimension 512).
[0248] Steps:
[0249] 1. Hospital A uses BERT to extract X1, and Hospital B uses ResNet to extract X2.
[0250] 2. Align the features through F i =W i ×X i +b i to unify the dimension to 256, and the feature distance converges after 5 rounds of iteration.
[0251] 3. Locally train a diagnostic model to generate β1 and β2.
[0252] 4. Add noise α i =β i +Lap(0.5), ε = 0.5.
[0253] 5. Aggregate θg = 0.6 × α1 + 0.4 × α2 (weights based on data volume).
[0254] After 10 iterations, the diagnostic accuracy rate reaches 92%.
[0255] Example 2: Smart home scenario
[0256] Background: In a smart home, Device A collects audio data and Device B collects video data.
[0257] Steps: Similar to Example 1, finally train a behavior recognition model with an accuracy rate of 94%.
[0258] The above are only preferred embodiments of the present invention, and do not impose any form of confidentiality restrictions on the present invention. Any simple modifications, equivalent changes, and modifications made to the above embodiments based on the technical content of the present invention and the technical essence of the present invention still fall within the scope of the technical solution of the present invention.
Claims
1. A distributed machine learning method for multi-modal data privacy protection based on federated learning, characterized in that, It includes the following steps: S1. Data preprocessing and feature extraction: On multiple distributed nodes, preprocess the multimodal data stored locally, extract the feature representations of each modality respectively, and perform feature matching; S2. Multimodal feature fusion and alignment: Use the feature alignment model to iteratively update through global and local parameters, map the heterogeneous features of each node to a unified feature space, and form a standardized feature representation; S3. Local model training: On each distributed node, use the standardized feature representation to train the local machine learning model to generate local model parameter updates; S4. Privacy protection processing: Use privacy protection technology to process the model, perform lightweight privacy protection processing on the local model parameter updates, and generate encrypted parameter updates; S5. Parameter aggregation: On the central server, receive the encrypted parameter updates uploaded by each distributed node, and perform weighted aggregation on the parameter updates through an adaptive aggregation algorithm to generate global model parameters; S6. Global model distribution: Distribute the global model parameters to each distributed node to update the local model, and repeat steps S3 to S6 until the model converges or reaches the preset number of training rounds.
2. The distributed machine learning method for multi-modal data privacy protection based on federated learning according to claim 1, wherein In step S1, on multiple distributed nodes, preprocess the multimodal data stored locally, extract the feature representations of each modality, calculate the feature similarity, determine the matching feature pairs, and ensure the efficiency of subsequent feature alignment and model training; among them, the multimodal data includes but is not limited to text, image, and audio data.
3. The distributed machine learning method for multi-modal data privacy protection based on federated learning according to claim 1, characterized in that, In step S2, based on the preset feature alignment model, map the different modality features extracted by each distributed node to a unified feature space, eliminate the influence of heterogeneity, and form a standardized feature representation, where the feature alignment is achieved through the following formula: F i = W i × X i + b i Among them, F i is the standardized feature representation of the i-th node; X i is the original feature of the i-th node; W i and b i are the weight matrix and bias vector of the alignment model, respectively; Denotes the square of the L2 norm, and the goal is to minimize the distance between different node features.
4. The distributed machine learning method for multi-modal data privacy protection based on federated learning according to claim 3, characterized in that The feature alignment is achieved through iterative update of global and local parameters, and the specific implementation of the alignment process includes: a) The central server initializes the global feature alignment model parameter W g and b g ; b) Distribute W g and b g to each node; c) Each distributed node optimizes the local alignment parameter W based on the local data X i and b i and calculates the feature representation through formula F i ; i d) Upload the local parameter updates to the central server and update the global parameters through weighted averaging: e) Iteratively execute steps b) to d) until the feature alignment model converges.
5. The distributed machine learning method for multi-modal data privacy protection based on federated learning according to claim 1, wherein In step S4, the lightweight privacy protection processing adds differential privacy noise through the following formula: Among them, α i is the encrypted parameter update of the i-th node, β i is the original parameter update, Lap(·) represents the Laplace distribution noise, Δf is the sensitivity of the parameter update, and ε i is the privacy budget of the i-th node, and the range is usually [0.1, 1.0].
6. The distributed machine learning method for multi-modal data privacy protection based on federated learning according to claim 5, wherein Privacy protection supports dynamic ε i Configuration and parameter compression, and the lightweight privacy protection process further includes: a) Calculate the sensitivity of parameter updates where δ i is the parameter update on the neighboring dataset, usually limited to 1 by gradient clipping; b) Add noise according to the formula where ε i is the privacy budget, which is dynamically selected by each node within the interval [0.1, 1.0] according to privacy requirements, with a default value of 0.5; c) Quantize and compress the parameters after adding noise to reduce communication overhead.
7. The distributed machine learning method for multi-modal data privacy protection based on federated learning according to claim 1, characterized in that, In step S5, the specific steps of parameter aggregation include: a) On the central server, receive the encrypted parameter updates uploaded by each distributed node and preliminarily check the data integrity and format; b) Standardize the parameter data and process outliers and missing values; c) Select an aggregation algorithm according to the parameter type and target; d) Perform weighted aggregation on the parameter updates through an adaptive aggregation algorithm, perform summary calculation of distributed data, and perform aggregation result verification, If the verification is qualified, output the aggregation result for subsequent use; If the verification is unqualified, through the anomaly detection and elimination mechanism, iteratively execute steps a) to d) to generate global model parameters.
8. The distributed machine learning method for multi-modal data privacy protection based on federated learning according to claim 7, characterized in that, The adaptive aggregation algorithm is calculated through the following formula: where θ g is the global model parameter, N is the total number of distributed nodes, ω i is the aggregation weight of the i-th node, n i is the data volume of the i-th node in the local samples, γ i is the computing power factor of the i-th node; defined as where ∈ i is the computing resource of the node; To ensure aggregation stability, perform abnormal parameter updates through the anomaly detection and elimination mechanism, specifically including: a) Calculate the computing power factor of each node where ∈ i is the computing resource of the i-th node; b) Assign the aggregation weights according to the formula ω i Allocate the aggregation weights; c) Detect abnormal parameter updates. If ( is the parameter mean and τ is the threshold), then reject this update.
9. The distributed machine learning method for multi-modal data privacy protection based on federated learning according to claim 8, wherein In step S6, the specific steps of model convergence or training are: a) After each round of training, use the local validation set to evaluate the performance of the global model. If the accuracy drops by more than 5%, then adjust ε i or ω i ; b) After the model converges, perform multi-modal task verification to ensure the stability and accuracy of parameter distribution.
10. The distributed machine learning method for multi-modal data privacy protection based on federated learning according to any one of claims 1-9, which is implemented by a module set on a central server, and the module includes: A data preprocessing module for preprocessing and feature extraction of local multi-modal data on multiple distributed nodes; A feature alignment module, which is used to map features of different modalities into a unified feature space through the formula F i = W i × X i + b i A local training module for training a local model based on the standardized feature representation and generating parameter updates; Privacy protection module, used to perform privacy protection on parameter updates through the formula ; A parameter aggregation module for weighted aggregation of encrypted parameters through the formula ; A model distribution module for distributing global model parameters to each distributed node; It further includes an anomaly detection module for identifying and removing abnormal parameter updates.
Citation Information
Cited By
Pathology detection and diagnosis method and system based on federal learning
CN120544942A
Federal learning utility optimization system and method for resisting data heterogeneity
CN121094169A