Cross-park enterprise data collaborative analysis method based on federal learning
By using dynamic feature alignment and adaptive nonlinear transformation, the difficulty of feature alignment in cross-campus data collaborative analysis was solved. A unified feature representation space across business domains was constructed, and a triple protection system was built by combining homomorphic encrypted transmission and gradient sparsity, thus achieving efficient cross-campus data collaborative analysis and model convergence.
Patent Information
- Application Number
- CN202511173958.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-21
- Publication Date
- 2025-11-14
Smart Images

Figure CN120956493A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of distributed machine learning and data security technology, and in particular to a cross-campus enterprise data collaborative analysis method based on federated learning. Background Technology
[0002] As the digitalization of global industrial parks accelerates, multinational corporations face an urgent need for cross-regional data collaborative analysis. Current mainstream solutions rely on centralized cloud computing architectures, requiring each park to transmit raw data to a central server for processing. However, increasingly stringent data sovereignty regulations prohibit the cross-border flow of sensitive information, resulting in isolated islands of business data scattered across different countries. Although federated learning technology allows for collaborative modeling without sharing raw data, significant bottlenecks remain in practical cross-park deployments.
[0003] The diverse business characteristics of different industrial parks result in highly heterogeneous data distribution. The time-series equipment data from manufacturing parks differs fundamentally from the transaction characteristics of financial parks. Traditional federated learning, employing static feature mapping rules, struggles to adapt to dynamically expanding business scenarios. When a new park joins, the feature engineering process often needs to be redesigned, causing model updates to lag behind business changes. Simultaneously, existing security mechanisms face a dilemma: homomorphic encryption, while ensuring privacy, incurs significant communication overhead, while lightweight noise injection can easily degrade model accuracy. More seriously, cross-park network environments present the risk of malicious nodes forging model updates, and detection methods based on fixed thresholds cannot handle adaptive attack strategies.
[0004] Furthermore, a key drawback of federated learning lies in the lack of a refined evaluation system for node contributions. In scenarios with non-independent and identically distributed data, the value of high-quality park nodes is often diluted by low-quality data, and traditional equal-weight aggregation strategies exacerbate model bias. On the other hand, convergence determination relies excessively on a single accuracy metric, failing to capture consistent changes in the distribution of data across multiple parks, often leading to premature termination of training or the accumulation of invalid communication rounds. These problems severely restrict the collaborative decision-making capabilities and model performance of multinational corporations. Summary of the Invention
[0005] To address the technical problems in existing technologies, such as difficulties in feature alignment due to data heterogeneity across different industrial parks, the inability to balance security and efficiency, inaccurate node contribution evaluation, and one-sided convergence determination mechanisms, this invention provides a collaborative analysis method for cross-industry enterprise data based on federated learning.
[0006] The technical solution provided by this invention is as follows:
[0007] This invention provides a cross-park enterprise data collaborative analysis method based on federated learning, comprising:
[0008] S1. Initialize the federated learning framework: The central server distributes the initial global model architecture and federated learning configuration parameters to the enterprise nodes in each park.
[0009] S2. Localized data preprocessing: Each enterprise node performs feature alignment processing on its local private dataset to generate standardized feature vectors;
[0010] S3. Collaboration Evaluation: Each node calculates the dynamic collaboration factor between its local dataset and the global data distribution;
[0011] S4. Local Model Training: Adjust the local training strategy based on dynamic co-factors and update model parameters using the local dataset;
[0012] S5. Security Parameter Aggregation: The central server collects model updates from each node through an encrypted channel and uses an aggregated offset threshold to filter valid updates.
[0013] S6. Global Model Generation: The filtered model updates are weighted and aggregated to generate a new global model;
[0014] S7. Iteration Termination Judgment: The process terminates when the global model meets the cross-park data convergence condition or reaches the maximum number of communication rounds.
[0015] Preferably, the feature alignment process in S2 includes:
[0016] S201. Extract the statistical feature vector of the local dataset;
[0017] S202. Perform dimension mapping with the global feature template issued by the central server;
[0018] S203. Embed heterogeneous data into a unified feature space through nonlinear transformation.
[0019] Preferably, the dynamic synergy factor in S3 is calculated using the following formula:
[0020]
[0021] Among them, Γ i Let σ(·) be the dynamic cooperation factor of node i; σ(·) is the local dataset D. i With global data distribution D g The KL divergence contraction function is defined as σ(x) = e -βx cosim(·) represents the local model gradient. With global gradient The cosine similarity is η; the gradient importance coefficient has a value range of [0.2, 0.8]; β is the divergence sensitivity parameter with a value range of [1.0, 3.0].
[0022] Preferably, the aggregation offset threshold in S5 is calculated using the following formula:
[0023]
[0024] Where, τ t μ is the aggregation offset threshold for the t-th round of communication; t This is the update amount for all nodes in this round of model updates. The mean vector of ; κ is the standard deviation scaling factor, with a value range of [1.5, 2.5]; N is the total number of nodes participating in federated learning.
[0025] Preferably, the local training strategy adjustment in S4 includes the following steps:
[0026] S401. Monitor the numerical relationship between the dynamic synergy factor and the preset threshold;
[0027] S402. When the dynamic collaboration factor is below the first threshold, perform adversarial example generation: calculate the Jacobian matrix of the input data based on the current local model parameters;
[0028] S403. Apply a perturbation to the original samples along the gradient direction of the loss function to generate an adversarial sample set;
[0029] S404. Merge the adversarial sample set with the original training data to form an augmented dataset;
[0030] S405. Add an L2 norm regularization term for the model parameters to the local loss function;
[0031] S406. When the dynamic coordination factor is higher than the second threshold, increase the number of local training iterations proportionally.
[0032] S407. Select the optimizer learning rate decay strategy based on the range of the dynamic co-factor value.
[0033] Preferably, the weighted aggregation of S6 includes the following steps:
[0034] S601. Construct a node aggregation weight allocation function: using dynamic collaborative factors as input variables;
[0035] S602. Standardize the node co-factor sequence to eliminate dimensional differences;
[0036] S603. Use S-shaped function mapping to convert standardized co-factors into initial weights;
[0037] S604. Normalize the initial weights so that the sum of the weights is 1;
[0038] S605. Maintain the node participation status record table and mark the historical aggregation participation status of each node;
[0039] S606, Detect isolated nodes that have not participated in aggregation for three consecutive rounds;
[0040] S607. Extract the most recent valid model update of the isolated node as the compensation benchmark;
[0041] S608: Integrate the compensation benchmark into the current round's global aggregation according to the attenuation coefficient.
[0042] Preferably, the aggregation of security parameters in S5 further includes:
[0043] S501. Generate homomorphic encryption key pairs and distribute the public key to each node;
[0044] S502, The node uses the public key encryption model to update parameters and generate ciphertext data packets;
[0045] S503: After receiving the encrypted data packet, the central server performs gradient sparsity processing.
[0046] S504. Calculate the L1 norm for each model update vector and sort them.
[0047] S505: Retain the top K largest absolute gradient dimensions to form a sparse gradient vector;
[0048] S506. Use the private key to decrypt the sparse gradient vector to obtain the plaintext update.
[0049] Preferably, the cross-park data convergence condition determination in S7 includes:
[0050] S701, Central Server builds cross-campus verification dataset;
[0051] S702. After each round of aggregation, perform validation set predictions in the new global model;
[0052] S703. Calculate the F1-score coefficient of variation of the prediction results for three consecutive rounds:
[0053] S704. The first convergence condition is triggered when the coefficient of variation remains below 0.5%.
[0054] S705. Extract the distribution of prediction results for the local test set of each node;
[0055] S706. Calculate the Jensen-Shannon divergence matrix of the predicted distribution between nodes;
[0056] S707. The second convergence condition is triggered when the mean of the divergence matrix drops to 10% of the initial calculated value.
[0057] Preferably, the dynamic adjustment of the β value in the KL divergence compression function includes:
[0058] S301, Central Server Maintenance Park Type - Parameter Mapping Table;
[0059] S302, Receive the park industry classification code reported by the node;
[0060] S303. When the code belongs to the manufacturing industry, the β baseline value is set to 2.5;
[0061] S304. When the code belongs to the financial industry, the β benchmark value is set to 1.2;
[0062] S305. The value should fluctuate within ±0.3 of the baseline value based on the size of the node data.
[0063] S306. Recalibrate the β value every five communication cycles.
[0064] Preferably, a step is included between S4 and S5:
[0065] S4a, Calculate the noise standard deviation scaling factor: λ i =1 / 1+Γ i ;
[0066] S4b, Obtain the preset reference noise level σ base ;
[0067] S4c generates a Gaussian distribution. Random noise;
[0068] S4d adds the noise vector element by element to the model update parameters;
[0069] S4e records the noise injection amount for subsequent aggregation compensation calculations.
[0070] The beneficial effects of the technical solution provided by this invention include at least the following:
[0071] (1) In this invention, a unified feature representation space across business domains is constructed through a dynamic feature alignment engine and adaptive nonlinear transformation. For heterogeneous data such as manufacturing sensor streams and financial transaction records, semantic matching and neural network encoding technologies are used to achieve seamless access to newly added industrial parks and automatic compensation of feature dimensions, significantly improving the efficiency of multi-source data fusion;
[0072] (2) In this invention, a triple protection system is constructed by combining homomorphic encrypted transmission, gradient sparsity, and dynamic anomaly filtering. While ensuring the confidentiality of updated parameters, malicious attacks are intercepted in real time through adaptive aggregation thresholds, and the communication load is reduced by selectively retaining gradient dimensions, effectively balancing security protection and system efficiency;
[0073] (3) In this invention, an innovative data-gradient dual-channel collaborative evaluation mechanism is designed, which quantifies node contributions by jointly using distribution similarity and model behavior. By combining the negative correlation strategy between noise injection and collaborative factors, high-quality nodes dominate the direction of model optimization, simultaneously improving the consistency of data distribution across multiple campuses and accelerating the model convergence process. Attached Figure Description
[0074] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0075] Figure 1 A flowchart illustrating a cross-park enterprise data collaborative analysis method based on federated learning, provided for an embodiment of the present invention;
[0076] Figure 2 A schematic diagram of the feature alignment process for a cross-park enterprise data collaborative analysis method based on federated learning, provided in an embodiment of the present invention;
[0077] Figure 3 A schematic diagram illustrating the local training strategy adjustment process of a cross-park enterprise data collaborative analysis method based on federated learning, provided for an embodiment of the present invention;
[0078] Figure 4 A schematic diagram of the security parameter aggregation process for a cross-park enterprise data collaborative analysis method based on federated learning, provided as an embodiment of the present invention;
[0079] Figure 5 This is a schematic diagram of the weighted aggregation process of a cross-park enterprise data collaborative analysis method based on federated learning, provided in an embodiment of the present invention. Detailed Implementation
[0080] The technical solution of the present invention will now be described with reference to the accompanying drawings.
[0081] In embodiments of the present invention, words such as "exemplarily," "for example," etc., are used to indicate that something is an example, illustration, or description. Any embodiment or design described as "exemplary" in the present invention should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of the word "exemplary" is intended to present the concept in a concrete manner. Furthermore, in embodiments of the present invention, the meaning expressed by "and / or" can be both, or either one.
[0082] In the embodiments of this invention, the terms "image" and "picture" may sometimes be used interchangeably. It should be noted that, without emphasizing the distinction between them, their intended meanings are consistent. Similarly, the terms "of," "corresponding (relevant)," and "corresponding" may sometimes be used interchangeably. It should be noted that, without emphasizing the distinction between them, their intended meanings are consistent.
[0083] In this embodiment of the invention, sometimes a subscript such as W1 may be mistakenly written as a non-subscript form such as W1. When the difference is not emphasized, the meaning they express is the same.
[0084] To make the technical problems, technical solutions and advantages of the present invention clearer, a detailed description will be given below in conjunction with the accompanying drawings and specific embodiments.
[0085] Reference manual attached Figure 1 The diagram illustrates a flowchart of a cross-park enterprise data collaborative analysis method based on federated learning, provided by an embodiment of the present invention.
[0086] This invention provides a cross-park enterprise data collaborative analysis method based on federated learning, and the processing flow may include the following steps:
[0087] S1. Initialize the federated learning framework: The central server distributes the initial global model architecture and federated learning configuration parameters to the enterprise nodes in each park.
[0088] It should be noted that the central server first deploys the federated learning coordination module, establishes a two-way certified communication link with enterprise nodes in each park through digital certificates, preloads a standardized global model architecture (such as ResNet-18 or Transformer infrastructure), and generates a configuration file containing communication protocol parameters (specifying the gRPC-over-TLS transport protocol, setting a 5-second heartbeat interval and 3 timeout retries), model initialization parameters (convolutional layer weights are initialized using the Kaiming normal distribution, and fully connected layer biases are set to zero vectors), and federated basic configuration (defining a maximum of 100 communication rounds and a minimum participation node ratio of 80% in each round). Subsequently, the initialization package is distributed to each node through a fragmented transmission mechanism. After receiving the package, each node verifies the validity of the certificate chain and creates a protected model sandbox environment locally.
[0089] Furthermore, the initialization package adopts a fragmented redundancy transmission protocol, setting the size of each data fragment to 1MB and attaching a CRC32 checksum. After receiving the fragment, the node performs a hash check (SHA-256). If the check fails, it automatically requests a retransmission. Three consecutive check failures trigger an alarm and suspend the federated learning process.
[0090] S2. Localized data preprocessing: Each enterprise node performs feature alignment processing on its local private dataset to generate standardized feature vectors.
[0091] It should be noted that the enterprise node calls the local data engine to scan the private dataset (including database tables and CSV file streams) and executes a multi-stage feature processing pipeline: First, it automatically identifies numerical / categorical feature fields for structured data, and uses a pre-trained ViT model to extract 128-dimensional feature vectors for unstructured data (such as images); then, it applies format preservation encryption (FPE) to fields containing PII, and performs z-score normalization on continuous numerical features; finally, it maps the extracted features to a unified feature space template issued by the server, fills missing dimensions with null markers, and caches the processed feature vectors in Protobuf binary format in an encrypted storage area.
[0092] In one possible implementation, such as Figure 2 As shown, the feature alignment process in S2 includes:
[0093] S201. Extract the statistical feature vector of the local dataset;
[0094] S202. Perform dimension mapping with the global feature template issued by the central server;
[0095] S203. Embed heterogeneous data into a unified feature space through nonlinear transformation.
[0096] It should be noted that when a node performs feature alignment, it first scans the local dataset to extract statistical feature vectors (including the mean / variance of numerical features and the mode distribution of categorical features); then it performs dimension mapping with the global feature template issued by the server, and unifies the heterogeneous feature names into 128-bit identifiers through hash encoding; for scenarios with missing dimensions, a multilayer perceptron (MLP) is used to perform nonlinear transformation: the original features are input into an MLP network with two hidden layers (dimensions of 64 and 32 respectively), the activation function of the output layer uses Tanh, and finally an embedding vector with the same dimension as the global template is generated, thus completing the spatial alignment of heterogeneous data.
[0097] Furthermore, the dimension mapping employs a two-level matching strategy: Level 1 matching directly maps fields with identical feature names; Level 2 matching establishes a mapping relationship for feature fields with semantic similarity > 85% (calculated based on Word2Vec word vector cosine similarity), and unmatched fields are marked as dimensions to be processed and enter the nonlinear transformation process.
[0098] S3. Collaboration Evaluation: Each node calculates the dynamic collaboration factor between its local dataset and the global data distribution.
[0099] It should be noted that the node-initiated distribution evaluation engine loads the statistical summary of the global data distribution (the mean and covariance matrix broadcast by the server) and performs a dual-channel analysis: in the data distribution comparison channel, it calculates the statistical difference measure (such as JS divergence or Wasserstein distance) between the local dataset and the global distribution and generates a standardized difference score; at the same time, in the model behavior observation channel, it records the rate of change of the angle between the gradient descent direction and the global average gradient; finally, the dual-channel outputs are fused to generate a dynamic collaborative index to quantify the potential contribution of node data to the federated task.
[0100] Furthermore, the dual-channel output employs adaptive weighted fusion: the data distribution comparison channel weights are set to... The weights of the model behavior observation channels are set to 1-α, where For local data volume, This is a global estimate of the total amount of data.
[0101] In one possible implementation, the dynamic synergy factor in S3 is calculated using the following formula:
[0102]
[0103] Among them, Γ i Let σ(·) be the dynamic cooperation factor of node i; σ(·) is the local dataset D. i With global data distribution D g The KL divergence contraction function is defined as σ(x) = e -βx cosim(·) represents the local model gradient. With global gradient The cosine similarity is η; the gradient importance coefficient has a value range of [0.2, 0.8]; β is the divergence sensitivity parameter with a value range of [1.0, 3.0].
[0104] It should be noted that during the synergy evaluation phase, two types of quantitative analysis are performed simultaneously when nodes calculate the dynamic synergy factor: firstly, the local data distribution D is calculated based on KL divergence. i With global distribution D g The original degree of difference is expressed by the exponential function σ(x) = e -βx Compress to the [0,1] interval (where β is dynamically set according to rules); simultaneously capture local gradients during forward propagation. Global gradients broadcast by the server Calculate the cosine similarity; finally, follow the formula. The fusion result (η is preset to 0.5) is calculated, written to the local log, and uploaded to the server.
[0105] In one possible implementation, dynamic adjustment of the β value in the KL divergence compression function includes:
[0106] S301, Central Server Maintenance Park Type - Parameter Mapping Table;
[0107] S302, Receive the park industry classification code reported by the node;
[0108] S303. When the code belongs to the manufacturing industry, the β baseline value is set to 2.5;
[0109] S304. When the code belongs to the financial industry, the β benchmark value is set to 1.2;
[0110] S305. The value should fluctuate within ±0.3 of the baseline value based on the size of the node data.
[0111] S306. Recalibrate the β value every five communication cycles.
[0112] It should be noted that the server maintains a park type-parameter mapping table (manufacturing code range 1000-1999, financial industry 2000-2999). Nodes report the park code during registration; if the code ∈ [1000, 1999], then β = 2.5 + 0.1 × (log 10 |D i |-3)(|D i (where β represents the local data volume), the financial industry coding is set as β = 1.2 - 0.05 × (log(data / ... 10 |D i |-4). After every 5 rounds of communication, the server adjusts the node data distribution based on the rate of change δ. d =||μ t -μ t-5 ||2 Recalibrate β value: If δ d If the value is greater than 0.1, then β:=β×1.05; otherwise, β:=β×0.98.
[0113] S4. Local Model Training: Adjust the local training strategy based on dynamic co-factors and update model parameters using the local dataset.
[0114] It should be noted that the training hyperparameters are dynamically configured based on the collaborative evaluation results: when the collaborative index is higher than the threshold θ1, full data training is enabled and the number of iterations is increased to 150% of the baseline value; when the index is lower than the threshold θ2, adversarial training mode is activated and an FGSM perturbation module is inserted during the backpropagation stage to generate adversarial examples; gradient pruning (threshold 2.0) is used during the training process to prevent gradient explosion, and the AdamW optimizer (weight decay coefficient 0.01) is selected. After training, the incremental parameters of the model are exported with Float16 precision.
[0115] In one possible implementation, such as Figure 3 As shown, adjusting the local training strategy in S4 includes the following steps:
[0116] S401. Monitor the numerical relationship between the dynamic synergy factor and the preset threshold;
[0117] S402. When the dynamic collaboration factor is below the first threshold, perform adversarial example generation: calculate the Jacobian matrix of the input data based on the current local model parameters;
[0118] S403. Apply a perturbation to the original samples along the gradient direction of the loss function to generate an adversarial sample set;
[0119] S404. Merge the adversarial sample set with the original training data to form an augmented dataset;
[0120] S405. Add an L2 norm regularization term for the model parameters to the local loss function;
[0121] S406. When the dynamic coordination factor is higher than the second threshold, increase the number of local training iterations proportionally.
[0122] S407. Select the optimizer learning rate decay strategy based on the range of the dynamic co-factor value.
[0123] It should be noted that when the dynamic synergy factor Γ i When Γ < 0.3, the node initiates an adversarial training mechanism: it calculates the Jacobian matrix of the input data loss to the model, applies a perturbation of magnitude ∈ = 0.01 to the original samples along the gradient sign direction to generate adversarial examples, and mixes the adversarial examples with the original data in a 1:1 ratio to form an augmented dataset. i If the value is greater than 0.7, the number of local training epochs will be increased from the baseline of 20 epochs to 30 epochs. During training, an L2 regularization term with a weight decay coefficient of 0.001 will be added to the loss function, and the result will be determined according to Γ. i Learning rate decay strategy based on value range: when Γ i Cosine annealing decay is used when the range is ∈[0.4,0.6], otherwise step decay is used.
[0124] In one possible implementation, a step is further included between S4 and S5:
[0125] S4a, Calculate the noise standard deviation scaling factor: λ i =1 / 1+Γ i ;
[0126] S4b, Obtain the preset reference noise level σ base ;
[0127] S4c generates a Gaussian distribution. Random noise;
[0128] S4d adds the noise vector element by element to the model update parameters;
[0129] S4e records the noise injection amount for subsequent aggregation compensation calculations.
[0130] It should be noted that after local training is completed, the node calculates the noise scaling factor λ. i =1 / 1+Γ i Load the preset reference noise σ base =0.05. Generate Gaussian noise vector. Add to model update parameters by element: Record the noise injection amount ||∈||2 and transmit it to the server with the update, which is used for reverse correction of weight allocation during subsequent aggregation compensation calculation.
[0131] S5. Security Parameter Aggregation: The central server collects model updates from each node through an encrypted channel and uses an aggregated offset threshold to filter valid updates.
[0132] It should be noted that the central server performs security operations through the aggregation gateway: the nodes first use the server's public key to encrypt the model increment and transmit the ciphertext through the TLS 1.3 channel; the server calculates the Mahalanobis distance matrix of all node updates in real time to isolate abnormal updates that deviate from the population distribution 3σ; then, based on the statistical distribution characteristics, the aggregation threshold is dynamically calculated, allowing only updates that conform to the Gaussian distribution to enter the aggregation pool, and the filtered ciphertext updates are temporarily stored in the Trusted Execution Environment (TEE).
[0133] Furthermore, the abnormal update isolation process implements a three-level handling procedure: Level 1 isolation removes the update from the aggregation pool; Level 2 isolation freezes the node's participation eligibility for 1-3 rounds; Level 3 isolation permanently disables nodes that exhibit abnormal behavior for 3 consecutive rounds and generates security audit logs which are uploaded to the monitoring module.
[0134] In one possible implementation, the aggregation offset threshold in S5 is calculated using the following formula:
[0135]
[0136] Where, τ t μ is the aggregation offset threshold for the t-th round of communication; t This is the update amount for all nodes in this round of model updates. The mean vector of N is denoted by N; k is the standard deviation scaling factor, with a value range of [1.5, 2.5]; N is the total number of nodes participating in federated learning.
[0137] It should be noted that the central server dynamically generates a threshold before each round of aggregation: collecting model update vectors from all nodes. Then, calculate its mean vector μ. tThen, calculate the sum of squared Euclidean distances between each updated vector and the mean, divide by the total number of nodes N to obtain the variance estimate; take the square root of this value, multiply by the standard deviation scaling factor κ (fixed at 2.0), and then weight it onto the mean vector, i.e., according to the formula... Output a scalar threshold that is used to filter out anomalous updates that deviate from the population distribution.
[0138] In one possible implementation, such as Figure 4 As shown, the aggregation of security parameters in S5 also includes:
[0139] S501. Generate homomorphic encryption key pairs and distribute the public key to each node;
[0140] S502, The node uses the public key encryption model to update parameters and generate ciphertext data packets;
[0141] S503: After receiving the encrypted data packet, the central server performs gradient sparsity processing.
[0142] S504. Calculate the L1 norm for each model update vector and sort them.
[0143] S505: Retain the top K largest absolute gradient dimensions to form a sparse gradient vector;
[0144] S506. Use the private key to decrypt the sparse gradient vector to obtain the plaintext update.
[0145] It should be noted that the server generates Paillier homomorphic encryption key pairs (key length 2048 bits), and the public key is distributed to each node after being signed by a CA. Nodes use the public-key encryption model to update and generate ciphertext data packets (ciphertext block size 256 bytes). After receiving these, the server calculates the L1 norm for each encryption update vector, sorts the gradients by absolute value of their dimensions, and retains the top 10% of gradients. All other dimensions are set to zero; the sparsified ciphertext is decrypted using the private key to obtain the plaintext sparse gradient vector for aggregation.
[0146] S6. Global Model Generation: The filtered model updates are weighted and aggregated to generate a new global model.
[0147] It should be noted that after decrypting the valid update within the TEE, multi-dimensional aggregation is performed: non-uniform weights are calculated based on the node's historical contribution (such as the quality score of the last 5 rounds of updates), and the geometric median method is used to calculate the center point of the update vector instead of the traditional weighted average. At the same time, nodes that lost updates due to network failures are injected with an exponentially decaying copy of their most recent valid update. The newly generated global model is then broadcast to all nodes after being verified by SHA-256.
[0148] In one possible implementation, such as Figure 5As shown, the weighted aggregation of S6 includes the following steps:
[0149] S601. Construct a node aggregation weight allocation function: using dynamic collaborative factors as input variables;
[0150] S602. Standardize the node co-factor sequence to eliminate dimensional differences;
[0151] S603. Use S-shaped function mapping to convert standardized co-factors into initial weights;
[0152] S604. Normalize the initial weights so that the sum of the weights is 1;
[0153] S605. Maintain the node participation status record table and mark the historical aggregation participation status of each node;
[0154] S606, Detect isolated nodes that have not participated in aggregation for three consecutive rounds;
[0155] S607. Extract the most recent valid model update of the isolated node as the compensation benchmark;
[0156] S608: Integrate the compensation benchmark into the current round's global aggregation according to the attenuation coefficient.
[0157] It should be noted that when the central server allocates aggregation weights, it first processes the node collaboration factor sequence {Γ}. i ,...,Γ N Perform Z-score standardization; input the standardized result into the Sigmoid function S(x) = 1 / (1+e^x). -x The initial weights are mapped to the target weights; the final weights are obtained through Softmax normalization. For nodes that have not participated in aggregation for three consecutive rounds, retrieve their most recent valid update (no more than five rounds of history), and apply it with a decay coefficient γ = 0.7. t (t is the absent round) After scaling, it is injected into the current aggregation pool and participates in the geometric median calculation together with the regular update.
[0158] Furthermore, the attenuation coefficient calculation follows the time decay law: γ t =e -λt Where λ = 0.2 is the decay rate constant, t is the number of absent rounds, and compensation stops when t > 5.
[0159] S7. Iteration Termination Judgment: The process terminates when the global model meets the cross-park data convergence condition or reaches the maximum number of communication rounds.
[0160] It should be noted that the server runs a convergence monitor that performs two types of judgments simultaneously: In performance convergence detection, the F1-score is tested through a cross-campus joint validation set (including a subset of data features from all campuses). If the fluctuation is less than δ (δ = 0.005) for three consecutive rounds, convergence is marked; In distribution convergence detection, the norm of the KL divergence matrix of the prediction results of each node is calculated. When this value drops to 10% of the initial value, termination is triggered; If either condition is met, a termination command is sent to all nodes and the final model is archived.
[0161] In one possible implementation, the cross-campus data convergence condition determination in S7 includes:
[0162] S701, Central Server builds cross-campus verification dataset;
[0163] S702. After each round of aggregation, perform validation set predictions in the new global model;
[0164] S703. Calculate the F1-score coefficient of variation of the prediction results for three consecutive rounds:
[0165] S704. The first convergence condition is triggered when the coefficient of variation remains below 0.5%.
[0166] S705. Extract the distribution of prediction results for the local test set of each node;
[0167] S706. Calculate the Jensen-Shannon divergence matrix of the predicted distribution between nodes;
[0168] S707. The second convergence condition is triggered when the mean of the divergence matrix drops to 10% of the initial calculated value.
[0169] It should be noted that the central server constructs a cross-campus validation set (extracting 5% non-overlapping data from each node), tests the F1-score after each round of aggregation, and calculates the coefficient of variation c for three consecutive rounds. v =σ / μ. When c v Performance convergence is determined when the result is less than 0.005. Prediction results from the test set of each node are extracted synchronously, and the Jensen-Shannon divergence D between each pair of nodes is calculated. JS (P i ||P j After forming an N×N divergence matrix, the matrix mean μ is calculated. JS .when ( When the initial round mean is used, the distribution is considered to converge.
[0170] Furthermore, the validation set construction satisfies the class balance constraint: for classification tasks, it ensures that the proportion of samples in each class is less than 5% of the global distribution error; for regression tasks, the Kolmogorov-Smirnov test (significance level α = 0.05) is performed to ensure data distribution consistency.
[0171] The beneficial effects of the technical solutions provided in the embodiments of the present invention include at least the following:
[0172] (1) In this invention, a unified feature representation space across business domains is constructed through a dynamic feature alignment engine and adaptive nonlinear transformation. For heterogeneous data such as manufacturing sensor streams and financial transaction records, semantic matching and neural network encoding technologies are used to achieve seamless access to newly added industrial parks and automatic compensation of feature dimensions, significantly improving the efficiency of multi-source data fusion;
[0173] (2) In this invention, a triple protection system is constructed by combining homomorphic encrypted transmission, gradient sparsity, and dynamic anomaly filtering. While ensuring the confidentiality of updated parameters, malicious attacks are intercepted in real time through adaptive aggregation thresholds, and the communication load is reduced by selectively retaining gradient dimensions, effectively balancing security protection and system efficiency;
[0174] (3) In this invention, an innovative data-gradient dual-channel collaborative evaluation mechanism is designed, which quantifies node contributions by jointly using distribution similarity and model behavior. By combining the negative correlation strategy between noise injection and collaborative factors, high-quality nodes dominate the direction of model optimization, simultaneously improving the consistency of data distribution across multiple campuses and accelerating the model convergence process.
[0175] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
[0176] The following points need to be explained:
[0177] (1) The accompanying drawings of the embodiments of the present invention only involve the structures involved in the embodiments of the present invention. Other structures can refer to the general design.
[0178] (2) For clarity, the thickness of layers or regions is enlarged or reduced in the drawings used to describe embodiments of the present invention; that is, these drawings are not drawn to actual scale. It is understood that when an element such as a layer, film, region, or substrate is referred to as being “above” or “below” another element, the element may be “directly” located “above” or “below” the other element, or there may be intermediate elements.
[0179] (3) Where there is no conflict, the embodiments of the present invention and the features in the embodiments can be combined with each other to obtain new embodiments.
[0180] The above are merely specific embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. The scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A cross-park enterprise data collaborative analysis method based on federated learning, characterized in that, include: S1. Initialize the federated learning framework: The central server distributes the initial global model architecture and federated learning configuration parameters to the enterprise nodes in each park. S2. Localized data preprocessing: Each enterprise node performs feature alignment processing on its local private dataset to generate standardized feature vectors; S3. Collaboration Evaluation: Each node calculates the dynamic collaboration factor between its local dataset and the global data distribution; S4. Local Model Training: Adjust the local training strategy based on dynamic co-factors and update model parameters using the local dataset; S5. Security Parameter Aggregation: The central server collects model updates from each node through an encrypted channel and uses an aggregated offset threshold to filter valid updates. S6. Global Model Generation: The filtered model updates are weighted and aggregated to generate a new global model; S7. Iteration Termination Judgment: The process terminates when the global model meets the cross-park data convergence condition or reaches the maximum number of communication rounds.
2. The cross-park enterprise data collaborative analysis method based on federated learning according to claim 1, characterized in that, The feature alignment process in S2 includes: S201. Extract the statistical feature vector of the local dataset; S202. Perform dimension mapping with the global feature template issued by the central server; S203. Embed heterogeneous data into a unified feature space through nonlinear transformation.
3. The cross-park enterprise data collaborative analysis method based on federated learning according to claim 1, characterized in that, The dynamic synergy factor in S3 is calculated using the following formula: Among them, Γ i Let σ(·) be the dynamic cooperation factor of node i; σ(·) is the local dataset D. i With global data distribution D g The KL divergence contraction function is defined as σ(x) = e -βx cosim(·) represents the local model gradient. With global gradient The cosine similarity is η; the gradient importance coefficient has a value range of [0.2, 0.8]; β is the divergence sensitivity parameter with a value range of [1.0, 3.0].
4. The cross-park enterprise data collaborative analysis method based on federated learning according to claim 1, characterized in that, The aggregation offset threshold in S5 is calculated using the following formula: Where, τ t μ is the aggregation offset threshold for the t-th round of communication; t This is the update amount for all nodes in this round of model updates. The mean vector of ; κ is the standard deviation scaling factor, with a value range of [1.5, 2.5]; N is the total number of nodes participating in federated learning.
5. A cross-park enterprise data collaborative analysis method based on federated learning according to claim 1, characterized in that, The local training strategy adjustment in S4 includes the following steps: S401. Monitor the numerical relationship between the dynamic synergy factor and the preset threshold; S402. When the dynamic collaboration factor is below the first threshold, perform adversarial example generation: calculate the Jacobian matrix of the input data based on the current local model parameters; S403. Apply a perturbation to the original samples along the gradient direction of the loss function to generate an adversarial sample set; S404. Merge the adversarial sample set with the original training data to form an augmented dataset; S405. Add an L2 norm regularization term to the local loss function for the model parameters; S406. When the dynamic coordination factor is higher than the second threshold, increase the number of local training iterations proportionally. S407. Select the optimizer learning rate decay strategy based on the range of the dynamic co-factor value.
6. The cross-park enterprise data collaborative analysis method based on federated learning according to claim 1, characterized in that, The weighted aggregation of S6 includes the following steps: S601. Construct a node aggregation weight allocation function: using dynamic collaborative factors as input variables; S602. Standardize the node co-factor sequence to eliminate dimensional differences; S603. Use S-shaped function mapping to convert standardized co-factors into initial weights; S604. Normalize the initial weights so that the sum of the weights is 1; S605. Maintain the node participation status record table and mark the historical aggregation participation status of each node; S606, Detect isolated nodes that have not participated in aggregation for three consecutive rounds; S607. Extract the most recent valid model update of the isolated node as the compensation benchmark; S608: Integrate the compensation benchmark into the current round's global aggregation according to the attenuation coefficient.
7. A cross-park enterprise data collaborative analysis method based on federated learning according to claim 1, characterized in that, The aggregation of security parameters in S5 also includes: S501. Generate homomorphic encryption key pairs and distribute the public key to each node; S502, The node uses the public key encryption model to update parameters and generate ciphertext data packets; S503: After receiving the encrypted data packet, the central server performs gradient sparsity processing. S504. Calculate the L1 norm for each model update vector and sort them. S505: Retain the top K largest absolute gradient dimensions to form a sparse gradient vector; S506. Use the private key to decrypt the sparse gradient vector to obtain the plaintext update.
8. A cross-park enterprise data collaborative analysis method based on federated learning according to claim 1, characterized in that, The cross-park data convergence condition determination in S7 includes: S701, Central Server builds cross-campus verification dataset; S702. After each round of aggregation, perform validation set predictions in the new global model; S703. Calculate the F1-score coefficient of variation of the prediction results for three consecutive rounds: S704. The first convergence condition is triggered when the coefficient of variation remains below 0.5%. S705. Extract the distribution of prediction results for the local test set of each node; S706. Calculate the Jensen-Shannon divergence matrix of the predicted distribution between nodes; S707. The second convergence condition is triggered when the mean of the divergence matrix drops to 10% of the initial calculated value.
9. A cross-park enterprise data collaborative analysis method based on federated learning according to claim 3, characterized in that, The dynamic adjustment of the β value in the KL divergence compression function includes: S301, Central Server Maintenance Park Type - Parameter Mapping Table; S302, Receive the park industry classification code reported by the node; S303. When the code belongs to the manufacturing industry, the β baseline value is set to 2.5; S304. When the code belongs to the financial industry, the β benchmark value is set to 1.2; S305. The value should fluctuate within ±0.3 of the baseline value based on the size of the node data. S306. Recalibrate the β value every five communication cycles.
10. A cross-park enterprise data collaborative analysis method based on federated learning according to claim 1, characterized in that, There are also steps between S4 and S5: S4a, Calculate the noise standard deviation scaling factor: λ i =1 / 1+Γ i ; S4b, Obtain the preset reference noise level σ base ; S4c generates a Gaussian distribution. Random noise; S4d adds the noise vector element by element to the model update parameters; S4e records the noise injection amount for subsequent aggregation compensation calculations.
Citation Information
Cited By
Federal learning-based cross-border multi-party data calculation method and system, and storage medium
CN121352032A
Cross-border multi-party data computing method, system and storage medium based on federated learning
CN121352032B
Distributed data collaboration method and device based on federal control
CN121396643A
Multi-source data management method and device based on federal mechanism
CN121658689A
A multi-source data management method and device based on a federal mechanism
CN121658689B