A data-heterogeneous-oriented adaptive privacy protection personalized federated learning method
Patent Information
- Application Number
- CN202610782141.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-02
- Publication Date
- 2026-09-08
- Estimated Expiration
- 2046-06-02
AI Technical Summary
[0006]本发明的主要目的在于提供一种面向数据异构的自适应隐私保护个性化联邦学习方法,以解决现有差分隐私联邦学习在数据异构环境下因参数刚性划分和统一梯度裁剪所导致的梯度失真、收敛退化问题,通过参数敏感度感知、动态分层与定制约束优化,在严格的差分隐私预算下实现高精度、高鲁棒性的个性化联邦协同训练
[0039] Compared with existing technologies, this invention effectively alleviates the problems of convergence degradation, gradient information loss and structural rigidity in existing differential privacy federated learning when facing non-independent and identically distributed data, thereby significantly improving the model's robustness and providing an intelligent solution for achieving an excellent privacy-utility tradeoff in distributed machine learning in heterogeneous environments.
Smart Images

Figure CN122310586B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of distributed machine learning and privacy protection technology, specifically relating to an adaptive privacy-preserving personalized federated learning method and system based on dynamic parameter hierarchical and customized constraints, which is suitable for secure model training, privacy protection and efficient federated collaboration in heterogeneous data environments. Background Technology
[0002] With the development of artificial intelligence technology, federated learning, as a privacy-preserving distributed machine learning paradigm, is gaining increasing attention in scenarios such as healthcare, financial services, and intelligent edge computing. In practical applications, because the data of each participant (client) often exhibits heterogeneous characteristics of non-independent identically distributed (Non-IID), traditional single global models are insufficient to meet the personalized needs of different clients. Therefore, personalized federated learning has emerged.
[0003] However, due to the privacy risks such as data leakage associated with gradient interactions during the training process in federated learning, the introduction of differential privacy mechanisms to provide strict mathematical privacy guarantees is particularly crucial. Achieving high-fidelity personalized model transmission and efficient convergence under low privacy budgets plays a vital role in system reliability and the widespread application of federated learning.
[0004] In recent years, although researchers have proposed various personalized federated learning schemes incorporating differential privacy, existing methods still suffer from structural bottlenecks, making it difficult to achieve a good privacy-utility tradeoff in complex heterogeneous data. First, existing methods typically rely on rigid parameter partitioning strategies (such as fixed partitioning of shared and personalized layers). This static approach, which fixes parameter roles during initialization, lacks data-driven flexibility, cannot adapt to varying data heterogeneity, and is prone to parameter oscillations during training. Second, in the noise injection process for differential privacy, existing methods generally use a uniform gradient pruning threshold, ignoring the significant differences in information richness and noise sensitivity among different model parameters. This "one-size-fits-all" constraint not only leads to severe gradient distortion in high-sensitivity features but also introduces excessive invalid noise into low-sensitivity regions, easily disrupting the "high-information-density hubs" containing key local features in the model. This further exacerbates the dilemma of "structural rigidity and gradient distortion" in federated learning, leading to model convergence degradation.
[0005] Therefore, there is an urgent need for a federated learning scheme that combines parameter sensitivity awareness, dynamic hierarchical architecture and adaptive constraint optimization to solve the distortion problem caused by rigid parameter partitioning and uniform pruning, and to achieve highly robust and accurate personalized federated collaboration under a strict differential privacy budget. Summary of the Invention
[0006] The main objective of this invention is to provide an adaptive privacy-preserving personalized federated learning method for heterogeneous data, which solves the gradient distortion and convergence degradation problems caused by rigid parameter partitioning and uniform gradient pruning in existing differential privacy federated learning in heterogeneous data environments. Through parameter sensitivity awareness, dynamic layering and customized constraint optimization, it achieves high-precision and high-robust personalized federated collaborative training under strict differential privacy budget.
[0007] As a holistic concept, the system aims to fully integrate Empirical Fisher information perception, dynamic hierarchical pyramid structure, and parameter customization constraints to break through the limitations of traditional rigid binary divisions by precisely quantifying parameter sensitivity. The system utilizes Empirical Fisher information to construct a soft mask, smoothly transitioning parameters into three tiers: personalized, bridging, and lagging. Furthermore, decoupled adaptive boundary control and regularization penalties are applied to parameters at different sensitivity levels, minimizing pruning distortion while preserving key local features. Based on this, combined with differential privacy pruning and noise-adding mechanisms, the system can perform secure updates on the client side.
[0008] Based on the first main aspect of the present invention, an adaptive privacy-preserving personalized federated learning method for heterogeneous data is provided, which includes the following four computer-executed phase steps: local parameter sensitivity evaluation phase, variable hierarchical pyramid construction and dynamic hierarchical phase, adaptive optimization phase based on parameter customization constraints, and differential privacy noise injection and secure aggregation phase.
[0009] During the local parameter sensitivity assessment phase, the client uses local private data and the global model to calculate empirical Fisher information for the model parameters;
[0010] In the variable hierarchical pyramid construction and dynamic hierarchical stage, the experience Fisher information is hierarchically normalized and a soft mask is generated, and the parameters are dynamically and smoothly divided into three parameter echelons: personalized, bridging and lagging.
[0011] In the adaptive optimization phase based on customized parameter constraints, a decoupled loss function is constructed for the three parameter echelons, and customized boundary controls such as proximal constraints, amplitude constraints and hybrid adjustment are applied respectively to perform local model training.
[0012] During the differential privacy noise injection and secure aggregation phase, L2 norm pruning is performed on the local parameter updates participating in the sharing, and Gaussian differential privacy noise is injected. Then, the data is uploaded to the server to complete secure aggregation and a new round of global model distribution.
[0013] The above scheme breaks through the limitations of traditional rigid parameter partitioning and unified regularization, and achieves a good balance between privacy protection and model utility in heterogeneous data environments.
[0014] Optionally, during the local parameter sensitivity assessment phase, this includes inputting client data into the computer system. The local dataset and the local personalized model parameters retained from the previous round The client obtains the empirical Fisher information vector for each model parameter dimension by calculating the square of the local log-likelihood loss gradient. The specific formula is as follows: ,in This means squaring by elements. Indicates the model parameters Find the gradient. This is a local loss function.
[0015] The above scheme provides a specific method for calculating empirical Fisher information, namely, approximating it using the squared magnitude of the local log-likelihood loss gradient. This method avoids the costly calculation of the complete Fisher matrix, efficiently approximating the local curvature of the loss function using the squared gradient value. A larger curvature indicates greater sensitivity of the parameters to local Non-IID data. Therefore, it provides a lightweight and accurate method for measuring parameter sensitivity, laying a computationally viable foundation for subsequent dynamic layering.
[0016] Optionally, in the variable-layered pyramid construction and dynamic-layering stages, for the first... The parameters of the layer, and their normalized empirical Fisher values are calculated as follows: ,in The first element in the empirical Fisher information vector representing the model parameters calculated by the client. The first layer within the layer The original empirical Fisher values corresponding to each parameter dimension. These are the minimum and maximum empirical Fisher values for this layer, respectively;
[0017] Subsequently, according to the preset control threshold The first soft mask is generated using a piecewise linear function. .
[0018] In the above scheme, normalization eliminates the differences in gradient scales between different network layers, making the sensitivity across layers comparable. This ensures the fairness and stability of subsequent thresholding, preventing layering bias caused by different gradient value ranges in deep networks.
[0019] Furthermore, the first soft mask is generated using a piecewise linear function. In the following manner:
[0020] when hour, The corresponding parameters are divided into highly sensitive personalized parameters. ;when hour, The corresponding parameters are divided into bridging parameters of medium sensitivity. ;otherwise The corresponding parameters are divided into low-sensitivity global parameters. ;
[0021] And calculate the complementary mask. .
[0022] The parameters are divided into three echelons: personalized (mask 1), bridging (intermediate value), and global (mask 0). This continuous soft mask replaces the traditional binary rigid partition, allowing parameters to transition smoothly between the three echelons, which greatly alleviates the parameter oscillation problem and provides a precise basis for subsequent customized constraints.
[0023] Optionally, the variable hierarchical pyramid construction and dynamic hierarchical stage also includes adaptive fusion initialization of local and global parameters: ;in, These are the local personalized model parameters retained from the previous round. These are the global model parameters that the server obtained in the previous communication round and distributed in this round. Indicates client The initial model parameters are processed after the variable hierarchical pyramid construction and dynamic hierarchical stage before starting this round of local training.
[0024] The above scheme utilizes soft masks to achieve adaptive fusion initialization of local and global models. Specifically, it uses a mask to weightedly fuse the local personalized model parameters retained from the previous round with the global model received in the current round. This initialization method preserves the client's historical personalized knowledge while aligning global information, ensuring the model is in a fused state before training begins, accelerating convergence and improving personalization performance.
[0025] Optionally, in the adaptive optimization phase based on parameter customization constraints, let the personalized parameters be... The bridging parameters are The global parameter is In each iteration of stochastic gradient descent, the client minimizes the following three decoupled custom loss functions:
[0026] 1) For personalized parameters Proximal constraint loss: ;
[0027] 2) Regarding global parameters Amplitude constraint loss: ;
[0028] 3) Regarding bridging parameters Mixed adjustment loss: ;
[0029] in, For cross-entropy loss, These are the parameter reference values at the start of local training. For differential privacy pruning threshold, and The regularization coefficient is . and It is the dynamic adjustment factor derived from the first soft mask and the complementary mask.
[0030] The above scheme designs decoupled constraint loss functions for each of the three parameter tiers: personalized parameters employ proximal constraints to preserve local feature trajectories; global parameters introduce amplitude constraints to make their update norm approach the pruning threshold, thereby minimizing the distortion caused by subsequent differential privacy pruning; bridging parameters use a soft-mask-derived adjustment factor to mix the above two constraints, achieving a dynamic balance. This customized optimization overcomes the destruction of important features and excessive constraints on low-sensitivity regions caused by uniform regularization, thus improving the training stability and final accuracy of the model under heterogeneous data.
[0031] Optionally, during the differential privacy noise injection and security aggregation phase, the client completes... After local training for one epoch, extract the parameter update amount of the globally shared part. The formulas for performing clipping and noise addition are as follows: ,in For differential privacy pruning threshold, This is a Gaussian noise term. The Gaussian noise multiplier. For the identity matrix, for the client In the The amount of local parameter updates to be shared in a round-robin fashion. This is the safety update vector after adding noise.
[0032] The above scheme provides the specific operations for the differential privacy noise injection stage, namely, after performing L2 norm pruning on the shared update vector, the variance is injected based on the differential privacy pruning threshold. With Gaussian noise multiplier The Gaussian noise was determined jointly. This operation strictly implemented sensitivity limits and privacy protection, ensuring that the overall privacy leakage risk of the update vector was controllable, and is a core technical component for meeting client-level differential privacy requirements.
[0033] Furthermore, the differential privacy noise injection and security aggregation stage also includes: the client will... Send to the server, and the server collects the selected client set. After the update, via Aggregation generates a new global model, achieving strict Client-level differential privacy guarantees; among which, These are the global model parameters that the server obtained in the previous communication round and issued in this round.
[0034] The above solution provides a secure aggregation method on the server side, which involves randomly sampling clients and then weighting and averaging the noisy updates uploaded by each client to update the global model. It also points out that the entire process, combined with the Ruili differential privacy combination theorem, can guarantee strict security. Differential privacy. This aggregation method not only protects the data privacy of each client but also ensures the convergence of the global model, making it suitable for large-scale, heterogeneous, and privacy-critical federated learning scenarios.
[0035] Based on a third key aspect of the present invention, an electronic device is provided, comprising one or more processors;
[0036] Storage device for storing one or more programs;
[0037] When the one or more programs are executed by the one or more processors, the one or more processors implement the aforementioned adaptive privacy-preserving personalized federated learning method for data heterogeneity.
[0038] Based on a fourth key aspect of the present invention, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed, implements the aforementioned adaptive privacy-preserving personalized federated learning method for data heterogeneity.
[0039] Compared with existing technologies, this invention effectively alleviates the problems of convergence degradation, gradient information loss and structural rigidity in existing differential privacy federated learning when facing non-independent and identically distributed data, thereby significantly improving the model's robustness and providing an intelligent solution for achieving an excellent privacy-utility tradeoff in distributed machine learning in heterogeneous environments. Attached Figure Description
[0040] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, obtaining other drawings based on these drawings without creative effort still falls within the scope of the present invention.
[0041] Figure 1 The following is an execution flowchart of an embodiment of the present invention: an adaptive privacy-preserving personalized federated learning method for heterogeneous data.
[0042] Figure 2A schematic diagram of a privacy-preserving personalized federated learning end-cloud collaborative system architecture is shown in one embodiment.
[0043] Figure 3 A schematic diagram of the variable hierarchical pyramid (VHP) feature extraction and dynamic hierarchical principle provided in one embodiment is shown.
[0044] Figure 4 A schematic diagram of an adaptive optimization principle based on parameter customization constraints (PCC) is shown in one embodiment. Detailed Implementation
[0045] The preferred embodiments of the present invention will be described in detail below to provide a clearer understanding of the purpose, features, and advantages of the invention. It should be understood that the following embodiments are not intended to limit the scope of the invention, but are merely illustrative of the essential spirit of the technical solution of the invention.
[0046] In the following description, certain specific details are set forth for the purpose of illustrating various disclosed embodiments in order to provide a thorough understanding of the various disclosed embodiments. However, those skilled in the art will recognize that embodiments may be practiced without one or more of these specific details. In other instances, well-known techniques associated with the invention may not have been shown or described in detail to avoid unnecessarily obscuring the description of the embodiments.
[0047] Throughout this specification, references to "an embodiment" or "an embodiment" indicate that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment. Therefore, the appearance of "in an embodiment" or "an embodiment" in various places throughout the specification does not necessarily refer to the same embodiment. Furthermore, a particular feature, structure, or characteristic may be combined in any manner in one or more embodiments.
[0048] The following is a description of the specific meanings of technical terms, English abbreviations, and formula parameters that may be used in this invention:
[0049] Empirical Fisher Information (EFFI): This approximates the sensitivity and information content of parameters to local data distribution characteristics by calculating the squared magnitude of the gradient of the loss function with respect to the model parameters. A larger value indicates that the parameters are more sensitive to local non-independent and identically distributed (Non-IID) data and contain richer personalized information.
[0050] Variable Hierarchical Pyramid (VHP): A dynamic parameter hierarchical architecture that smoothly divides model parameters into three echelons—personalized, bridging, and global—by inputting normalized empirical Fisher information into a piecewise linear function to generate a soft mask, replacing the traditional rigid binary partition.
[0051] Soft mask: A continuous value between 0 and 1 (rather than a hard value of 0 or 1) used to characterize the degree to which a parameter belongs to different tiers (personalized, bridging, global). Soft masks allow parameters to transition smoothly between the three tiers, avoiding parameter oscillations.
[0052] Personalized parameters: These are parameters with high empirical Fisher values, highly sensitive to local data features, and are mainly protected by proximal constraints during training to preserve the client's local knowledge trajectory. They do not participate in or participate very little in global sharing.
[0053] Bridging parameters: Parameters with moderate empirical Fisher values, which are mixed with personalized and global constraints through adjustment factors α and β derived from soft masks, to achieve a dynamic balance between local feature preservation and global norm alignment, and a portion of their updates are shared globally.
[0054] Global Parameters: Parameters with low empirical Fisher values are less sensitive to local data. They are subject to explicit amplitude constraints, causing their update norm to converge toward the differential privacy pruning threshold. Their updates are fully involved in global sharing and aggregation.
[0055] Parameter-Customized Constraint (PCC): A local training optimization mechanism for three parameter echelons. It abandons the traditional unified regularization and applies proximal constraints to individual parameters, amplitude constraints to global parameters, and hybrid adjustments to bridging parameters, thereby achieving decoupled adaptive boundary control.
[0056] Proximal Constraint: This is a loss term applied to personalized parameters to limit the extent to which personalized parameters deviate from their local reference values, thereby preserving the local knowledge trajectory.
[0057] Magnitude Constraint: Applied to the loss term of the global parameters, it guides the update norm of the global parameters to approach the differential privacy pruning threshold, so as to minimize the distortion caused by subsequent L2 pruning operations.
[0058] Hybrid Regulation: This applies to the loss term of the bridging parameters and achieves the dual goals of personalized preservation and global norm alignment through a soft mask-derived factor.
[0059] Differential Privacy Clipping Threshold: The boundary value used for L2 norm clipping. It serves as both the target value for magnitude constraints in the local optimization stage and the upper limit of the clipping norm for sensitivity restrictions on the shared update vector before differential privacy noise addition.
[0060] L2 Norm Clipping: This operation scales the shared parameter update vector to ensure that its L2 norm does not exceed a preset threshold, thus limiting the sensitivity of updates from a single client.
[0061] Gaussian Differential Privacy Noise: Noise that satisfies differential privacy and is injected after cropping to control the privacy budget.
[0062] Rényi Differential Privacy (RDP): A technical framework for combinatorial privacy analysis. This invention utilizes the RDP combinatorial theorem to accurately track the cumulative privacy loss throughout the federated learning process.
[0063] like Figure 1 As shown, in one feasible embodiment, an adaptive privacy-preserving personalized federated learning method for heterogeneous data includes the following four computer-executed phases: local parameter sensitivity evaluation phase, variable hierarchical pyramid construction and dynamic hierarchical phase, adaptive optimization phase based on parameter customization constraints, and differential privacy noise injection and secure aggregation phase.
[0064] Step S110: In the local parameter sensitivity assessment phase, the client uses local private data and the global model to calculate the Empirical Fisher information of the model parameters, and quantifies the sensitivity of different parameters to data features by evaluating the local curvature approximation.
[0065] Step S120: In the variable hierarchical pyramid construction and dynamic hierarchical stage, the experience Fisher information is hierarchically normalized and a soft mask is generated. The parameters are dynamically and smoothly divided into three echelons: personalized, bridging and lagging, to achieve adaptive fusion initialization of local and global knowledge.
[0066] Step S130: In the adaptive optimization stage based on parameter customization constraints, uniform regularization is abandoned. Decoupled loss functions are constructed for the three parameter echelons, and customized boundary controls such as proximal constraints, amplitude constraints and hybrid adjustment are applied respectively to perform local model training.
[0067] Step S140: In the differential privacy noise injection and secure aggregation stage, L2 norm clipping is performed on the local parameter updates participating in the sharing, and Gaussian differential privacy noise is injected. Then, it is uploaded to the server to complete secure aggregation and a new round of global model distribution.
[0068] The following examples provide a further detailed description of the above steps.
[0069] In one embodiment, during the local parameter sensitivity evaluation phase, the client, in the early stages of local training iterations, utilizes local private data and the global model received from the server to extract empirical sensitivity features of the model parameters. First, the squared magnitude of the gradient of the loss function with respect to each model parameter is calculated, which is used as Empirical Fisher information to approximately evaluate local curvature, thereby accurately quantifying the sensitivity of different model parameters to the distribution characteristics of local data and the amount of information contained therein.
[0070] In the variable hierarchical pyramid construction and dynamic hierarchical stage, the extracted empirical Fisher features are used as input to construct the variable hierarchical pyramid. First, the empirical Fisher information of each layer's parameters is subjected to extreme value normalization. Then, a continuous soft mask is generated using a set threshold and a piecewise linear function. Based on the soft mask, the local model parameters are dynamically and smoothly divided into three parameter tiers: personalized parameters with high empirical Fisher values, global parameters with low empirical Fisher values, and bridging parameters with medium empirical Fisher values. The soft mask is then used to adaptively fuse and initialize the previous round of local personalized state with the current global model.
[0071] In the adaptive optimization phase based on parameter customization constraints, the system abandons traditional unified regularization for the three dynamically divided parameter tiers mentioned above, and instead performs local model training based on parameter customization constraints. By constructing decoupled loss functions for individual parameters, global parameters, and bridging parameters respectively, adaptive boundary control is implemented: individual parameters retain their local feature optimization trajectory; global parameters are subject to explicit magnitude constraints, causing their update norm to approach a preset pruning threshold to minimize pruning distortion; and bridging parameters are adjusted using a soft mask-based adjustment factor to dynamically balance the need for local feature preservation and global norm alignment.
[0072] In the differential privacy noise injection and secure aggregation phase, the system extracts the parameter update vectors (i.e., the update parts of global parameters and bridging parameters) that need to be shared globally after local joint optimization, and performs sensitivity restrictions and differential privacy processing on the client side. First, L2 norm pruning is performed on the parameter update vectors to ensure that their maximum update magnitude does not exceed a preset pruning threshold. Then, rigorously calibrated Gaussian differential privacy noise is injected into the pruned update vectors. Finally, the noisy update vectors are uploaded to the server, where secure aggregation and the distribution of the next round of global models are completed, ensuring high-precision personalized federated collaboration under strict privacy budgets.
[0073] As a possible implementation, in step S110, during the local parameter sensitivity evaluation phase, the input is: client... Local private dataset and the initial global model received from the server in the current round. Output: Empirical Fisher information vector for each model parameter dimension. .
[0074] At this stage, the system does not directly calculate the costly true Fisher information matrix, but instead uses the squared magnitude of the local log-likelihood loss gradient for an empirical approximation.
[0075] Specifically, for each parameter dimension, the calculation formula is as follows: ,in This means squaring by elements. This is a local loss function. This empirical sensitivity index effectively approximates the local curvature of the loss function plane. The larger the curvature (i.e., the empirical Fisher value), the more sensitive the parameter is to the characteristics of local non-independent and identically distributed (Non-IID) data, and the richer the amount of local personalized information it contains.
[0076] As a possible implementation, step S120 is as follows:
[0077] In the variable hierarchical pyramid construction and dynamic hierarchical stages, the input is: the empirical Fisher information vector. The global model parameters obtained by the server in the previous communication round and distributed in this round. and the local personalized model parameters retained from the previous round Output: Locally initialized model and soft mask .
[0078] First, to eliminate the difference in gradient scale between different network layers, the gradient scale of the first layer is... Local extremum normalization of the empirical Fisher values of the layer: ,in and These are the minimum and maximum empirical Fisher values for this layer, respectively.
[0079] Subsequently, control thresholds were set. Soft masks are constructed using piecewise linear functions. :when hour, (Classified into highly sensitive personalized parameters) );when hour, (Bridging parameters classified as medium sensitivity) );otherwise (Global parameters divided into low-sensitivity categories) The complementary mask is calculated as follows: .
[0080] Finally, adaptive fusion initialization of local and global parameters is achieved through soft masking: .
[0081] As a possible implementation, step S130 is as follows:
[0082] In the adaptive optimization phase based on parameter customization constraints, the input is: the initialized parameters. The corresponding pre-training reference value and differential privacy pruning threshold Output: Updated model parameters after local training. When performing stochastic gradient descent (SGD) updates on the client side, the system abandons a uniform regularization term and designs decoupled constraint loss functions for parameters in different echelons:
[0083] For personalized parameters Proximal constraints are used to preserve local knowledge trajectories: ;
[0084] For global parameters An explicit amplitude constraint is introduced to bring the update norm closer to the clipping threshold, thereby minimizing subsequent clipping distortion. ;
[0085] For bridging parameters Adjustment factor generated based on soft mask and (in ), perform mixed adjustment loss: The above. For cross-entropy loss, and This is the regularization hyperparameter.
[0086] As a possible implementation, step S140 is as follows:
[0087] During the differential privacy noise injection and security aggregation phase, the input is: the amount of local parameter updates to be shared by the client. (That is, the update vector that participates in the globally shared data after excluding personalized parameters), pruning threshold and Gaussian noise multiplier Output: A new round of models after global aggregation. .
[0088] First, the client-side sensitivity limit for shared updates is implemented (L2 norm pruning), and Gaussian noise, strictly calibrated according to the privacy budget, is injected: ,in It is an identity matrix.
[0089] After adding noise, the client will safely update the vector. Uploaded to the server. The server randomly samples a portion of the clients to form a set. Perform a secure aggregation to update the global model: .
[0090] This mechanism, combined with the Rényi Differential Privacy (RDP) combinatorial theorem, ensures that the entire federated learning process meets strict requirements. Client-level differential privacy enables secure and high-precision collaborative training in highly heterogeneous data environments.
[0091] The following is in conjunction with the appendix Figure 2-4 The present invention will be further described below.
[0092] Figure 2 A schematic diagram of the privacy-preserving personalized federated learning cloud collaboration system architecture provided for embodiments of the present invention. This diagram illustrates the overall macro-level workflow of the system:
[0093] ① The server distributes the new global model to each client (e.g., client-side applications). and client );
[0094] ②The client uses local data to perform local training based on the received global model, generating an updated local model;
[0095] ③ The client extracts the parameter update amount that needs to be shared globally and performs a clipping operation on it to limit the sensitivity of parameter updates;
[0096] ④ Subsequently, calibrated differential privacy noise (adding noise) is injected into the cropped parameters to prevent privacy leakage;
[0097] ⑤ Each client uploads the noise-added security parameters to the server. The server then securely aggregates the collected parameters to generate a new global model for the next round. This architecture achieves collaborative training of federated models with strict differential privacy protection while ensuring that data from each node does not leave its local machine.
[0098] Figure 3 This diagram illustrates the principle of Variable Hierarchical Pyramid (VHP) feature extraction and dynamic hierarchical layering provided in this embodiment of the invention. The diagram demonstrates the mechanism by which local parameters are dynamically layered:
[0099] First, empirical Fisher information is calculated based on the local model (calculation). This allows for precise quantification of the sensitivity and information content of different parameters to local data.
[0100] Subsequently, the experience Fisher information is input into the Variable Hierarchical Pyramid (VHP) module, combined with the set threshold. Perform regularization (normalization) processing;
[0101] Finally, a soft mask is generated through parameter selection, smoothly and dynamically dividing the parameters into three hierarchical tiers: the global parameters represented by the topmost circle (corresponding to the mask). The bridging parameters (corresponding to the mask) are represented by the semicircle of the middle layer. ), and personalized parameters represented by the lower square (corresponding mask) ).
[0102] This mechanism breaks the limitations of traditional rigid partitioning and provides a structural foundation for a smooth transition in subsequent adaptive optimization.
[0103] Figure 4 This diagram illustrates the adaptive optimization principle based on customizable parameter constraints (PCC) provided in this embodiment of the invention. It shows how, after completing the dynamic layering of the Virtual Hierarchical Structure (VHP), the system applies a decoupling constraint mechanism to parameters at different levels: the parameter echelons (upper, middle, and lower layers) output by the VHP module with soft mask attributes are correspondingly fed into the customizable parameter constraints (PCC) module for processing.
[0104] Apply a penalty to the set of personalized parameters (blocks). (Proximal constraints) are implemented to preserve local features from being violated.
[0105] Apply a penalty to the global parameter set (circles). (Amplitude constraint, causing its update to be directed towards the pruning threshold) (by bringing them closer together) helps reduce cutting distortion caused by subsequent operations;
[0106] Apply a penalty to the bridging parameter set (semicircle). (Right now This enables a hybrid dynamic adjustment that preserves local personalization while aligning with the global norm.
[0107] Ultimately, these three customized boundary controls work together to guide the adaptive training process of the local model.
[0108] To further verify the effectiveness and reliability of the method proposed in this invention, the following is a theoretical proof analysis of the invention in terms of differential privacy guarantee and model convergence:
[0109] 1. Analysis of Differential Privacy Theory Guarantee
[0110] In step S140 of the present invention, the sensitivity clipping and noise addition mechanism executed by the client satisfies strict client-level requirements. - Differential privacy. Its theoretical proof is as follows:
[0111] In this invention, Rényi Differential Privacy (RDP) is used for compact combinatorial analysis. For any client... Its globally shared update vector Crop threshold Limitations, i.e. Given adjacent datasets (differences of only one client data point), the L2 sensitivity of this update operation is strictly defined as follows: .
[0112] According to the RDP theorem of the Gaussian mechanism, for a sensitivity of... The function, injection scale is Gaussian noise, satisfying -RDP, where .
[0113] Because RDP possesses linear combinatorial properties, after... After one round of global communication iterations, the total privacy budget consumption is: Ultimately, according to RDP to the standard... -DP's transformation theorem, the method of this invention is applicable to any It can meet strict requirements -DP, where the optimal The value is:
[0114]
[0115] This proof demonstrates that the present invention, through precise cropping and noise addition, can mathematically provide a strict guarantee against privacy breaches.
[0116] 2. Model Convergence and Technical Performance Analysis
[0117] Existing technologies using uniform gradient clipping often lead to convergence difficulties, while this invention significantly reduces convergence error by introducing customizable parameter constraints (PCC). Its convergence theoretical analysis is as follows:
[0118] Suppose that the heterogeneity of non-independent and identically distributed (Non-IID) data leads to an upper bound on the difference between the client-side gradient and the global gradient. After Round global communication and After local iteration, the lower bound of the squared average gradient norm of the globally shared parameters in this invention is constrained by the following three components:
[0119]
[0120] The first term is the optimization initialization error, the second term is the stochastic gradient variance, and the third term... The first term is Client Drift caused by data heterogeneity, and the fourth term is the NoiseFloor, which introduces error margins due to differential privacy noise. The dimension for shared parameters. Indicates the first Globally shared parameters after round-robin global communication This represents the difference between the initial loss and the optimal loss. This represents the learning rate of the client-side local stochastic gradient descent. This represents the upper bound of the variance of the client-side local stochastic gradient. This indicates the number of clients participating in the aggregation in each round of global communication. This represents the Lipschitz smoothing constant of the loss function.
[0121] In this invention, the custom constraint module applies an explicit magnitude constraint loss to the globally shared parameters. From an optimization theory perspective, this is equivalent to introducing a proximal operator, explicitly limiting the magnitude of the local update's deviation from the reference state. Mathematically, this significantly reduces the effective heterogeneity constant. The value of is determined. Therefore, compared to traditional differential privacy federated learning methods, this invention not only limits noise error to a smaller dimension of shared parameters, but also... Meanwhile, the heterogeneity drift term was actively reduced through constraint optimization. This ensures that the model has a faster convergence speed and higher final accuracy while injecting the same differential privacy noise.
[0122] The adaptive privacy-preserving personalized federated learning method of this invention can be widely applied in multiple industrial and technological fields, such as healthcare, financial services, and intelligent edge computing and the Internet of Things.
[0123] In the healthcare field, applications can revolve around collaborative training of medical models across institutions. For example, multiple hospitals may want to jointly train a deep learning model for early disease screening, such as a tumor detection network based on medical imaging, without requiring raw patient data. Due to differences in equipment, patient populations, and data acquisition protocols among hospitals, local data typically exhibits non-independent identically distributed (Non-IID) characteristics.
[0124] After adopting the solution of this invention, each hospital, as a client, first uses local image data to calculate Fisher information of model parameters, and automatically identifies personalized parameters (such as high-level texture feature extraction layer) that are sensitive to local lesion features and global parameters corresponding to common features, such as low-level edge detection layer.
[0125] After generating a soft mask using a variable hierarchical pyramid, the system applies proximal constraints to personalized parameters to preserve the hospital's distinctive diagnostic patterns, applies amplitude constraints to global parameters to make their update norm approach a preset differential privacy pruning threshold, and dynamically adjusts the bridging parameters between the two.
[0126] Subsequently, the updates to the global parameters and bridging parameters are trimmed and Gaussian noise is injected before being uploaded to the central server. The server then securely aggregates the data and distributes it to each hospital.
[0127] Under strict differential privacy protection, this process not only prevents the distortion of key features due to excessive noise, but also ensures that the model can absorb the common knowledge of multi-center data, ultimately generating a personalized screening model with high accuracy and privacy protection for local data from each hospital.
[0128] In the financial services sector, this invention can be used by multiple banks or credit institutions to jointly build anti-fraud models. Different institutions exhibit significant differences in customer transaction behavior and fraud patterns (for example, some institutions primarily engage in small online transactions, while others focus on large cross-border transactions), resulting in highly heterogeneous data. Traditional federated learning, if using a unified global model, struggles to adapt to both types of patterns simultaneously.
[0129] This invention allows institutions to automatically quantify the sensitivity of different parameters (such as transaction frequency feature weights and amount fluctuation feature weights) to local fraud patterns during local training using empirical Fisher information. Low-sensitivity parameters are shared as global parameters to learn common fraud indicators across institutions. High-sensitivity parameters are retained as personalized parameters in local proximal constraints to avoid interference from global noise.
[0130] Crucially, the custom constraint mechanism of this invention is specifically designed to impose amplitude constraints on global parameters that converge towards the pruning threshold. This significantly reduces the accuracy loss caused by subsequent differential privacy noise addition, enabling the anti-fraud model to maintain a high recall rate for suspicious transactions even with a limited privacy budget.
[0131] Meanwhile, the adaptive adjustment of bridging parameters ensures that emerging cross-regional fraud patterns can be smoothly integrated into the global knowledge. The entire training process does not require any original customer data to leave the local machine, and the server only collects updates to the parameters after adding noise, fully meeting the legal compliance requirements of the financial industry for customer privacy protection.
[0132] In intelligent edge computing and IoT scenarios, this invention can be deployed in federated learning systems composed of a large number of terminal devices (such as smartphones, industrial sensors, and autonomous vehicles). These devices are typically limited by communication bandwidth, battery life, and the spatiotemporal variations of local data distribution.
[0133] This invention enables devices to adjust the soft mask assignment of each model parameter in real time based on changes in the current local data distribution through the dynamic layering capability of the variable layered pyramid: when the device environment undergoes a sudden change (such as an autonomous vehicle changing from sunny weather to rainy or foggy weather), the Fisher information of relevant parameters increases rapidly, and the system automatically classifies them into the personalized or bridging echelon to enhance local adaptability; when the environment tends to stabilize, some parameters smoothly return to the global echelon, reducing the number of uploaded update dimensions to save bandwidth.
[0134] Furthermore, this invention performs targeted norm constraint pruning on the shared parameters before adding noise in differential privacy, avoiding gradient vanishing due to over-pruning. This is especially important for control models with extremely low noise tolerance (such as trajectory prediction networks in autonomous driving).
[0135] This invention enables edge devices to perform high-frequency federated collaborative training with extremely low privacy budgets, while maintaining the model's rapid response to dynamic environments. It provides a practical and feasible privacy-preserving machine learning solution for real-time decision-making systems such as smart cities, Industry 4.0, and the Internet of Vehicles.
[0136] It should be understood that the program code used to implement the methods of the present invention can be written in any combination of one or more programming languages. This program code can be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing device, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code can be executed entirely on the machine, partially on the machine, as a standalone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0137] The acquisition, storage, and application of user personal information involved in the technical solution of this invention all comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0138] It should be understood that the various forms of processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this invention does not impose any limitations on them.
[0139] The technical terms, principles, or means related to the technical solutions of the present invention mentioned in the above embodiments, which are not described in detail above, are all well-known technologies or common practices that are known to those skilled in the art.
[0140] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely illustrative of the principles of the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the present invention as claimed. The scope of protection of this invention is defined by the appended claims and their equivalents.
Claims
1. An adaptive privacy-preserving personalized federated learning method for heterogeneous data, characterized in that, The method includes the following four computer-executed phases: local parameter sensitivity assessment phase, variable hierarchical pyramid construction and dynamic hierarchical phase, adaptive optimization phase based on parameter customization constraints, and differential privacy noise injection and secure aggregation phase. During the local parameter sensitivity assessment phase, the client uses local private data and the global model to calculate empirical Fisher information for the model parameters; In the variable hierarchical pyramid construction and dynamic hierarchical stage, the experience Fisher information is hierarchically normalized and a soft mask is generated, and the parameters are dynamically and smoothly divided into three parameter echelons: personalized, bridging and lagging. In the adaptive optimization phase based on parameter customization constraints, a decoupled loss function is constructed for the three parameter echelons, and proximal constraints, amplitude constraints and hybrid adjustment customized boundary controls are applied respectively to perform local model training. In the differential privacy noise injection and secure aggregation stage, L2 norm pruning is performed on the local parameter updates participating in the sharing and Gaussian differential privacy noise is injected. Then, it is uploaded to the server to complete secure aggregation and a new round of global model distribution. In the variable-layered pyramid construction and dynamic-layering stages, for the first... The parameters of the layer, whose empirical Fisher values are denoted as Its normalized empirical Fisher value is calculated as follows: ,in, These are the minimum and maximum empirical Fisher values for this layer, respectively; Subsequently, according to the preset control threshold The first soft mask is generated using a piecewise linear function. ; The first soft mask is generated using a piecewise linear function. In the following manner: when hour, The corresponding parameters are divided into highly sensitive personalized parameters. ;when hour, The corresponding parameters are divided into bridging parameters of medium sensitivity. ;otherwise The corresponding parameters are divided into low-sensitivity global parameters. ; And calculate the complementary mask. ; In the adaptive optimization phase based on customized parameter constraints, let the customized parameter be... The bridging parameters are The global parameter is In each iteration of stochastic gradient descent, the client minimizes the following three decoupled custom loss functions: 1) For personalized parameters Proximal constraint loss: ; 2) Regarding global parameters Amplitude constraint loss: ; 3) Regarding bridging parameters Mixed adjustment loss: ; in, For cross-entropy loss, These are the parameter reference values at the start of local training. For differential privacy pruning threshold, and The regularization coefficient is . and It is the dynamic adjustment factor derived from the first soft mask and the complementary mask.
2. The adaptive privacy-preserving personalized federated learning method for heterogeneous data as described in claim 1, characterized in that, The local parameter sensitivity assessment phase includes inputting client data into the computer system. The local dataset and the local personalized model parameters retained from the previous round The client obtains the empirical Fisher information vector for each model parameter dimension by calculating the square of the local log-likelihood loss gradient. The specific formula is as follows: ,in This means squaring by elements. Indicates the model parameters Find the gradient. This is a local loss function.
3. The adaptive privacy-preserving personalized federated learning method for heterogeneous data as described in claim 1, characterized in that, The variable hierarchical pyramid construction and dynamic hierarchical stage also includes adaptive fusion initialization of local and global parameters: ;in, These are the local personalized model parameters retained from the previous round. These are the global model parameters that the server obtained in the previous communication round and distributed in this round. Indicates the client The initial model parameters are processed after the variable hierarchical pyramid construction and dynamic hierarchical stage before starting this round of local training.
4. The adaptive privacy-preserving personalized federated learning method for heterogeneous data as described in claim 1, characterized in that, During the differential privacy noise injection and security aggregation phase, the client completes... After local training for one epoch, extract the parameter update amount of the globally shared part. The formulas for performing clipping and noise addition are as follows: ,in For differential privacy pruning threshold, This is a Gaussian noise term. The Gaussian noise multiplier. For the identity matrix, for the client In the The amount of local parameter updates to be shared in a round-robin fashion. This is the safety update vector after adding noise.
5. The adaptive privacy-preserving personalized federated learning method for heterogeneous data as described in claim 4, characterized in that, The differential privacy noise injection and security aggregation stage also includes: the client will... Send to the server, and the server collects the selected client set. After the update, via Aggregation generates a new global model, achieving strict Client-level differential privacy guarantees; among which, These are the global model parameters that the server obtained in the previous communication round and issued in this round.
6. An electronic device, characterized in that, Includes one or more processors; Storage device for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the adaptive privacy-preserving personalized federated learning method for data heterogeneity as described in any one of claims 1-5.
7. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed, the computer program implements the adaptive privacy-preserving personalized federated learning method for data heterogeneity as described in any one of claims 1-5.
Citation Information
Patent Citations
Layered differential privacy federal learning method based on Fisher information matrix
CN118332601A
Internet of vehicles data privacy protection method based on block chain and ternary federated learning
CN119046973A