Efficient federated learning aggregation method and device with robustness, verifiability and privacy

By adopting the two-party core principal component analysis and density space clustering algorithm for privacy protection in the federated learning system, combining linear homomorphic hashing and commitment solutions for verification, and using customer-level differential privacy enhancement technology, the federated learning system's robustness, verifiability and privacy under the strong threat model is solved, and efficient and secure federated learning aggregation is achieved.

CN120217429AActive Publication Date: 2025-06-27BEIJING UNIV OF POSTS & TELECOMM

Patent Information

Application Number
CN202510278084.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-10
Publication Date
2025-06-27
Estimated Expiration
2045-03-10

AI Technical Summary

Technical Problem

Existing federated learning systems are difficult to balance robustness, verifiability, privacy and efficiency under strong threat models, especially when faced with problems such as poisoning attacks, privacy leaks and reduced system efficiency.

Method used

An efficient federated learning aggregation method is proposed, through the analysis of the principal component of the privacy protection of two-party cores and a density spatial clustering algorithm that tolerate differential privacy, combined with linear homomorphic hashing and promise schemes for aggregation integrity verification, and adopts gradient lossless customer-level differential privacy enhancement technology.

Benefits of technology

The unity of robustness, verifiability and privacy under the strong threat model is achieved, which significantly improves system efficiency and avoids expensive encryption operations and high communication overhead.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120217429A_ABST
    Figure CN120217429A_ABST
Patent Text Reader

Abstract

The invention provides an efficient federated learning aggregation method and device with robustness, verifiability and privacy, and belongs to the field of distributed deep learning security. The method comprises the following specific steps: a trusted authorization mechanism sends verification preparation information to a client; the client locally trains a model, calculates auxiliary information and shares the auxiliary information to the double servers; under the gradient lossless differential privacy enhancement algorithm, the double servers convert Boolean sharing into lossless gradient arithmetic sharing, and client-level differential privacy is added at the same time; the double servers execute kernel principal component PCA dimensionality reduction and DP tolerance clustering; the double servers aggregate the gradient and send global update information to the client; the client completes lightweight enhanced verification information transmission and verifies aggregation integrity by using the received information to ensure strong verifiability; and if the verification is passed, returning to the step 2, and continuing the next round of training until the model converges or reaches a preset training round.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of distributed deep learning security, and relates to the security protection, result verification, privacy protection and model efficiency of federated learning. Specifically, it is an efficient federated learning aggregation method and device with robustness, verifiability and privacy. Background Art

[0002] Federated learning is a distributed machine learning architecture with privacy protection, which allows clients to collaboratively train a globally shared model without sharing local data. Nowadays, federated learning has been applied to privacy protection fields such as drug discovery, face recognition, and medical diagnosis.

[0003] FedAvg[1] and FedSGD[2] are common aggregation algorithms in federated learning, but they are vulnerable to poisoning attacks, thus affecting the global model. Label flipping attacks [3] and scaling attacks [4] change labels, while outlier attacks [5] and Gaussian noise attacks [6] modify gradients, resulting in model failure or poor performance. Backdoor attacks [4] implant triggers for specific data, and adaptive poisoning attacks [7] further enhance the concealment. In addition, these attacks can be combined with Sybil attacks [8] to control multiple clients to enhance the effect of poisoning attacks.

[0004] To defend against poisoning attacks, new robust aggregation algorithms have been proposed. For example, Dnc [9] is a spectral method based on singular value decomposition to detect and remove outliers. FLTrust [3] calculates the noise gradient through a clean small dataset to overcome the limitation of honest majority. Bulyan

[10] combines Krum and Trimmed Mean, and calculates the trimmed median of the remaining updates after filtering malicious updates.

[11] uses k-means to cluster and group gradient updates. However, there is a conflict between hiding gradient updates and removing malicious gradients by calculating update similarity. These methods may lead to privacy leakage [12, 13], making them vulnerable to membership inference attacks and attribute inference attacks

[14] . Although differential privacy can provide certain protection, attackers can still infer private data through repeated queries or analysis [12, 15, 16]. To better protect privacy, it is necessary to combine differential privacy with secure multi-party computing and homomorphic encryption technologies

[17] .

[0005] To balance the robustness and privacy of poisoning attacks, in a single-server architecture, RoFL

[18] enforces according to Bulletproofs and l ∝Defense. BREA

[19] filters malicious gradients through pairwise distances between updates and enhances privacy based on verifiable outlier detection. However, multi-round interactive verification significantly reduces system efficiency, and the assumption that the server is honest is unrealistic. In a distributed trust architecture, PPRAgg [6] protects privacy through homomorphic encryption and random obfuscation techniques and uses cosine similarity to evaluate the reputation of Byzantine nodes. However, this method has high computational overhead and cannot resist enhanced malicious gradients. Prio

[20] effectively filters malicious gradient updates based on non-interactive proof protocols (SNIPs) of arithmetic secret sharing and provides strong privacy protection when the client colludes with at most one server. However, the application of SNIPs results in low overall efficiency of Prio. Prio+

[21] replaces arithmetic secret sharing with boolean secret sharing, avoids using zero-knowledge proofs in Prio, and further restricts the operating ability of malicious clients while improving efficiency. However, similar to other federated learning systems (such as LSFL

[22] , SecFedDMC

[23] ), Prio+ only guarantees privacy in the semi-honest server model. Although subsequent research

[24] optimized SNIPs, their computational cost is still high. For this reason, ELSA [5] guarantees privacy and robustness under a strong threat model through a lightweight protocol and an l2-norm defense mechanism. However, a carefully designed local model can easily bypass this defense [25, 26, 27], resulting in training failure.

[0006] The above studies ignore the possibility of the server forging the aggregation result. In studies that take into account verifiability and privacy, VerifyNet

[27] combines bilinear pairs and homomorphic hash functions to achieve verifiable secure aggregation. VERIFL

[28] improves VerifyNet by introducing linear homomorphic hashing and an ambiguous commitment scheme, and can significantly reduce communication overhead while resisting collusion forgery by clients and servers. VFL

[29] uses Lagrange interpolation to verify the correctness of the aggregated gradients but ignores the problem of client dropout. VOSA

[30] designs a dynamic group management mechanism that can handle user dropout and support these users to participate in future training rounds. PVFL

[31] proposes a perturbation mechanism that can be offset during aggregation, achieving verification independent of disconnection and dimension. However, most integrity verification methods rely on masked routes, there is a conflict between confusing gradient updates and malicious gradient detection, and key negotiation also brings complex computational overhead. It can be seen that there is currently a gap in the field of federated learning research in building a robust, verifiable, private, and efficient federated learning secure aggregation framework to achieve multiple security requirements.

[0007] The references are as follows:

[0008] [1] Mcmahan H B, Moore E, Ramage D, et al. Communication-Efficient Learning of Deep Networks from Decentralized Data[J]. 2016. DOI: 10.48550 / arXiv.1602.05629.

[0009] [2] Mcmahan H B, Moore E, Ramage D, et al. Federated Learning of Deep Networks using Model Averaging[J]. 2016. DOI: 10.48550 / arXiv.1602.05629.

[0010] [3] Cao X, Fang M, Liu J, et al. FLTrust: Byzantine-robust Federated Learning via Trust Bootstrapping.[C] / / Network and Distributed System Security Symposium. Internet Society, 2021. DOI: 10.14722 / NDSS.2021.24434.

[0011] [4] Bagdasaryan E, Veit A, Hua Y, et al. How To Backdoor Federated Learning[J]. 2018. DOI: 10.48550 / arXiv.1807.00459.

[0012] [5] Rathee M, Shen C, Wagh S, et al. Elsa: Secure aggregation for federated learning with malicious actors[C] / / 2023 IEEE Symposium on Security and Privacy(SP). IEEE, 2023: 1961 - 1979.

[0013] [6] Ma X, LI Q, JIANG Q, et al. Byzantine-robust federated learning over non-IID data[J]. Journal on Communications, 2023, 44(6): 138-153.

[0014] [7] Qi X, Xie T, Li Y, et al. Revisiting the assumption of latent separability for backdoor defenses[C] / / The eleventh international conference on learning representations. 2023.

[0015] [8] Fung C, Yoon C J M, Beschastnikh I. The Limitations of Federated Learning in Sybil Settings[C] / / Recent Advances in Intrusion Detection. 2020.

[0016] [9] Shejwalkar V, Houmansadr A. Manipulating the byzantine: Optimizing model poisoning attacks and defenses for federated learning[C] / / NDSS. 2021.

[0017]

[10] Guerraoui R, Rouault S. The hidden vulnerability of distributed learning in byzantium[C] / / International conference on machine learning. PMLR, 2018: 3521-3530.

[0018]

[11] Yu L, Wu L. Towards byzantine-resilient federated learning via group-wise robust aggregation[J]. Federated Learning: Privacy and Incentive, 2020: 81-92.

[0019]

[12] Shokri R, Stronati M, Song C, et al. Membership inference attacks against machine learning models[C] / / 2017 IEEE symposium on security and privacy(SP). IEEE, 2017: 3-18.

[0020]

[13] Lee D D, Pham P, Largman Y, et al. Advances in neural information processing systems 22[J]. Tech Rep, 2009.

[0021]

[14] Hu K, Gong S, Zhang Q, et al. An overview of implementing security and privacy in federated learning[J]. Artificial Intelligence Review, 2024, 57(8): 204.

[0022]

[15] Yeom S, Giacomelli I, Fredrikson M, et al. Privacy risk in machine learning: Analyzing the connection to overfitting[C] / / 2018 IEEE 31st computer security foundations symposium(CSF). IEEE, 2018: 268-282.

[0023]

[16] Carlini N, Liu C, Erlingsson et al.The secret sharer:Evaluatingand testing unintended memorization in neural networks[C] / / 28th USENIXsecurity symposium(USENIX security 19).2019:267-284.

[0024]

[17] Domingo-Enrich C,Mroueh Y.Auditing Differential Privacy in HighDimensions with the Kernel Quantum R\'enyi Divergence[J].arXiv preprintarXiv:2205.13941,2022.

[0025]

[18] Burkhalter L,Lycklama H,Viand A,et al.Rofl:Attestable robustnessfor secure federated learning[J].arXiv preprint arXiv:2107.03311,2021,21.

[0026]

[19] So J,Güler B,Avestimehr A S.Byzantine-resilient secure federatedlearning[J].IEEE Journal on Selected Areas in Communications,2020,39(7):2168-2181.

[0027]

[20] Corrigan-Gibbs H,Boneh D.Prio:Private,robust,and scalablecomputation of aggregate statistics[C] / / 14th USENIX symposium on networkedsystems design and implementation(NSDI 17).2017:259-282.

[0028]

[21] Addanki S, Garbe K, Jaffe E, et al. Prio+: Privacy preserving aggregate statistics via boolean shares[C] / / International Conference on Security and Cryptography for Networks. Cham: Springer International Publishing, 2022: 516 - 539.

[0029]

[22] Zhang Z, Wu L, Ma C, et al. LSFL: A lightweight and secure federated learning scheme for edge computing[J]. IEEE Transactions on Information Forensics and Security, 2022, 18: 365 - 379.

[0030]

[23] Mu Xutong, Cheng Ke, Song Anxiao, et al. Privacy - preserving federated learning against Byzantine attacks[J]. Chinese Journal of Computers, 2024, 47(04): 842 - 861.

[0031]

[24] Boneh D, Boyle E, Corrigan - Gibbs H, et al. Zero - knowledge proofs on secret - shared data via fully linear PCPs[C] / / Annual International Cryptology Conference. Cham: Springer International Publishing, 2019: 67 - 97.

[0032]

[25] Xie C, Koyejo O, Gupta I. Fall of empires: Breaking byzantine - tolerant sgd by inner product manipulation[C] / / Uncertainty in Artificial Intelligence. PMLR, 2020: 261 - 270.

[0033]

[26] Tolpegin V,Truex S,Gursoy M E,et al.Data poisoning attacksagainst federated learning systems[C] / / Computer security–ESORICs 2020:25thEuropean symposium on research in computer security,ESORICs 2020,guildford,UK,September 14–18,2020,proceedings,part i 25.Springer InternationalPublishing,2020:480-501.

[0034]

[27] Zhang J,Chen B,Cheng X,et al.PoisonGAN:Generative poisoningattacks against federated learning in edge computing systems[J].IEEE Internetof Things Journal,2020,8(5):3310-3322.

[0035]

[27] Xu G,Li H,Liu S,et al.VerifyNet:Secure and verifiable federatedlearning[J].IEEE Transactions on Information Forensics and Security,2019,15:911-926.

[0036]

[28] Guo X,Liu Z,Li J,et al.V eri fl:Communication-efficient and fastverifiable aggregation for federated learning[J].IEEE Transactions onInformation Forensics and Security,2020,16:1736-1751.

[0037]

[29] Fu A, Zhang X, Xiong N, et al. VFL: A verifiable federated learning with privacy-preserving for big data in industrial IoT[J]. IEEE Transactions on Industrial Informatics, 2020, 18(5): 3316-3326.

[0038]

[30] Wang Y, Zhang A, Wu S, et al. VOSA: Verifiable and oblivious secure aggregation for privacy-preserving federated learning[J]. IEEE Transactions on Dependable and Secure Computing, 2022, 20(5): 3601-3616.

[0039]

[31] Zhou H, Yang G, Huang Y, et al. Privacy-preserving and verifiable federated learning framework for edge computing[J]. IEEE Transactions on Information Forensics and Security, 2022, 18: 565-580. Summary of the Invention

[0040] Aiming at the problem that it is difficult to balance multiple security requirements in the above-mentioned federated learning, the present invention proposes an efficient federated learning aggregation method and device with robustness, verifiability, and privacy. The specific design includes: (1) effectively filtering malicious gradients and resisting noise through privacy-preserving two-party kernel principal component analysis and two-party density space clustering algorithm tolerating differential privacy; (2) using the aggregation integrity verification under distributed trust, combined with linear homomorphic hashing and commitment scheme, to ensure that the aggregation result cannot be forged even when a server and a client collude; (3) adopting gradient-lossless client-level differential privacy enhancement, through lossless gradient sharing and differential privacy injection, to ensure gradient privacy protection; (4) the overall framework is lightweight, avoiding expensive encryption operations, with low verification communication overhead, and significantly improving the system efficiency. This framework efficiently realizes the unity of robustness, verifiability, and privacy under a strong threat model, and solves the deficiencies of existing research.

[0041] The efficient federated learning aggregation method and device with robustness, verifiability, and privacy are specifically divided into the following steps:

[0042] Step 1: Verification Preparation: The trusted authorization agency sends the linear homomorphic hash function preparation information LHHpp and the commitment scheme preparation information COMpp required for verification to the client.

[0043] Step 2: Training and Uploading: The client independently trains the model locally with local data, calculates the auxiliary information, and shares this information with the server through secret sharing.

[0044] Step 3: Client-level Differential Privacy Enhancement with Gradient Losslessness: Servers S (1) and S (0) convert the boolean sharing into lossless gradient arithmetic sharing. S (1) Adds client-level differential privacy to ensure strong privacy. The detailed process is as follows:

[0045] (1) Conversion of Lossless Arithmetic Sharing: The two servers use COT (correlated oblivious transfer)-based bit multiplication and bit combination to convert the boolean sharing of gradient updates into lossless arithmetic sharing. To reduce communication overhead, the server only needs to verify the auxiliary correlation information generated by the client. A malicious server may try to break the gradient privacy by sending malformed messages. Since each client performs independent operations in the first stage, honest clients can pre-simulate the interaction with the two servers. The present invention ensures malicious privacy security in the case of collusion by verifying the consistency between the one-time polling verification session record and the actual interaction.

[0046] (2) Server-based Client-level Differential Privacy Injection: Since the defense interaction is not independent among clients, a client-level differential privacy mechanism based on the semi-honest server S (1) is designed to ensure that even if the malicious server S (0) obtains the gradient update, it is still difficult to break the data privacy. First, the query function Q and its bounded sensitivity S Q are defined as follows:

[0047]

[0048] where, is the gradient obtained through batch data training on the dataset d i . U t represents the samples of all clients in the t-th round, and w t-1 is the global model weight in the (t - 1)-th round.

[0049] To solve the problem of calculating sensitivity without prior knowledge, norm clipping is required. Specifically, a threshold τ is set for updating, and the model update g is replaced by

[0050] Since the server S (1) cannot clip the plaintext gradient, each client must normalize the l2 norm in advance and cooperate with the server S (0) to verify whether the fixed update norm ||g||2 ≤ τ holds, ensuring that the uploaded gradient update is consistent with the result after norm clipping. Since the verification process is independent of the client, privacy is ensured through a single round of polling. The sensitivity St t and the noise scale σ t are expressed as:

[0051]

[0052] where η t is the learning rate and ε is the privacy budget - the smaller the privacy budget, the greater the noise.

[0053] Adding the noise to the locally shared update g (1) results in the local model g ′ that satisfies client-level differential privacy, and its expression is:

[0054]

[0055] Step 4: Two-party kernel principal component analysis with privacy protection and two-party density space clustering algorithm tolerant to differential privacy: The dual servers perform kernel principal component PCA dimensionality reduction and DP-tolerant clustering to ensure strong protection against malicious updates. The detailed process is as follows:

[0056] (1) Kernel matrix calculation: The Laplace kernel function is used to amplify and separate data differences in the high-dimensional space, efficiently achieving linear separability:

[0057]

[0058] where exp is the exponential function, σ is the standard deviation, and ||·|| represents the Manhattan distance between z i and z j . The calculation of k(z i , z j ) depends on SecMSB and SecMul. The server S (b) finally obtains the kernel matrix K (b) .

[0059] (2) Kernel matrix centering: Before performing eigenvalue decomposition, the kernel matrix needs to be normalized:

[0060]

[0061] Among them, 1 N is an N×N matrix, and each of its elements is equal to

[0062] (3) Eigenvalue decomposition: The kernel matrix is decomposed into a set of eigenvalues and eigenvectors. Using Theorem 1, the eigen decomposition is transformed into simpler secure operations, thus reducing the computational and communication overhead.

[0063] Theorem 1: If matrices A and B satisfy B = P -1 AP, then the two matrices have the same eigenvalues. If matrix A has an eigenvector v with the corresponding eigenvalue λ, then P -1 v is the eigenvector of matrix B corresponding to the eigenvalue λ.

[0064] Two-party kernel principal component analysis with privacy protection uses the arithmetic sharing P of the random matrix P (b) to blind the secret sharing of the kernel matrix so as to obtain After the server recovers it, calculate its eigenvalues and eigenvectors Based on Theorem 1, server S (b) securely calculates the eigenvalues of the kernel matrix and the eigenvector sharing

[0065] (4) Reconstruct the principal components: Before selecting the principal component vectors, perform normalization processing to obtain the new eigenvector sharing Select the top n eigenvectors in descending order as the principal components

[0066] (5) Data projection: To achieve dimensionality reduction, perform secure multiplication between the original local model and the principal components to obtain the dimensionality-reduced local model sharing

[0067] (6) Pre-adjust the neighborhood radius eps: Only in the first round of training, adaptively adjust eps based on the linear correlation between the neighborhood radius eps and the noise scale σ t The expression is eps = α·σ t +β, where α and β are constants.

[0068] The values of the constants α and β are determined based on the training results of the auxiliary dataset D by server S (1) in the offline phase. To obtain n sets of model updates W, server S(1) During the offline phase, the enhanced dataset D will be divided into n parts for training. The neighborhood radius eps is calculated by the ordered k-distance graph method, obtaining the model update sets from the clean model update set W and the model update set with noise level σ t scale Since σ t is known, the values of α and β can be approximated using the following equations:

[0069]

[0070] (7) On the defined neighborhood radius eps, the server S (1) and the server S (0) execute the density space clustering algorithm based on secure multi-party computation between them, removing malicious gradient updates from the gradient updates uploaded by all clients.

[0071] Step Five: Aggregation and Download: The server S (1) and the server S (0) aggregate the gradients and send the global update information to the clients.

[0072] Step Six: Aggregation Integrity Verification under Distributed Trust: After completing the lightweight enhanced verification information transmission, the clients use the received information to verify the aggregation integrity to ensure strong verifiability. The detailed process is as follows:

[0073] (1) Lightweight Enhanced Verification Transmission: Each client uses LHHpp and Compp obtained from the trusted authority (TA) to locally compute the verification information V (hash value h, commitment value c, random number r) of the gradient update g i and sends its arithmetic sharing to the servers S (0) and S (1) . After aggregation, the server sends the set of received verification information sharings to each client for verification. The length of the verification information is fixed and no longer depends on the dimension of the gradient update, and the communication overhead is limited to O(|C|), where C is the set of participating clients.

[0074] (h i,0 ,h i,1 ) = SS(HH.hash(g i ))

[0075] (c i,0 ,c i,1 ) = SS(COM.commit(h i ))

[0076] (r i,0 ,r i,1 ) = SS(r i )

[0077] (2) Integrity verification of the aggregation result: First, execute the Decommit process to verify whether the hash value of the gradient update has been tampered with.

[0078]

[0079] Second, to verify the aggregation result, it is necessary to associate it with the hash value and reduce it to a discrete logarithm problem. Taking federated average aggregation as an example (with equal weight coefficients, i.e., ), the verification can be completed by checking whether the product of the hash value of the aggregation result and the multi-party hash value h * (k) is the same. After passing the integrity verification, a new round of training can begin.

[0080]

[0081] Step Seven: Repeat Execution: If the verification passes, return to Step Two and continue the next round of training until the model converges or reaches the preset number of training rounds.

[0082] The advantages of the present invention are as follows:

[0083] (1) In response to the severe challenge of the difficulty in balancing multiple security requirements in federated learning, the present invention innovatively proposes a new aggregation framework with robustness, verifiability, privacy, and efficiency.

[0084] (2) In terms of strong robustness of aggregation, the present invention proposes new algorithms for two-party kernel principal component analysis with privacy protection and two-party density space clustering tolerant to differential privacy. Combining secure two-party computing with kernel principal component analysis significantly improves the robustness under high-dimensional gradient attacks and achieves efficient malicious gradient filtering by maintaining the quasi-clustering boundary.

[0085] (3) In terms of strong verifiability of aggregation, the present invention designs a new mechanism for verifying the integrity of aggregation under distributed trust, using linear homomorphic hashing and commitment schemes to prevent the server and client from colluding to forge the aggregation result while ensuring the efficiency of verification.

[0086] (4) In terms of strong privacy of individual gradients, the present invention proposes a new strategy for enhancing client-level differential privacy with gradient losslessness. By sharing gradients losslessly and injecting client-level differential privacy, it ensures that gradient privacy can be protected even if malicious clients collude with the server.

[0087] (5) The present invention adopts a lightweight design overall, avoiding expensive encryption operations and using linear homomorphic hashing functions to reduce communication overhead, significantly improving the system efficiency. Description of the Drawings

[0088] Figure 1 Schematic diagram of an efficient federated learning aggregation method and device with robustness, verifiability, and privacy for the present invention

[0089] Figure 2 Flowchart of an efficient federated learning aggregation method and device with robustness, verifiability, and privacy for the present invention

[0090] Figure 3 Schematic diagram of client-level differential privacy enhancement with lossless gradients for the present invention

[0091] Figure 4 Schematic diagram of two-party kernel principal component analysis with privacy protection for the present invention

[0092] Figure 5 Schematic diagram of two-party density space clustering with tolerance for differential privacy for the present invention

[0093] Figure 6 Schematic diagram of aggregation integrity verification under distributed trust for the present invention

[0094] Figure 7 Experimental graph of time overhead comparison of each verification method of the present invention under different scales of gradients and numbers of clients Detailed implementation manners

[0095] The following combines the accompanying drawings and specific embodiments to describe the implementation manners of the present invention in detail and clearly.

[0096] The present invention proposes an efficient federated learning aggregation method and device with robustness, verifiability, and privacy, having significant advantages in robustness, verifiability, efficiency, and privacy protection. Through two-party kernel principal component analysis with privacy protection and density space clustering algorithm with tolerance for differential privacy, it can effectively filter malicious gradients and resist noise interference. Combining linear homomorphic hashing and commitment scheme, it realizes aggregation integrity verification under distributed trust, and even if the server colludes with the client, the result cannot be forged. At the same time, it adopts client-level differential privacy enhancement technology with lossless gradients to ensure gradient privacy security. The overall design is lightweight, avoiding high-cost encryption operations and having low communication overhead, significantly improving the system efficiency. This method realizes the unity of robustness, verifiability, and privacy under a strong threat model, solving the deficiencies of existing research.

[0097] The specific steps of the efficient federated learning aggregation method and device with robustness, verifiability, and privacy under the implementation method are as Figure 1 and Figure 2 shown, including the following processes:

[0098] Step 1. Verification Preparation: The trusted authorization agency TA sends the linear homomorphic hash function preparation information LHHpp and the commitment scheme preparation information COMpp required for verification to the client c respectively.

[0099] Step 2. Training and Uploading: The client c independently trains the model locally with local data. Using norm clipping, the model update g is replaced by Meanwhile, the auxiliary information Aux is calculated and shared with the server S (1) and S (0) through secret sharing. The auxiliary information includes OT correlation, square correlation, Beaver triple correlation, and the session record information generated by the server for each client during the lossless arithmetic sharing conversion stage.

[0100] Step 3. Client-level Differential Privacy Enhancement with Gradient Losslessness ( Figure 3 ) : The servers S (1) and S (0) convert the boolean sharing into lossless gradient arithmetic sharing. S (1) Adds client-level differential privacy to ensure strong privacy. The detailed process is as follows:

[0101] (1) Lossless Arithmetic Sharing Conversion: The two servers use COT (correlated oblivious transfer)-based bit multiplication and bit combination to convert the boolean sharing of the gradient update into lossless arithmetic sharing. By verifying the consistency between the session record and the actual interaction through a one-time polling, malicious privacy security in the case of collusion is ensured.

[0102] (2) Server-based Client-level Differential Privacy Injection: The server S (1) collaborates with the server S (0) to verify whether the fixed update norm ||g||2 ≤ τ holds, ensuring that the uploaded gradient update is consistent with the result after norm clipping.

[0103] The server S (1) adds noise to the locally shared update g (1) to obtain the local model g ′ that satisfies client-level differential privacy, and its expression is:

[0104]

[0105] Step 4. Two-party Kernel Principal Component Analysis with Privacy Protection and Two-party Density Space Clustering Algorithm Tolerant to Differential Privacy ( Figure 4 and Figure 5 ) : The dual servers perform kernel principal component PCA dimensionality reduction and DP tolerant clustering to ensure strong protection against malicious updates. The detailed process is as follows:

[0106] (1) Kernel matrix calculation: Server S (0) and Server S (1) Based on secure multi-party computation and the Laplacian kernel function, amplify and separate data differences in high-dimensional space, and efficiently achieve linear separability:

[0107]

[0108] where exp is the exponential function, σ is the standard deviation, and ||·|| represents the Manhattan distance between z i and z j . The calculation of k(z i , z j ) depends on SecMSB and SecMul. Server S (b) Finally obtains the kernel matrix K (b) .

[0109] (2) Kernel matrix centering: Before performing eigenvalue decomposition, Server S (0) and Server S (1) Normalize the kernel matrix:

[0110]

[0111] where 1 N is an N×N matrix whose each element is equal to

[0112] (3) Eigenvalue decomposition: The kernel matrix is decomposed into a set of eigenvalues and eigenvectors. Use the arithmetic sharing P of the random matrix P (b) to blind the secret sharing of the kernel matrix to obtain Server S (0) and Server S (1) After recovery , Server S (1) Calculates its eigenvalues and eigenvectors Subsequently, Server S (b) Securely calculates the eigenvalues and eigenvector sharing of the kernel matrix

[0113] (4) Reconstruct the principal components: Before selecting the principal component vectors, Server S (0) and Server S (1) Perform normalization processing to obtain the new eigenvector sharing Select the top n eigenvectors in descending order as the principal components

[0114] ​(5) Data Projection: To achieve dimensionality reduction, server S (0) and server S (1) perform secure multiplication between the original local model and the principal components to obtain the shared local model after dimensionality reduction

[0115] (6) Pre-adjust the neighborhood radius eps: Only in the first round of training, adaptively adjust eps based on the linear correlation between the neighborhood radius eps and the noise scale σ t whose expression is eps = α·σ t +β, where α and β are constants, and are determined based on the training results of the auxiliary dataset D by server S (1) in the offline phase.

[0116] (7) On the defined neighborhood radius eps, server S (0) and server S (1) execute the density space clustering algorithm based on secure multi-party computation between them, and remove malicious gradient updates from the gradient updates uploaded by all clients.

[0117] Step Five: Aggregation and Download: Server S (0) and server S (1) aggregate the gradients and send the global updates to the clients participating in the aggregation.

[0118] Step Six: Aggregation Integrity Verification under Distributed Trust( Figure 6 ): After completing the lightweight enhanced verification information transmission, the clients use the received information to verify the aggregation integrity to ensure strong verifiability. The detailed process is as follows:

[0119] (1) Lightweight Enhanced Verification Transmission: Each client uses LHHpp and Compp obtained from the trusted authority (TA) to locally compute the verification information V (hash value h = HH.hash(g i ), commitment value c = COM.Commit(h i ), random number r) for the gradient update g i ), and arithmetic share its (h i,0 ,h i,1 ), (c i,0 ,c i,1 ), (r i,0 ,r i,1 ) to server S (0) and S (1) . After aggregation, the server sends the set of the received verification information shares to each client for verification.

[0120] (2) Aggregation result integrity verification: First, execute the COM.Decommit(c i , h i , r i ) process to determine whether it is 1 to verify whether the hash value of the gradient update has been tampered with.

[0121] Secondly, to verify the aggregation result, it is necessary to establish an association between it and the hash value and reduce it to a discrete logarithm problem. Taking federated average aggregation as an example (with equal weight coefficients, i.e., ), the verification can be completed by checking whether the hash value of the aggregation result is the same as the product of the multi-party hash values . After passing the integrity verification, a new round of training can begin.

[0122] Step Seven, Repeat Execution: If the verification passes, go back to Step Two and continue the next round of training until the model converges or reaches the preset number of training rounds.

[0123] Experimental Setup: The method is implemented based on PyTorch. The experimental environment is 2 NVIDIA RTX 2080Ti GPUs, Intel Xeon Gold 5218 CPUs, 64GB of RAM, and Ubuntu 18.04.6 LTS. When training on large-scale datasets, NVIDIA A100 GPUs are used. The datasets include MNIST, Fashion-MNIST, CIFAR10, ImageNet, and THUCNews, and non-IID data is generated through the Dirichlet distribution. MLP is used for MNIST, CNN for Fashion-MNIST, LeNet5 for CIFAR10, ResNet-50 for ImageNet, and ALBERT, RoBERTa, and BERT for THUCNews. The attack methods include LF-attack, GS-attack, Scaling attack, Sybil attack, and Outlier attack. The comparison baselines cover Bulyan, BREA, LSFL, PPRAgg, Prio, ELSA, etc. The evaluation metrics include DAR, FPR, TACC, PSR, SDO, and SVO. 100 clients, 45 malicious clients, 100 rounds of training, local Adam optimizer, learning rate 0.01, batch size 128, Dirichlet distribution parameter β = 5, and the verification scheme is based on technologies such as NIST P-256 elliptic curve, SHA-256 hash, and Shamir secret sharing are set in the experiment. In the experiment, the method of the present invention is abbreviated as EARVP.

[0124] Experimental Results:

[0125] (1) Malicious Gradient Defense Comparative Experiment: All attacks are combined with the Sybil attack, and the attack effect is exacerbated through multi-client collaboration. As shown in the following table, Prio and Prio+ rely on weak l p norm defense and are ineffective against most poisoning attacks; BREA selects similar models based on the Krum algorithm, and Bulyan combines Krum and Trimmed-Mean, with performance superior to BREA. However, their test accuracies (TACC) are 1.86% and 4.85% lower than that of EARVP respectively, and the poisoning success rate (PSR) is relatively high. ELSA uses the l2 norm threshold and performs well only under the Outlier attack, with insufficient overall robustness. LSFL adopts the K-nearest neighbor method and performs poorly under non-IID data. PPRAgg relies on unreliable reference gradients and is vulnerable to 45% of malicious gradients. SecFedDMC uses the RPCA+K-means algorithm and has a poor effect on the LF attack on CIFAR10, with a poisoning success rate nearly 10% higher. Although the accuracy rate in Outlier attack detection reaches 99.33%, the missed detection affects the overall test accuracy. In contrast, EARVP performs best on ImageNet and other datasets. By normalizing with the l_2 norm and gradient clipping, it reduces the poisoning success rate by 1.56% and improves the test accuracy by 3.61%, significantly outperforming SecFedDMC.

[0126] Table 1. Comparison of the effects of applying poisoning attack detection methods on different datasets and different attacks

[0127]

[0128]

[0129]

[0130] (2) Overall time overhead comparison experiment: To evaluate the defense overhead of the gradient-lossless client-level differential privacy enhancement, privacy-preserving two-party kernel principal component analysis, and two-party density space clustering algorithm tolerant to differential privacy in EARVP, we measured the time of key steps (excluding local training). As shown in the following table, Prio uses a non-interactive proof protocol of secret sharing and affine aggregation coding, which is both robust and privacy-preserving, but has the highest time overhead. PPPAgg relies on homomorphic encryption and obfuscation techniques, with complex calculations and the second-highest running time. LSFL, Prio+, ELSA, and SecFedDMC use secure multi-party computation for lightweight gradient detection. Among them, ELSA has the shortest running time due to its simple l2-norm defense, but the protection is weak; LSFL improves the detection effect by increasing the gradient comparison time; SecFedDMC enhances the detection through dimensionality reduction and clustering. EARVP is superior to SecFedDMC in terms of detection effect and has a similar running time, demonstrating high efficiency. Its computational overhead is feasible in practical applications and is suitable for wide deployment.

[0131] Table 2. Comparison of time overhead (SDO) of different security defense methods under different datasets

[0132]

[0133] Secondly, we further explored the time overhead (SVO) comparison between the aggregation integrity verification under distributed trust and other advanced aggregation result verification methods. As Figure 7 shown, whether increasing the gradient update dimension or the number of clients, the overall verification time overhead of the aggregation integrity verification under distributed trust is always lower than that of other algorithms. When |U| = 100 and |G| = 1000, its verification overhead is only 0.76% of VerifyNet, 2.42% of VERIFL, 3.13% of PVFL, and 58.89% of VOSA.

Claims

1. An efficient federated learning aggregation method and device with robustness, verifiability and privacy, characterized by: The following steps are taken: Step 1: Verification preparation: The trusted authority sends the linear homomorphic hash function preparation information LHHpp and commitment scheme preparation information COMpp required for verification to the client. Step 2: Training and uploading: The client independently trains the model locally using local data, calculates auxiliary information, and shares this information with the server through secret sharing. Step 3: Gradient lossless client-level differential privacy enhancement: Server S (1) and S (0) Convert Boolean sharing to lossless gradient arithmetic sharing. (1) Add client-level differential privacy to ensure strong privacy. The breakdown process is as follows: (1) Lossless arithmetic sharing conversion: The two servers use bit multiplication and bit combination based on COT (correlated oblivious transfer) to convert the Boolean sharing of gradient updates into lossless arithmetic sharing. In order to reduce communication overhead, the server only needs to verify the auxiliary related information generated by the client. A malicious server may attempt to destroy gradient privacy by sending malformed messages. Since each client performs independent operations in the first phase, honest clients can simulate interactions with the two servers in advance. The present invention verifies the consistency between session records and actual interactions through a one-time polling, ensuring malicious privacy security in the case of collusion. (2) Server-based client-level differential privacy injection: Since the defense interactions are not independent between clients, a semi-honest server S is designed. (1) The client-level differential privacy mechanism ensures that even a malicious server S (0) After obtaining the gradient update, it is still difficult to violate data privacy. First, query the function Q and its bounded sensitivity S Q The definition is as follows: in, is obtained by i The gradient obtained by batch training on U t represents the samples of all clients in round t, w t-1 is the global model weight at round t-1. In order to solve the problem of calculating sensitivity without prior knowledge, norm clipping is used. Specifically, a threshold τ is set for update, and the model update g is replaced by Since the server S (1) It is impossible to clip the plaintext gradient. Each client must normalize the l2 norm in advance and communicate with the server S (0) Collaborate to verify whether the fixed update norm ||g||2≤τ holds, and ensure that the uploaded gradient update is consistent with the norm clipped result. Since the verification process is independent of the client, privacy is ensured through a single poll. The sensitivity S of the tth round t and the noise scale σ t It is expressed as: Among them, η t is the learning rate and ε is the privacy budget - the smaller the privacy budget, the greater the noise. The noise Updates added to local share (1) In the above example, we can obtain the local model g that satisfies client-level differential privacy. ′ , whose expression is: Step 4: Privacy-preserving two-party kernel principal component analysis and differential privacy-tolerant two-party density space clustering algorithm: Dual servers perform kernel principal component PCA dimensionality reduction and DP-tolerant clustering to ensure strong protection against malicious updates. The subdivision process is as follows: (1) Kernel matrix calculation: Use the Laplace kernel function to amplify and separate data differences in high-dimensional space and efficiently achieve linear separability: Where exp is the exponential function, σ is the standard deviation, and ||·|| represents z i and z j The Manhattan distance between them. k(z i ,z j ) depends on SecMSB and SecMul. Server S (b) Finally, the kernel matrix K is obtained (b) . (2) Kernel matrix centering: Before eigenvalue decomposition, the kernel matrix needs to be normalized: Among them, 1 N is an N×N matrix, each element of which is equal to (3) Eigenvalue decomposition: The kernel matrix is ​​decomposed into a set of eigenvalues ​​and eigenvectors. Using Theorem 1, the eigendecomposition is transformed into a simpler security operation, thus reducing computational and communication overhead. Theorem 1: If matrices A and B satisfy B = P -1 AP, then the two matrices have the same eigenvalues. If the matrix A has an eigenvector v and the corresponding eigenvalue is λ, then P -1 v is the eigenvector of the matrix B corresponding to the eigenvalue λ. Privacy-preserving two-party kernel principal component analysis using arithmetic sharing of random matrices P (b) To blind the kernel matrix The secret sharing of Server Recovery Then, calculate its eigenvalue and the eigenvector Based on Theorem 1, server S (b) Safely Compute the Kernel Matrix The eigenvalue of Shared with feature vectors (4) Reconstructing the principal component: Before selecting the principal component vector, normalize it to obtain a new feature vector sharing Select the first n eigenvectors in descending order as principal components (5) Data projection: To achieve dimensionality reduction, a safe multiplication is performed between the original local model and the principal components. Get the local model sharing after dimensionality reduction (6) Pre-adjust the neighborhood radius eps: Only in the first round of training, based on the neighborhood radius eps and the noise scale σ t The linear correlation between them adaptively adjusts eps, which is expressed as eps = α·σ t +β, where α and β are constants. The values ​​of constants α and β are based on the server S (1) The training results of the auxiliary dataset D in the offline phase are determined. In order to obtain n sets of model updates W, the server S (1) The enhanced dataset D is divided into n parts for training in the offline phase. The neighborhood radius eps is calculated by the ordered k-distance graph method, and the updated set W of the clean model and the noise σ are obtained. t Model update set of scale Due to σ t is known, the values ​​of α and β can be approximated using the following equations: (7) In the defined neighborhood radius eps, server S (1) and server S (0) A density space clustering algorithm is executed based on secure multi-party computing to remove malicious gradient updates from the gradient updates uploaded by all clients. Step 5: Aggregation and downloading: Server S (1) and server S (0) Aggregate gradients and send global updates to clients. Step 6: Aggregate integrity verification under distributed trust: After completing the transmission of lightweight enhanced verification information, the client uses the received information to verify the aggregate integrity to ensure strong verifiability. The subdivision process is as follows: (1) Lightweight enhanced verification transmission: Each client uses LHHpp and Compp obtained from the trusted authority (TA) to locally calculate the gradient update g i The verification information V (hash value h, commitment value c, random number r) is sent to the server S. (0) and S (1) After aggregation, the server sends the shared set of received verification information to each client for verification. The length of the verification information is fixed and no longer depends on the dimension of the gradient update. The communication overhead is limited to O(|C|), where C is the set of participating clients. (h i,0 ,h i,1 )=SS(HH.hash(g i )) (c i,0 ,c i,1 )=SS(COM.Commit(h i )) (r i,0 ,r i,1 )=SS(r i ) (2) Aggregation result integrity verification: First, the Decommit process is performed to verify whether the hash value of the gradient update has been tampered with. Secondly, in order to verify the aggregation result, it is necessary to associate it with the hash value and reduce it to a discrete logarithm problem. Take the federated average aggregation as an example (with equal weight coefficients, i.e. ), verification can be done by checking the aggregation results The hash value and the multi-party hash value h * After the integrity verification, a new round of training can be started. Step 7. Repeat: If the verification passes, return to step 2 and continue the next round of training until the model converges or reaches the preset training round.

2. The method and device for efficient federated learning aggregation with robustness, verifiability and privacy as claimed in claim 1, characterized in that: The gradient lossless client-level differential privacy enhancement algorithm designed in step 3 ensures gradient privacy protection through lossless gradient sharing and differential privacy injection: Gradient-lossless client-level differential privacy enhancement: Server S (1) and S (0) Convert Boolean sharing to lossless gradient arithmetic sharing. (1) Add client-level differential privacy to ensure strong privacy. The breakdown process is as follows: (1) Lossless arithmetic sharing conversion: The two servers use bit multiplication and bit combination based on COT (correlated oblivious transfer) to convert the Boolean sharing of gradient updates into lossless arithmetic sharing. In order to reduce communication overhead, the server only needs to verify the auxiliary related information generated by the client. A malicious server may attempt to destroy gradient privacy by sending malformed messages. Since each client performs independent operations in the first phase, honest clients can simulate interactions with the two servers in advance. The present invention verifies the consistency between session records and actual interactions through a one-time polling, ensuring malicious privacy security in the case of collusion. (2) Server-based client-level differential privacy injection: Since the defense interactions are not independent between clients, a semi-honest server S is designed. (1) The client-level differential privacy mechanism ensures that even a malicious server S (0) After obtaining the gradient update, it is still difficult to violate data privacy. First, query the function Q and its bounded sensitivity S Q The definition is as follows: in, is obtained by i The gradient obtained by batch training on U t represents the samples of all clients in round t, w t-1 is the global model weight at round t-1. In order to solve the problem of calculating sensitivity without prior knowledge, norm clipping is used. Specifically, a threshold τ is set for update, and the model update g is replaced by Since the server S (1) It is impossible to clip the plaintext gradient. Each client must normalize the l2 norm in advance and communicate with the server S (0) Collaborate to verify whether the fixed update norm ||g||2≤τ holds, and ensure that the uploaded gradient update is consistent with the norm clipped result. Since the verification process is independent of the client, privacy is ensured through a single poll. The sensitivity S of the tth round t and the noise scale σ t It is expressed as: Among them, η t is the learning rate and ε is the privacy budget - the smaller the privacy budget, the greater the noise. The noise Add updates to local share (1) In the above example, we can obtain the local model g that satisfies client-level differential privacy. ′ , whose expression is: 。 3. The method and device for efficient federated learning aggregation with robustness, verifiability and privacy as claimed in claim 1, characterized in that: The privacy-preserving two-party kernel principal component analysis and the two-party density space clustering algorithm tolerant to differential privacy designed in step 4 can effectively filter malicious gradients and resist noise: Privacy-preserving two-party kernel principal component analysis and differential privacy-tolerant two-party density space clustering algorithm: Dual servers perform kernel principal component PCA dimensionality reduction and DP-tolerant clustering to ensure strong protection against malicious updates. The subdivision process is as follows: (1) Kernel matrix calculation: Use the Laplace kernel function to amplify and separate data differences in high-dimensional space and efficiently achieve linear separability: Where exp is the exponential function, σ is the standard deviation, and ||·|| represents z i and z j The Manhattan distance between them. k(z i ,z j ) depends on SecMSB and SecMul. Server S (b) Finally, the kernel matrix K is obtained (b) . (2) Kernel matrix centering: Before eigenvalue decomposition, the kernel matrix needs to be normalized: Among them, 1 N is an N×N matrix, each element of which is equal to (3) Eigenvalue decomposition: The kernel matrix is ​​decomposed into a set of eigenvalues ​​and eigenvectors. Using Theorem 1, the eigendecomposition is transformed into a simpler security operation, thus reducing computational and communication overhead. Theorem 1: If matrices A and B satisfy B = P -1 AP, then the two matrices have the same eigenvalues. If the matrix A has an eigenvector v and the corresponding eigenvalue is λ, then P -1 v is the eigenvector of the matrix B corresponding to the eigenvalue λ. Privacy-preserving two-party kernel principal component analysis using arithmetic sharing of random matrices P (b) To blind the kernel matrix The secret sharing of Server Recovery Then, calculate its eigenvalue and the eigenvector Based on Theorem 1, server S (b) Safely Compute the Kernel Matrix The eigenvalue of Shared with feature vectors (4) Reconstructing the principal component: Before selecting the principal component vector, normalize it to obtain a new feature vector sharing Select the first n eigenvectors in descending order as principal components (5) Data projection: To achieve dimensionality reduction, a safe multiplication is performed between the original local model and the principal components. Get the local model sharing after dimensionality reduction (6) Pre-adjust the neighborhood radius eps: Only in the first round of training, based on the neighborhood radius eps and the noise scale σ t The linear correlation between them adaptively adjusts eps, which is expressed as eps = α·σ t +β, where α and β are constants. The values ​​of constants α and β are based on the server S (1) The training results of the auxiliary dataset D in the offline phase are determined. In order to obtain n sets of model updates W, the server S (1) The enhanced dataset D is divided into n parts for training in the offline phase. The neighborhood radius eps is calculated by the ordered k-distance graph method, and the updated set W of the clean model and the noise σ are obtained. t Model update set of scale Due to σ t is known, the values ​​of α and β can be approximated using the following equations: (7) In the defined neighborhood radius eps, server S (1) and server S (0) A density space clustering algorithm is executed based on secure multi-party computing to remove malicious gradient updates from the gradient updates uploaded by all clients.

4. The method and device for efficient federated learning aggregation with robustness, verifiability and privacy as claimed in claim 1, characterized in that: The aggregation integrity verification algorithm under distributed trust designed in step 4, combined with linear homomorphic hashing and commitment scheme, ensures that the aggregation result cannot be forged even when a server and client collude: Aggregate integrity verification under distributed trust: After completing the transmission of lightweight enhanced verification information, the client uses the received information to verify the aggregate integrity to ensure strong verifiability. The subdivision process is as follows: (1) Lightweight enhanced verification transmission: Each client uses LHHpp and Compp obtained from the trusted authority (TA) to locally calculate the gradient update g i The verification information V (hash value h, commitment value c, random number r) is sent to the server S. (0) and S (1) After aggregation, the server sends the shared set of received verification information to each client for verification. The length of the verification information is fixed and no longer depends on the dimension of the gradient update. The communication overhead is limited to O(|C|), where C is the set of participating clients. (h i,0 ,h i,1 )=SS(HH.hash(g i )) (c i,0 ,c i,1 )=SS(COM.Commit(h i )) (r i,0 ,r i,1 )=SS(r i ) (2) Aggregation result integrity verification: First, the Decommit process is performed to verify whether the hash value of the gradient update has been tampered with. Secondly, in order to verify the aggregation result, it is necessary to associate it with the hash value and reduce it to a discrete logarithm problem. Take the federated average aggregation as an example (with equal weight coefficients, i.e. ), verification can be done by checking the aggregation results The hash value and the multi-party hash value h * After the integrity verification, a new round of training can be started. 。

Citation Information

Patent Citations

  • Robust federated learning privacy protection system based on block chain

    CN117113413A

  • Byzantine robust federated learning-oriented user data privacy protection system and method

    CN117395067A

  • Federal learning model based on local differential privacy and transfer learning technology

    CN119337970A

  • Federated learning method against backdoor attack

    WO2025039338A1

Cited By

  • Safety clustering federated learning method and system based on generalization parameters

    CN121257784A

  • Non-IID federated learning backdoor attack defense method and system and medium

    CN121262004A

  • A non-iid federated learning backdoor attack defense method, system and medium

    CN121262004B

  • Distributed model integrity verification method and system based on probability driving

    CN121367621A