Block chain and fully homomorphic encryption-based edge federated learning privacy protection method
By adopting blockchain and fully homomorphic encryption technology in federated learning, combining elliptic curve digital signature and unsupervised model parameter update identification mechanism, the shortcomings of federated learning solutions in resisting malicious attacks and protecting privacy are solved, and higher security and privacy protection are achieved.
Patent Information
- Application Number
- CN202510306723.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-07-01
AI Technical Summary
The existing federated learning solutions still have shortcomings in resisting attacks from malicious users, avoiding malicious aggregation behavior of servers, and protecting privacy.
The edge federated learning privacy protection method based on blockchain and fully homomorphic encryption is adopted, and transparent training process is promoted through blockchain technology, combined with the CKKS fully homomorphic encryption scheme for data encryption, and the elliptic curve digital signature algorithm is used to generate and verify digital signatures to ensure the traceability and integrity of the data. At the same time, an unsupervised model parameter update recognition mechanism was designed to detect malicious clients through the Euclidean distance between model updates.
Effectively resist poisoning attacks, reduce the risk of centralization and malicious aggregation, improve global model performance, and greatly reduce computing and communication overhead through a fully homomorphic encryption solution to ensure the security of data privacy.
Smart Images

Figure CN120238312A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of privacy and security of federated learning, and particularly relates to a privacy protection method for edge federated learning based on blockchain and fully homomorphic encryption Background Technique
[0002] As a new framework for distributed machine learning, federated learning can jointly complete the task of collaborative training of machine learning models by multiple local devices while only sharing local model parameters, and can effectively avoid the data leakage problem caused by the direct transmission of original local data from local devices to the central server. Although federated learning has provided a method for decentralized data training and model aggregation, due to the fact that computing nodes are often located in untrusted environments, it is still vulnerable to security and privacy leakage, specifically facing the following security problems:
[0003] The first problem is that the federated learning scheme may be affected by poisoning attacks. The second problem is that the federated learning scheme is vulnerable to malicious aggregation by the server and single-point failure threats. The third problem is privacy leakage. Although federated learning transmits model parameters instead of original data, existing research has shown that partial privacy information of client data can be deduced by using model parameters. Therefore, both model parameters and data information related to the model need to be strictly protected to avoid deducing the client model and client data
[0004] In summary, while existing federated learning schemes provide efficient aggregation, they still have deficiencies in resisting attacks from malicious users, avoiding malicious aggregation behavior of the server, and protecting privacy Summary of the Invention
[0005] In view of this, the purpose of the present invention is to propose a method for identifying malicious nodes in federated learning based on blockchain and unsupervised recognition mechanism. Specifically, the solution in this paper uses Euclidean distance to measure the similarity between the model updates predicted by each client or the global model updates and the received model updates in each training round. In addition, the CKKS fully homomorphic encryption scheme is used as a protection means in this paper, blockchain technology is used to promote a transparent training process, and the elliptic curve digital signature algorithm is used to generate and verify digital signatures to ensure the traceability and integrity of data
[0006] To achieve the above purpose, the present invention provides the following technical solutions for implementation:
[0007] A privacy protection method for edge federated learning based on blockchain and fully homomorphic encryption, comprising the following steps:
[0008] Client local training and encryption step: The client According to the local dataset And the issued global model Perform pre-training locally to obtain local gradients , and finally perform CKKS fully homomorphic encryption on it. The key pair used is .
[0009] Client blockchain registration steps: Register the identity with the central server in the blockchain and obtain a personal digital account and a pair of personal digital signatures , which can be used to trade the above-mentioned encrypted local gradients and verify the transactions of all parties.
[0010] Client blockchain transaction steps: In the stage where the model needs to be aggregated, "trade" the encrypted local gradients to the central server. All relevant transaction details will be packaged into blocks and stored in the blockchain.
[0011] Model aggregation verification steps: The central server selects one of the good users to be the model aggregator and transaction verifier. Aggregate the locally encrypted gradients "traded" by all parties and verify the legality of each transaction.
[0012] Smart contract steps: After obtaining the aggregated global model, the client still needs to calculate a series of relevant data and upload it to the server to complete malicious identification. When obtaining the global model, the tokens in the account need to be used as a deposit. The smart contract will automatically execute relevant calculations after the client receives the latest global model and prediction model, eliminating client intervention. With the time-consuming proof mechanism under the Intel SGX trusted hardware technology, ensure that the calculation process and results are real and reliable. Finally, encrypt and upload the calculated data through the key pair for encryption and upload.
[0013] Malicious node identification steps: The relevant data in the previous step will be the basis for malicious identification in this step. Calculate the Hessian vector product through L-BFGS. Its calculation result is relatively close for good clients, but quite different for malicious clients. Finally, use Gap statistics to determine the number of clusters. If the clients can be grouped into more than one cluster according to the Gap statistics data, then use K-means to divide the clients into two clusters. Finally, classify the clients with a larger mean value in the cluster as malicious clients. When at least one client is classified as malicious in a certain iteration, the detection ends, and the server deletes the clients classified as malicious and starts the next round of training.
[0014] The client local training and encryption steps specifically include:
[0015] Assume that in the round of training, each client participating in the training aggregation uses its local dataset Perform training. The client obtains the global model from the new blockchain block , and uses to decrypt the global model to obtain . Then randomly select from samples for training. After obtaining the latest local gradient, use to encrypt to obtain . In the subsequent aggregation process, the transaction will be uploaded.
[0016] The specific steps for the client to register on the blockchain include:
[0017] The central server uses the Secp256k1 elliptic curve. The client selects an initial point on this curve , a random integer , and its corresponding private key is . And the corresponding public key can be obtained from the formula . Among them represents a point on the elliptic curve, and the corresponding public key is .
[0018] The specific steps for the client's blockchain transaction include:
[0019] The client uniformly sends the encrypted gradient parameters to the central server in the form of "transactions". The central server verifies the transactions and places the relevant transaction details in the block.
[0020] The specific steps for model aggregation verification include:
[0021] The central server randomly selects one from the good users to be the verifier, aggregator, and block packer. Although the central server acts as an authoritative institution, it does not have the right to aggregate the local model parameters. It is only responsible for updating the ledger and does not participate in model aggregation.
[0022] The model aggregation process can be described by the formula .
[0023] The main way to verify transactions is to use the public key to verify the validity of the signature. The process of generating a digital signature is as follows:
[0024] Step 1: Regenerate a random number , and use the formula to obtain a point on the elliptic curve. At this time, the abscissa of this point is denoted as .
[0025] Step 2: Calculate the hash value of the data to be signed, denoted as .
[0026] Step 3: Calculate according to the formula , where is the base of the modular arithmetic and needs to be specified in advance.
[0027] The length of the signature is usually 40 bytes. The first 20 bytes are , and the last 20 bytes are . The combination of the two is the final digital signature.
[0028] The generated digital signature can be verified according to the formula . If 's abscissa is equal to , then the signature verification is successful. During the signature verification process, the private key is not required at all.
[0029] Therefore, from the above, it can be seen that this algorithm calculates the hash value based on the data to be signed to ensure the integrity of the data, and can be verified through the public key to ensure the compliance of the data source.
[0030] The specific steps of the smart contract include:
[0031] When obtaining the global model, the tokens of the account need to be used as a deposit. After the smart contract receives the latest global model and prediction model on the client side, it will automatically execute relevant calculations, eliminating the intervention of the client. Coupled with the time-consuming proof mechanism under the Intel SGX trusted hardware technology, it is ensured that the calculation process and results are real and reliable.
[0032] The specific steps for malicious node identification include:
[0033] The central server uses the L-BFGS algorithm to approximate the integral Hessian matrix, and calculates the approximate value in each training round, denoted as . Specifically, assume that the global model difference and global model update difference in the th training round are represented by and , where the global model update is aggregated from the model updates of the clients. As the number of training rounds increases, the calculation of and becomes more and more complex. To reduce the calculation and storage costs, a historical record period needs to be defined to regulate the amount of calculation.
[0034] Assume that the training has been carried out for a certain number of rounds, that is, the training round is greater than At this time, there is already a certain amount of historical data in the server, and only the th round of training round to the th round of training round round global model differences and global model difference updates can be represented by and respectively. The specific calculation steps are as follows:
[0035] Calculate .
[0036] Calculate , and obtain the diagonal matrix and the lower triangular submatrix .
[0037] Calculate .
[0038] Calculate to get .
[0039] Calculate .
[0040] Obtain .
[0041] Therefore, it can be predicted that the client is updated iteratively according to the formula in the th round of training. Among them, is the prediction model update of the central server for the client . For a good client, the predicted model update is relatively close to the actual model update , but for a malicious client, there is a large difference. On this basis, the proposed scheme in this paper uses the Euclidean distance to measure the consistency between the predicted model update and the received model update , uses to define client Euclidean distance vectors in the th round of training, and the formula can be obtained. Finally, the Gap statistic is used for the average value to determine the number of clusters. If the clients can be grouped into more than one cluster according to the Gap statistic data, then K-means will be used to group the clients according to their average value It is divided into two clusters. Finally, the clients with larger means in the clusters are classified as malicious clients. When at least one client is classified as malicious in a certain iteration, the detection ends, and the server deletes the clients classified as malicious and starts the next round of training.
[0042] Compared with the prior art, the advantages and beneficial effects of the present invention are as follows: The present invention designs a blockchain-based federated learning framework, which promotes the transparency of the training aggregation process by using blockchain in combination with the elliptic curve digital signature algorithm. The relevant calculation parameters are auditable and traceable, avoiding problems such as single-point failures. In addition, according to the consensus protocol, good clients are randomly selected to perform tasks such as verification, aggregation, and creating new blocks, reducing the degree of centralization and the risk of malicious aggregation; a unsupervised model parameter update recognition mechanism based on the Euclidean distance between model updates is proposed. Malicious clients are detected through the consistency between model updates, and malicious updates are removed, thereby resisting poisoning attacks and improving the performance of the global model; the CKKS fully homomorphic encryption scheme is used to provide privacy protection, which not only greatly reduces the computational and communication overhead, but also effectively prevents the server and clients from snooping and cracking the encrypted model parameters.
[0043] Other advantages, objectives, and features of the present invention will be described to some extent in the subsequent specification, and to some extent, will be obvious to those skilled in the art based on the study of the following text, or can be taught from the practice of the present invention. The objectives and other advantages of the present invention can be realized and obtained through the following specification. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be described in detail preferably with reference to the accompanying drawings, where:
[0045] Figure 1 is a system architecture diagram.
[0046] Figure 2 is a schematic diagram of the recognition mechanism. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0047] The following uses specific specific examples to illustrate the implementation manners of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific implementation manners. The details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the diagrams provided in the following embodiments only illustrate the basic concept of the present invention schematically. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.
[0048] Among them, the attached drawings are only for illustrative purposes, showing only schematic diagrams rather than physical diagrams, and should not be construed as a limitation to the present invention: To better illustrate the embodiments of the present invention, some components in the attached drawings will be omitted, enlarged, or reduced, which do not represent the dimensions of the actual product; for those skilled in the art, it is understandable that some well-known structures and their descriptions in the attached drawings may be omitted. The same or similar reference numerals in the attached drawings of the embodiments of the present invention correspond to the same or similar components.
[0049] The present invention provides a privacy protection method for edge federated learning based on blockchain and fully homomorphic encryption, which utilizes the characteristics of blockchain to endow federated learning with anti-tampering, anti-single point of failure, and data flow transparency, and combines the CKKS fully homomorphic encryption scheme to encrypt relevant calculation parameters, reducing the risk of privacy leakage. For the problem of malicious user poisoning, an unsupervised model parameter update recognition mechanism is designed to detect malicious clients through the consistency between model updates, identify and eliminate malicious updates, thereby resisting poisoning attacks. The architecture diagram of the overall system is as Figure 1 shown;
[0050] As Figure 1 shown in the system architecture diagram, the present invention consists of a central server, a user group, a blockchain system, relevant keys, etc. The secret keys held by the central server and the clients are all different from each other. The keys of the same color in the figure only represent key pairs with the same purpose. The users in the user group are terminals with certain computing power and storage capacity, and each client has a local dataset .
[0051] In addition, in the present invention, assuming that the central server is absolutely trustworthy and curious, then the complete training process of the r-th round can be summarized as follows:
[0052] Training preparation: The central server sends training tasks to the user group, and at the same time registers all the clients in the blockchain system, and generates a unique signature key pair for each client based on ECDSA , and this digital signature is used for "transactions" and participating in blockchain transactions. Then, a genesis block is created according to the training task, and the blocks are synchronized. The main contents of the genesis block include global model initialization parameters , the total number of training rounds , the number of local training times of the client , the historical update record period , the public keys of all clients for digital signatures and other information.
[0053] Step 1: After synchronizing the blockchain content, the client trains on the local dataset . Iterate After that, for the iterated model parameters, use the encryption key pair to perform CKKS homomorphic encryption to address the risk of local gradient privacy leakage. This encryption and decryption key pair is held locally by the client and is not publicly disclosed.
[0054] Step 2: The client "transacts" the encrypted model parameters to the central server in the form of a transaction and performs a digital signature. The central server will verify the digital signature of the "transaction" according to the public key in the genesis block and will give corresponding tokens as a reward if the verification passes. Through the transaction, the central server has the encrypted model parameters of all clients, but cannot decrypt the parameters to obtain the original text, which to a certain extent resists the curious behavior of the server to ensure data security.
[0055] Step 3: The server randomly selects one from the good users and asks this client to act as the verifier, aggregator, and block packer in this training round. The selected client will confirm the above transaction again, then perform the model parameter aggregation task, and finally generate a new block. All clients can obtain the transaction details and the latest global aggregated model.
[0056] Step 4: All clients obtain the new global model from the new block. During the acquisition process, each client has to pay a deposit to the smart contract to ensure participation in the subsequent malicious client identification data upload link. At the same time, the server will, based on the identification data of the previous rounds, make predictions on the models of all clients in this round on the basis of the previous global model parameters and call the result the prediction model, and "transact" it to the client.
[0057] Step 5: After the client obtains the new global model and the prediction model, it needs to calculate the Euclidean distance between the local model and the prediction model. Finally, encrypt the relevant calculated parameters, local operation time, etc. using the encryption key pair This key pair is different from the other encryption key pair mentioned above. This key pair is pre-allocated and is only held by the central server and the client.
[0058] Step 6: The client "transacts" the encrypted relevant calculation results and local operation time to the central server. After receiving the calculation results, the central server verifies the digital signature and operation time, and then performs malicious model parameter update identification on the calculation results to find out abnormal model update data and malicious clients. And start the next round of training from Step 2 until the model converges or reaches the maximum number of training rounds.
[0059] The above is a complete training process. The main symbols used in this article are shown in Table 1.
[0060]
[0061] Table 1 Symbol Explanation
[0062] Next, from the perspective of security, the security of the solution of the embodiment of the present invention is theoretically analyzed.
[0063] From the perspective of formal proof, Theorem 1 can be obtained: A malicious client cannot obtain the sensitive information of the client.
[0064] Proof: In the solution of this article, the blockchain discloses transaction details and the aggregated model to all clients, and the key information therein is protected by the CKKS fully homomorphic encryption scheme. The sum of the aggregated local models can be expressed as . Since the clients are not fully colluded, the model parameters of the good clients cannot be deduced, where . Suppose the first malicious clients conduct a conspiracy attack, and we can get . When , the model parameters of the good clients can be deduced inversely. However, there are still two points to note: First, it is unrealistic for there to be malicious clients among the clients, because such a federated learning task is meaningless. Second, assuming there are malicious clients, theoretically mastering local encryption key pairs , the encrypted model parameters of the good clients can be deduced inversely, but the local encryption key pairs of the good clients
[0065] are still unobtainable, so the ciphertext of the model parameters cannot be deciphered.
[0066] The informal proof is shown in Table 2:
[0067]
[0068] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. A privacy protection method for edge federated learning based on blockchain and fully homomorphic encryption, characterized by: The following steps are involved: Client local training and encryption steps: Client Based on local dataset And the global model issued Perform pre-training locally to obtain local gradients , and finally perform CKKS fully homomorphic encryption on it. The key pair used is ; Client blockchain registration steps: register identity with the central server in the blockchain and obtain personal digital account and personal digital signature pair , which can be used to trade the above-mentioned encrypted local gradients and verify transactions between parties; Client blockchain transaction steps: At the stage where the model aggregation is required, the encrypted local gradients are "traded" to the central server, and all relevant transaction details are packaged into blocks and stored in the blockchain; Model aggregation verification step: The central server selects one good user to become the model aggregator and transaction verifier; aggregates the local encrypted gradients from the "transactions" of all parties, and verifies the legitimacy of each transaction; Smart contract steps: After obtaining the aggregated global model, the client still needs to calculate a series of related data and upload it to the server to complete malicious identification; when obtaining the global model, the account token needs to be used as a deposit. After the client receives the latest global model and prediction model, the smart contract will automatically perform related calculations to eliminate client intervention, and then use the time consumption proof mechanism under Intel SGX trusted hardware technology to ensure that the calculation process and results are real and reliable; finally, the calculated data is sent to the server through the key pair. Perform encrypted upload; Malicious node identification step: The relevant data in the previous step will become the basis for malicious identification in this step; the Hessian vector product is calculated by L-BFGS; the calculation results are relatively close for good clients, but are quite different for malicious clients; finally, the Gap statistics are used to determine the number of clusters; if the clients can be grouped into more than one cluster according to the Gap statistics, K-means will be used to divide the clients into two clusters; finally, the clients with larger means in the cluster are classified as malicious clients; when at least one client is classified as malicious in a certain iteration, the detection ends, the server deletes the clients classified as malicious and starts the next round of training.
2. The edge federated learning privacy protection method based on blockchain and homomorphic encryption according to claim 1 is characterized by: The client local training and encryption steps specifically include: Assume that in In each training round, each client participating in the training aggregation uses its local dataset Training is performed; the client obtains the global model from the new block of the blockchain ,use For the global model Decrypt and get , and then randomly from Select from Samples Train and get the latest local gradient and use right Encrypted , the subsequent aggregation stage will upload the transaction .
3. The edge federated learning privacy protection method based on blockchain and homomorphic encryption according to claim 1 is characterized in that: The client blockchain registration steps specifically include: The central server uses the Secp256k1 elliptic curve, and the client selects an initial point on the curve , a random integer , and its corresponding private key is , and by the formula The corresponding public key can be obtained; Represented as a point on the elliptic curve, the corresponding public key is .
4. The edge federated learning privacy protection method based on blockchain and homomorphic encryption according to claim 1 is characterized by: The client blockchain transaction steps specifically include: The client sends the encrypted gradient parameters in the form of "transactions" to the central server, which verifies the transactions and places the relevant transaction details in the block.
5. The edge federated learning privacy protection method based on blockchain and homomorphic encryption according to claim 1 is characterized in that: The model aggregation verification step specifically includes: The central server randomly selects one good user to become a validator, aggregator, and block packager. Although the central server plays the role of an authority, it does not have the right to aggregate local model parameters. It is only responsible for updating the ledger and does not participate in model aggregation. The model aggregation process can be expressed by the formula to describe; The main way to verify a transaction is to use a public key to verify the validity of the signature; the process of generating a digital signature is as follows: Step 1: Regenerate a random number , using the formula A point on the elliptic curve can be found , the horizontal coordinate of this point is marked as ; Step 2 Calculate the hash value of the data to be signed, denoted as ; Step 3 According to the formula Calculate, where It is the base number of the model budget and needs to be specified in advance; The signature is usually 40 bytes long, the first 20 bytes are , the following 20 bytes are , the combination of the two is the final digital signature; The generated digital signature can be based on the formula Verify; if The horizontal axis is equal to , then the signature verification is successful. In the signature verification process, the private key does not need to be involved at all; Therefore, from the above, it can be concluded that the algorithm will calculate the hash value according to the data that needs to be signed to ensure the integrity of the data, and it can be verified through the public key to ensure the compliance of the data source.
6. The edge federated learning privacy protection method based on blockchain and homomorphic encryption according to claim 1 is characterized by: The smart contract steps specifically include: When obtaining the global model, the account tokens need to be used as a deposit. After the client receives the latest global model and prediction model, the smart contract will automatically perform relevant calculations to eliminate client intervention. It is supplemented by the time-consuming proof mechanism under Intel SGX trusted hardware technology to ensure that the calculation process and results are real and reliable.
7. The edge federated learning privacy protection method based on blockchain and homomorphic encryption according to claim 1 is characterized by: The malicious node identification step specifically includes: The central server uses the L-BFGS algorithm to approximate the integral Hessian matrix and calculates it in each training round. The approximate value of Specifically, assuming that The global model difference and global model update difference in rounds of training are used and Represents, where the global model update is aggregated from the client model updates. As the number of training rounds increases, and The calculation will become more complicated; to reduce the calculation and storage costs, a historical record period needs to be defined To reduce the amount of calculation; Assume that the training has been carried out for a certain number of rounds, that is, the training round Greater than In this case, the server already has a certain amount of historical data, and only the first Training rounds to Between training rounds The global model difference and global model difference update can be done using and To express; the specific calculation steps are as follows: (1) Calculation ; (2) Calculation , and obtain the diagonal matrix and the lower triangular submatrix ; (3) Calculation ; (4) Calculation get ; (5) Calculation ; (6) ; Therefore, it is predictable that the client In round training, the formula To update iteratively; Central server to client For good clients, the predicted model update is Update with actual model It is relatively close, but it is quite different for malicious clients; on this basis, this paper uses Euclidean distance to measure the update of the prediction model. and received model updates Consistency between To define Client No. The Euclidean distance vector in round training can be obtained by the formula Finally, the average The Gap statistic is used to determine the number of clusters; if the clients can be grouped into more than one cluster based on the Gap statistic, K-means is used to group the clients based on their mean values. Divide into two clusters; finally, classify the clients with larger mean in the cluster as malicious clients; when at least one client is classified as malicious in a certain iteration, the detection ends, the server deletes the clients classified as malicious and starts the next round of training.