An adaptive privacy-preserving personalized federated learning method based on blockchain
By introducing blockchain technology and adaptive privacy protection mechanisms into federated learning, central server dependence and privacy security issues are solved, and decentralized, personalized and efficient federated learning is achieved.
Patent Information
- Application Number
- CN202410381356.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-01
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2044-04-01
AI Technical Summary
The existing federated learning technology relies on central servers and is vulnerable to poisoning attacks, model reversal and other attacks, and has problems such as privacy data leakage and performance bottlenecks.
Adaptive privacy protection personalized federated learning method based on blockchain is adopted, decentralized through the blockchain architecture, design new consensus algorithms and edge node election mechanisms, group and aggregate user local models with similar data distributions, and generate personalized models.
Decentralized federated learning is realized, reducing communication overhead and privacy data leakage risks, meeting users' personalized needs, and improving model accuracy.
Smart Images

Figure CN118246572B_ABST
Abstract
Description
Technical Field
[0001] The invention discloses an adaptive privacy protection personalized federated learning method based on blockchain, belonging to the technical field of blockchain privacy protection. Background Art
[0002] In recent years, deep learning has been widely used. In order to obtain more accurate prediction models, deep learning usually requires a large amount of data for training. If this data contains sensitive information, such as personal identity, health status, etc., then there may be serious privacy and security issues during the model training process, such as data leakage, abuse or unauthorized use. Federated learning, as an emerging distributed machine learning technology, allows users to obtain a shared global model through collaborative training without sharing data.
[0003] Although federated learning has been proven to be an effective distributed learning mechanism, it requires a trusted central server to aggregate local models uploaded by users, which is a typical centralized structure. Constrained by this structure, federated learning is vulnerable to attacks such as poisoning and model inversion, resulting in problems such as reduced model accuracy and privacy data leakage. At the same time, the performance bottleneck of the server itself and the risk of malicious tampering with the aggregation results will also have an adverse impact on federated learning. As a distributed ledger, blockchain has the characteristics of decentralization, immutability, and traceability. Therefore, in recent years, a large number of scholars have proposed using blockchain technology to achieve a fully decentralized federated learning architecture solution. Summary of the invention
[0004] The present invention proposes an adaptive privacy-preserving personalized federated learning method based on blockchain, which realizes decentralized and personalized federated learning, solves the problem of federated learning's dependence on central servers through blockchain architecture, and uses participating users with similar data distribution to cooperate with each other to provide users with exclusive personalized models. A new consensus algorithm is designed based on the blockchain architecture, and the local models of users with similar data distribution are grouped and aggregated by the way that edge nodes elect leaders, so as to realize mutual cooperation among users. In addition, when the edge node elects a leader, the method verifies its priority through all adjacent nodes of the edge node, which reduces the communication overhead to a certain extent.
[0005] The technical solution of the present invention is as follows:
[0006] An adaptive privacy-preserving personalized federated learning method based on blockchain. The learning method is based on a blockchain jointly maintained by edge nodes such as servers or base stations in cellular networks, and a group of distributed clients U={C1,C2,...,C n}, n represents the total number of participating users;
[0007] It is characterized by comprising the following steps:
[0008] (1) Participating users first download the trained initial global model M from the latest generated block. g And the relevant parameter information involved in this training;
[0009] (2) Participating user C i The global model M g Extract feature representation of data from local datasets R i , and use this model to train the local model
[0010] (3) Next, user C i The feature representation R i With local model Encrypt and upload the encrypted ciphertext to the edge node connected to it;
[0011] (4) Finally, after the edge nodes on the blockchain communicate with each other, they elect a leader through a consensus algorithm and send the encrypted R i and Sent to the leader; the leader collects all R i Calculate and determine the similarity of user data, and group users with similar data distribution into a group. Each group can represent a specific data distribution type. After grouping, the leader aggregates the local models uploaded by all participating users in each group and generates a personalized model M for each group. p To better adapt to the data distribution of the group, the personalized model M p The number of M is consistent with the number of groups; p This is the personalized model suitable for this data type.
[0012] Preferably, the specific steps of the above step (1) are as follows:
[0013] Before the system runs, the initial global model M is obtained through a specified number of iterations. g From and save in the latest generated block, at the same time, the block also includes the participating user C i The relevant information includes:
[0014] (1-1) Parameters of local training for participating users, including learning rate η, number of local training rounds LocalEpoch and local training batch LocalBatchsize;
[0015] (1-2) Specify the Relu layer of the initial global model involved in user feature extraction representation;
[0016] (1-3) The pseudo-random generator and the public keys Pk of all participating users are used to encrypt the feature representations and local models generated by the participating users in step ②.
[0017] Preferably, the specific steps of the above step (2) are as follows:
[0018] Participating User C i Input your own local data into the global model M g , and from the global model M g Select a channel in the specified Relu layer and extract the sparsity of local data in the channel.
[0019] Sparsity is calculated by the following equation:
[0020]
[0021] Where sp represents sparsity, calculate D i The average value of all sample sparsity in is regarded as the participating user C i The sparsity sp(D i ), and define it as the feature representation R of the user i :
[0022]
[0023] At the same time, participating user C i The global model M will also be initialized using local data based on relevant training information. g Train and get the local model Local Model The training definition of is as follows:
[0024]
[0025] Among them, D i represents the local data of user i, M q Represents a personalized model.
[0026] Preferably, the specific steps of the above step (3) are as follows:
[0027] When participating user C i Calculate the feature representation R i With local model After that, a shared key will be generated by the Diffie-Hellman algorithm; the algorithm is as follows:
[0028] For the i-th participating user Ci , the public keys of all participating users have been downloaded, and the user generates a shared key S for it and all other users in pairs through KA.agree in the Diffie-Hellman algorithm i,j :
[0029] S i,j =KAagree(Sk i ,Pk j )
[0030] Where 1≤j≤n,Sk i Represents the private key of user i, Pk j represents the public key of user j;
[0031] Afterwards, let the shared key S i,j As the input of the pseudo-random generator, the mask is generated. By adding the mask to generate the ciphertext, the user will represent the feature R i With local model encryption:
[0032]
[0033]
[0034] Where Z is the specified modulus;
[0035] After encryption is completed, the ciphertext is uploaded to the edge node connected to it.
[0036] Preferably, the leader in the above step (4) aggregates the local models uploaded by all participating users in each group and references the pure proof-of-stake mechanism in the Algorand consensus protocol, and uses a new consensus algorithm to elect and verify the leader in the edge nodes. The leader is responsible for completing the grouping aggregation of the model.
[0037] Preferably, the new consensus algorithm is as follows:
[0038] The new consensus algorithm is divided into two parts. In the first part, edge nodes elect a leader, and in the second part, the leader aggregates the groups to obtain a personalized model.
[0039] The first part of the process is as follows:
[0040] After obtaining the local model uploaded by the client connected to it, the edge node on the blockchain will elect a leader from the edge nodes holding the legal local model, select a specified number of stakes from the total amount of stakes owned by all edge nodes, and use the selected stakes in each edge node as an indicator for electing a leader. Define τ as the number of stakes selected from the total amount of stakes. Each stake u may be selected, and W is the sum of the number of all stakes. The probability of any stake being selected is τ / W. For an edge node with w stakes, it first uses its private key S k Generate hash hash and proof Proof with seed Seed; and for w+1 subintervals divided on the interval [0,1), when a value J is calculated that satisfies:
[0041]
[0042] This means that the edge node has J rights selected, that is, the priority value of the edge node is J, and hashlen is the length of the hash. After calculating the priorities of all edge nodes, the edge nodes will communicate with each other, and each edge node will verify its Proof priority through its neighbor nodes. Each edge node compares the priority in turn, and the edge node with the highest priority becomes the leader.
[0043] The second part of the process is as follows:
[0044] When a leader is elected, all edge nodes will send their feature representations and local models to the leader. The leader will compare the similarities between data by calculating the Euclidean distance of the feature vectors. For the feature vector R of the i-th participating user, i and the feature vector R of the jth client j , their similarity is calculated by the square difference formula:
[0045]
[0046] Calculate the data similarity between any two participating users;
[0047] The leader will randomly select a feature vector R uploaded by a participating user v As a comparison object, the similarity between the feature vector and all other feature vectors is calculated to find similar users and group them;
[0048] First, a comparison threshold is set. The leader determines whether to group two users into one group based on whether the similarity of the two user feature vectors is less than the threshold. Through the above method, the leader divides each participating user into several groups and defines the grouping set. q represents the total number of groups,
[0049] After grouping, the leader aggregates the local models uploaded by the participating users in each group, and uses the federated averaging method to aggregate the models in the group to generate a personalized model M. p :
[0050] in, Represents the total number of participating users in the group, |D k ∣ represents the number of samples owned by the participating user k in the group, represents the encrypted local model of participating user k, Represents the sum of the sample numbers of all users in the group. This method is used to aggregate and generate personalized models for each group to obtain a personalized model set
[0051] The beneficial effects of the present invention are:
[0052] This paper proposes an adaptive privacy-preserving personalized federated learning framework based on blockchain, which combines blockchain technology and model personalization technology for the first time and applies it to federated learning. While ensuring the privacy and security of user data, it also meets the specific needs of participating users.
[0053] A new consensus algorithm is proposed based on the Algorand consensus protocol, which adaptively groups and aggregates the local models of participating users with similar data distribution, reducing the communication cost to a certain extent. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] Figure 1 It is a schematic diagram of the method of the present invention. DETAILED DESCRIPTION
[0055] The present invention will be further described below in conjunction with the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solution of the present invention, and cannot be used to limit the protection scope of the present invention.
[0056] like Figure 1 As shown in FIG, an adaptive privacy-preserving personalized federated learning method based on blockchain is presented. The learning method is based on a blockchain jointly maintained by edge nodes such as servers or base stations in cellular networks, and a group of distributed clients U = {C1, C2, ..., C n}, n represents the total number of participating users;
[0057] It is characterized by comprising the following steps:
[0058] (1) Participating users first download the trained initial global model M from the latest generated block. gAnd the relevant parameter information involved in this training;
[0059] (2) Participating user C i The global model M g Extract feature representation of data from local datasets R i , and use this model to train the local model
[0060] (3) Next, user C i The feature representation R i With local model Encrypt and upload the encrypted ciphertext to the edge node connected to it;
[0061] (4) Finally, after the edge nodes on the blockchain communicate with each other, they elect a leader through a consensus algorithm and send the encrypted R i and Sent to the leader; the leader collects all R i Calculate and determine the similarity of user data, and group users with similar data distribution into a group. Each group can represent a specific data distribution type. After grouping, the leader aggregates the local models uploaded by all participating users in each group and generates a personalized model M for each group. p To better adapt to the data distribution of the group, the personalized model M p The number of M is consistent with the number of groups; p This is the personalized model suitable for this data type.
[0062] Preferably, the specific steps of the above step (1) are as follows:
[0063] Before the system runs, the initial global model M is obtained through a specified number of iterations. g From and save in the latest generated block, at the same time, the block also includes the participating user C i The relevant information includes:
[0064] (1-1) Parameters of local training for participating users, including learning rate η, number of local training rounds LocalEpoch and local training batch LocalBatchsize;
[0065] (1-2) Specify the Relu layer of the initial global model involved in user feature extraction representation;
[0066] (1-3) The pseudo-random generator and the public keys Pk of all participating users are used to encrypt the feature representations and local models generated by the participating users in step ②.
[0067] Preferably, the specific steps of the above step (2) are as follows:
[0068] Participating User C i Input your own local data into the global model M g , and from the global model M g Select a channel in the specified Relu layer and extract the sparsity of local data in the channel.
[0069] Sparsity is calculated by the following equation:
[0070]
[0071] Where sp represents sparsity, calculate D i The average value of all sample sparsity in is regarded as the participating user C i The sparsity sp(D i ), and define it as the feature representation R of the user i :
[0072]
[0073] At the same time, participating user C i The global model M will also be initialized using local data based on relevant training information. g Train and get the local model Local Model The training definition of is as follows:
[0074]
[0075] Among them, D i represents the local data of user i, M q Represents a personalized model.
[0076] Preferably, the specific steps of the above step (3) are as follows:
[0077] When participating user C i Calculate the feature representation R i With local model After that, a shared key will be generated by the Diffie-Hellman algorithm; the algorithm is as follows:
[0078] For the i-th participating user C i , the public keys of all participating users have been downloaded, and the user generates a shared key S for it and all other users in pairs through KA.agree in the Diffie-Hellman algorithm i,j :
[0079] S i,j=KAagree(Sk i ,Pk j )
[0080] Where 1≤j≤n,Sk i Represents the private key of user i, Pk j represents the public key of user j;
[0081] Afterwards, let the shared key S i,j As the input of the pseudo-random generator, the mask is generated. By adding the mask to generate the ciphertext, the user will represent the feature R i With local model encryption:
[0082]
[0083]
[0084] Where Z is the specified modulus;
[0085] After encryption is completed, the ciphertext is uploaded to the edge node connected to it.
[0086] Preferably, the leader in the above step (4) aggregates the local models uploaded by all participating users in each group and references the pure proof-of-stake mechanism in the Algorand consensus protocol, and uses a new consensus algorithm to elect and verify the leader in the edge nodes. The leader is responsible for completing the grouping aggregation of the model.
[0087] Preferably, the new consensus algorithm is as follows:
[0088] The new consensus algorithm is divided into two parts. In the first part, edge nodes elect a leader, and in the second part, the leader aggregates the groups to obtain a personalized model.
[0089] The first part of the process is as follows:
[0090] After obtaining the local model uploaded by the client connected to it, the edge node on the blockchain will elect a leader from the edge nodes holding the legal local model, select a specified number of stakes from the total amount of stakes owned by all edge nodes, and use the selected stakes in each edge node as an indicator for electing a leader. Define τ as the number of stakes selected from the total amount of stakes. Each stake u may be selected, and W is the sum of the number of all stakes. The probability of any stake being selected is τ / W. For an edge node with w stakes, it first uses its private key S k Generate hash hash and proof Proof with seed Seed; and for w+1 subintervals divided on the interval [0,1), when a value J is calculated that satisfies:
[0091]
[0092] This means that the edge node has J rights selected, that is, the priority value of the edge node is J, and hashlen is the length of the hash. After calculating the priorities of all edge nodes, the edge nodes will communicate with each other, and each edge node will verify its Proof priority through its neighbor nodes. Each edge node compares the priority in turn, and the edge node with the highest priority becomes the leader.
[0093] The second part of the process is as follows:
[0094] When a leader is elected, all edge nodes will send their feature representations and local models to the leader. The leader will compare the similarities between data by calculating the Euclidean distance of the feature vectors. For the feature vector R of the i-th participating user, i and the feature vector R of the jth client j , their similarity is calculated by the square difference formula:
[0095]
[0096] Calculate the data similarity between any two participating users;
[0097] The leader will randomly select a feature vector R uploaded by a participating user v As a comparison object, the similarity between the feature vector and all other feature vectors is calculated to find similar users and group them;
[0098] First, a comparison threshold is set. The leader determines whether to group two users into one group based on whether the similarity of the two user feature vectors is less than the threshold. Through the above method, the leader divides each participating user into several groups and defines the grouping set. q represents the total number of groups,
[0099] After grouping, the leader aggregates the local models uploaded by the participating users in each group, and uses the federated averaging method to aggregate the models in the group to generate a personalized model M. p :
[0100] in, Represents the total number of participating users in the group, |D k ∣ represents the number of samples owned by the participating user k in the group, represents the encrypted local model of participating user k, Represents the sum of the sample numbers of all users in the group. This method is used to aggregate and generate personalized models for each group to obtain a personalized model set
[0101] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the technical principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.
Claims
1. An adaptive privacy-preserving personalized federated learning method based on blockchain. The learning method is based on a blockchain jointly maintained by edge nodes such as servers or base stations in cellular networks, and a group of distributed clients U={C1,C2,...,C n }, n represents the total number of participating users; Features The steps include: (1) Participating users first download the trained initial global model M from the latest generated block. g And the relevant parameter information involved in this training; (2) Participating user C i The global model M g Extract feature representation of data from local datasets R i , and use this model to train the local model (3) Next, user C i The feature representation R i With local model Encrypt and upload the encrypted ciphertext to the edge node connected to it; (4) Finally, after the edge nodes on the blockchain communicate with each other, they elect a leader through a consensus algorithm and send the encrypted R i and Sent to the leader; the leader collects all R i Calculate and determine the similarity of user data, and group users with similar data distribution into one group. Each group can represent a specific data distribution type. After grouping, the leader aggregates the local models uploaded by all participating users in each group and generates a personalized model M for each group. p To better adapt to the data distribution of the group, the personalized model M p The number of M is consistent with the number of groups; p This is the personalized model suitable for this group of data types.
2. According to claim 1, the adaptive privacy-preserving personalized federated learning method based on blockchain is characterized in that The specific steps of step (1) are as follows: Before the system runs, the initial global model M is obtained through a specified number of iterations. g From and save in the latest generated block, at the same time, the block also includes the participating user C i The relevant information includes: (1-1) Parameters of local training for participating users, including learning rate η, number of local training rounds LocalEpoch and local training batch LocalBatchsize; (1-2) Specify the Relu layer of the initial global model involved in user feature extraction representation; (1-3) The pseudo-random generator and the public keys Pk of all participating users are used to encrypt the feature representations and local models generated by the participating users in step (1-2).
3. According to claim 1, the adaptive privacy-preserving personalized federated learning method based on blockchain is characterized in that The specific steps of step (2) are as follows: Participating User C i Input your own local data into the global model M g , and from the global model M g Select a channel in the specified Relu layer and extract the sparsity of local data in the channel. Sparsity is calculated by the following equation: Where sp represents sparsity, calculate D i The average value of all sample sparsity in is regarded as the participating user C i The sparsity sp(D i ), and define it as the feature representation R of the user i : At the same time, participating user C i The global model M will also be initialized using local data based on relevant training information. g Train and get the local model Local Model The training definition of is as follows: Among them, D i represents the local data of user i, M q Represents a personalized model.
4. According to claim 1, the adaptive privacy-preserving personalized federated learning method based on blockchain is characterized in that The specific steps of step (3) are as follows: When participating user C i Calculate the feature representation R i With local model After that, a shared key will be generated by the Diffie-Hellman algorithm; the algorithm is as follows: For the i-th participating user C i , the public keys of all participating users have been downloaded, and the user generates a shared key S for it and all other users in pairs through KA.agree in the Diffie-Hellman algorithm i,j : S i,j =KAgree(Sk i ,Pk j ) Where 1≤j≤n,Sk i Represents the private key of user i, Pk j represents the public key of user j; Afterwards, let the shared key S i,j As the input of the pseudo-random generator, the mask is generated. By adding the mask to generate the ciphertext, the user will represent the feature R i With local model encryption: Where Z is the specified modulus; After encryption is completed, the ciphertext is uploaded to the edge node connected to it.
5. According to claim 1, the adaptive privacy-preserving personalized federated learning method based on blockchain is characterized in that In step (4), the leader aggregates the local models uploaded by all participating users in each group and references the pure proof-of-stake mechanism in the Algorand consensus protocol, and uses a new consensus algorithm to elect and verify the leader in the edge nodes. The leader is responsible for completing the group aggregation of the model.
6. The method of adaptive privacy-preserving personalized federated learning based on blockchain according to claim 5, characterized in that The new consensus algorithm is as follows: The new consensus algorithm is divided into two parts. In the first part, edge nodes elect a leader, and in the second part, the leader aggregates the groups to obtain a personalized model. The first part of the process is as follows: After obtaining the local model uploaded by the client connected to it, the edge node on the blockchain will elect a leader from the edge nodes holding the legal local model, select a specified number of stakes from the total amount of stakes owned by all edge nodes, and use the selected stakes in each edge node as an indicator for electing a leader. Define τ as the number of stakes selected from the total amount of stakes. Each stake u may be selected, and W is the sum of the number of all stakes. The probability of any stake being selected is τ / W. For an edge node with w stakes, it first uses its private key S k Generate hash hash and proof Proof with seed Seed; and for w+1 subintervals divided on the interval [0,1), when a value J is calculated that satisfies: This means that the edge node has J rights selected, that is, the priority value of the edge node is J, and hashlen is the length of the hash. After calculating the priorities of all edge nodes, the edge nodes will communicate with each other, and each edge node will verify its Proof priority through its neighbor nodes. Each edge node compares the priority in turn, and the edge node with the highest priority becomes the leader. The second part of the process is as follows: When a leader is elected, all edge nodes will send their feature representations and local models to the leader. The leader will compare the similarities between data by calculating the Euclidean distance of the feature vectors. For the feature vector R of the i-th participating user, i and the feature vector R of the jth client j , their similarity is calculated by the square difference formula: Calculate the data similarity between any two participating users; The leader will randomly select a feature vector R uploaded by a participating user v As a comparison object, the similarity between the feature vector and all other feature vectors is calculated to find similar users and group them; First, a comparison threshold is set. The leader determines whether to group two users into one group based on whether the similarity of the two user feature vectors is less than the threshold. Through the above method, the leader divides each participating user into several groups and defines the grouping set. q represents the total number of groups, After grouping, the leader aggregates the local models uploaded by the participating users in each group, and uses the federated averaging method to aggregate the models in the group to generate a personalized model M. p : in, represents the total number of participating users in the group, |D k ∣ represents the number of samples owned by the participating user k in the group, represents the encrypted local model of participating user k, Represents the sum of the sample numbers of all users in the group. This method is used to aggregate and generate personalized models for each group to obtain a personalized model set
Citation Information
Patent Citations
Federal learning method of lightweight grouping consensus based on block chain
CN115796261A
Personalized federal learning method and device based on block chain
CN116432777A