Medical block chain data sharing method based on dual quality thresholds and federated learning
By introducing blockchain technology and a dual-quality threshold aggregation mechanism in federated learning, combined with random response and differential privacy technology, the problems of privacy leakage and malicious node poisoning in federated learning are solved, and more efficient and secure medical data sharing is achieved.
Patent Information
- Application Number
- CN202411901370.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-23
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2044-12-23
AI Technical Summary
There are problems in federated learning with the risk of privacy leakage and the untrustworthy central servers, especially when facing poisoning attacks from malicious nodes, it is difficult to ensure the stability of model prediction performance.
A medical blockchain data sharing method based on dual-quality thresholds and federated learning is adopted, and a central server is replaced by blockchain, and privacy protection is protected by combining random response and differential privacy technologies. A dual-quality threshold aggregation mechanism and reputation incentive evaluation mechanism are designed to filter malicious nodes and low-quality nodes.
It effectively reduces the impact of noise on model performance, improves federated learning training efficiency, enhances resistance to malicious nodes, and promotes honest participation of nodes through reputation incentive mechanisms.
Smart Images

Figure CN119995888A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer technology and relates to a medical blockchain data sharing method based on dual quality thresholds and federated learning. Background Art
[0002] The rapid development of the Internet of Things (IoT) has transformed many application fields, and the integration of IoT technology and various medical devices has given birth to the Internet of Medical Things (IoMT). In the medical field, a large amount of medical data is generated every day. With the help of these medical data, high-quality deep learning models can be trained to provide diagnostic assistance to doctors, thereby improving the efficiency of medical services. However, relying solely on the data of a single institution to build a model often encounters challenges such as limited distribution of training samples and lack of data. In addition, since most medical data involves patient personal information and medical condition privacy, they need to be highly confidential. Federated learning (FL) can solve the above problems. As a distributed machine learning framework, federated learning supports data sharing of medical institutions without leaving the local area. Multiple medical institutions use local medical data for model training, and then upload the model parameters to the central aggregator for aggregation to obtain a global model. FL solves the problem of data silos and does not require their own data to be moved out of the local area, reducing the risk of data privacy leakage. However, the current FL has the following two problems in practical applications.
[0003] The first problem is that the parameter exchange method in FL still has the risk of privacy leakage. Attackers can obtain model-specific private information through inference attacks and model inversion attacks.
[0004] The second problem is that FL relies on a central server for model aggregation, which cannot guarantee the central server's credibility and is prone to single point failure. In addition, the current FL cannot effectively resist poisoning attacks by malicious nodes. Label flipping attacks, as a type of poisoning attack, can modify the labels of local training data and then update and upload the poisoned local model for aggregation, ultimately reducing the prediction accuracy of the global model or even failing to converge. This is fatal to medical scenarios that have high requirements for model prediction performance, and incorrect medical diagnoses will lead to serious consequences.
[0005] There are currently multiple solutions to solve the problem. Homomorphic encryption (HE) can protect intermediate results, but it will bring heavy computational overhead. At the same time, since encrypted data usually takes up more space than plaintext data, the use of homomorphic encryption will increase the cost of network transmission. Differential privacy (DP) can add noise to the original gradient or parameters to achieve privacy protection. However, the addition of noise will affect the accuracy of the model training results, which is unacceptable in medical scenarios with high accuracy requirements. Privacy and model performance need to be weighed.
[0006] Regarding the second problem, since blockchain has the characteristics of decentralization and immutability, it can be combined with FL to solve the problem of single point failure and replace the central server of traditional FL for model aggregation, but it still cannot resist poisoning attacks by malicious nodes. Existing methods focus on reducing the weight of malicious nodes in aggregation, but cannot fundamentally solve the problem. In addition, there is a lack of effective reputation evaluation and incentive mechanisms to punish malicious nodes and low-quality nodes. Therefore, it is necessary to comprehensively consider the above issues and design a medical data sharing method that takes into account both security and practicality. Summary of the invention
[0007] In view of this, the purpose of the present invention is to provide a medical blockchain data sharing method based on dual quality thresholds and federated learning, which can protect privacy data while reducing the impact of noise on model performance and improving the efficiency of federated learning training, and resist poisoning attacks by malicious nodes, and finally perform reputation incentive management on these nodes.
[0008] In order to achieve the above object, the present invention provides the following technical solutions:
[0009] A medical blockchain data sharing method based on dual quality thresholds and federated learning includes the following steps:
[0010] S1: The task publishing node publishes the federated learning training task and uploads the initial model to the blockchain. All nodes register their information on the blockchain and randomly select nodes to form a committee.
[0011] S2: Participating nodes download the latest global model parameters from the blockchain, use local data sets to train the model, and use random response and differential privacy techniques to protect the privacy of model gradients;
[0012] S3: The participating nodes that have completed the training upload the local model parameters to the blockchain in the form of transactions; the committee selects the node with the highest reputation as the mining node and calculates the training quality of each participating node through the local model parameters to evaluate whether the node can participate in the aggregation of global model parameters, and then calculates the participation of the participating node in the aggregation in the past. Finally, the contribution value is calculated through the node training quality, local data set and participation, and the global model parameters are aggregated according to the weight of the contribution value;
[0013] S4: The mining node calculates the reputation value and reward and punishment incentives of each node based on its performance in this round of iterative training and its previous performance, and packages all the reputation incentive information and global model parameters of all nodes into a new block;
[0014] S5: After receiving the new block, other committee nodes will verify it. After more than half of the committee nodes pass the verification, the new block will be saved on the blockchain and the node composition of the committee will be adjusted.
[0015] Furthermore, the specific process of step S1 includes: the system model consists of four parts: task publishing node, blockchain, task participating node, and committee; the task publishing node is responsible for formulating federated learning tasks and reward mechanisms to attract node participation, and at the same time, the initial model needs to be uploaded to the blockchain through transactions; the blockchain replaces the original central server in federated learning, because it can provide data that cannot be tampered with and is open and transparent, ensuring that all key data such as task release, participant information, and model parameter reputation are securely stored, increasing the transparency and credibility of the system; the task participating node is the entity that provides data, responsible for obtaining tasks from the chain and performing model training locally; the committee members are selected from the task participants who perform well in each round, and are responsible for packaging blocks, aggregating models, and reaching consensus; after being selected into the committee, the node does not participate in the model training task. All nodes need to register information and create accounts on the blockchain when they first join the system, and the initial committee members are randomly selected from the nodes.
[0016] Furthermore, the specific process of step S2 includes: the publishing node will initially set the task initial global model parameters Upload to the blockchain; first, the participating nodes download the latest global model parameters of the e-th round iteration from the blockchain Perform local model training.
[0017] Extract a small batch of data d for each iteration from the local dataset D and calculate each data d in d i Gradient In order to prevent the influence of individual gradients being too large on model updates, gradient clipping is required. The gradient after clipping is:
[0018]
[0019] Where C is the gradient clipping threshold. The above formula means that when the 2-norm of the gradient is less than the clipping threshold C, the gradient is retained as the original value. When the 2-norm of the gradient is greater than the clipping threshold C, the gradient is limited to C. Because if ||g i ||2, the sensitivity cannot be calculated, and DP noise cannot be added. According to the definition of sensitivity, the sensitivity is calculated as:
[0020]
[0021] Where D is the size of the local data set. The noise scale can be calculated by calculating the sensitivity. This paper uses Gaussian noise, and the noise scale is:
[0022]
[0023] where ε k =ε / E, where ε is the privacy budget, E is the number of local training rounds, and ε k It refers to the privacy budget consumed by exposing the current round of information of the node, and δ is the probability of information leakage. The random response is introduced into the differential privacy mechanism to interfere with the gradient. The gradient after interference is:
[0024]
[0025] Finally, update the parameters of the local training model for the kth round:
[0026]
[0027] In the first iteration To replace the initial global model parameters γ is the learning rate.
[0028] Further, in step S3, the committee selects the node with the highest reputation as the mining node to download all transactions from the blockchain for model aggregation, which specifically includes the following steps:
[0029] S31: The mining node calculates the local model training quality of the participating nodes and divides the set according to different quality thresholds.
[0030] S32: Calculate the participation of the participating nodes in the past in the aggregation, and then calculate their contribution values through the node training quality, local data set and participation.
[0031] S33: Calculate the node aggregation weights of different sets according to the contribution values and aggregate the model parameters.
[0032] Further, in step S31, after the participating nodes upload the local training model parameters to the blockchain, the committee members summarize them according to the proportion of local data sets of each participating node as weights, and calculate their accuracy ACC in the test set. all , then summarize the model parameters that do not include participating node i and calculate its accuracy ACC noi , then the model training quality of participating node i in the tth round of training can be expressed as:
[0033]
[0034] The model performance is expressed as accuracy. If the training quality of participating node i This means that the node's participation in aggregation improves the performance of the model.
[0035] In order to resist poisoning attacks by malicious nodes, we set two quality thresholds Q Hth and Q Lth , and divide the nodes into three sets; below Q Lth The nodes are grouped into the set Set low , higher than Q Lth Lower than Q Hth The nodes are classified into the set, which is higher than Q Hth The nodes are grouped into a collection.
[0036] Furthermore, in step S32, in addition to considering the size of the local device data set of the participating nodes, we also consider the training quality of the participating nodes after each round of training and the participation of a node. The node contribution value is expressed as:
[0037]
[0038] in represents the contribution value of participating node i in round t, represents the proportion of the local device data set size of participating node i, represents the quality of the participating node i in the tth round of training, The node participation we introduced. Participation We use the exponentially weighted moving average method (EWMA) to calculate, which is specifically expressed as:
[0039]
[0040] in represents the participation degree of node i in the previous round, θ represents the attenuation factor, which is generally taken as θ≥0.9. and In the tth round of aggregation, whether the participating node i enters the set Set hight or a Set accept Indicator function participating in aggregation, if the node enters the set Set hight Participate in aggregation, then If the node enters the set Set accept Participate in aggregation, then If the node does not participate in the aggregation, then α and β are expressed as indicator functions and The weight parameter, and α+β=2, α>β, indicates that the node enters the set Set hight Participate in aggregation and enter the set Set accept The degree of participation in aggregation is different. The more times a participating node enters the high-quality collection, the higher its participation degree.
[0041] Further, in step S33, the set Set hight and Set accept The aggregation weights of the participating nodes are expressed as:
[0042]
[0043] Aggregate the set Set according to the aggregation weight hight and Set accept Model parameters of midside nodes:
[0044]
[0045] in Represents the local model parameters uploaded by participating node i in round t, and finally aggregates the global model parameters:
[0046]
[0047] Where λ and η are sets Set hight Aggregated model parameters and Set accept Aggregated model parameters The weight of the global model parameter aggregation, in order to improve the performance of the model so that the high-quality set Set hight Aggregated model parameters carry more weight.
[0048] Further, in step S4, the mining node calculates the reputation value and reward and punishment incentives of each node through its performance in this round of iterative training and in combination with its previous performance, and packages all the reputation incentive information and global model parameters of all nodes into a new block, which specifically includes the following steps:
[0049] S41: Calculate the reputation and incentives of participating nodes;
[0050] S42: Calculate the reputation and incentives of committee members. The mining node packages the global model parameters and node reputation incentives and other information into a new block for verification by other committee members.
[0051] Furthermore, in step S41, in order to reasonably reflect the node performance reputation, we comprehensively consider the performance reputation and historical performance reputation of the participating nodes in the tth round of model aggregation. The change in the income of the participating nodes in each round changes accordingly according to the change value of their reputation in each round. The final reputation of node i and the change in revenue per round Y i t It can be expressed as:
[0052]
[0053] in represents the historical cumulative reputation of node i, and Represents the change in reputation of node i after each round of iteration.
[0054] is the performance reputation of node i in round t, and is also the main influencing value of node reputation change. The node reputation is calculated based on whether the node participates in aggregation in each round as the judgment standard, and then:
[0055]
[0056] u i is the probability that node i successfully sends the local model parameters, and τ is the reputation compensation factor. The failure to receive the local model parameters may also be due to unstable network connection, which is not necessarily the responsibility of the participating nodes, so compensation is needed. ψ and ζ are the sensitivity of the quality contribution to the reputation of node i in round t, and in order to increase the penalty for malicious nodes and poorly performing nodes, the sensitivity is set to ψ < ζ.
[0057] is the historical performance reputation of node i, which will be updated according to the conditions such as whether the node enters different sets in each iteration and whether it participates in aggregation. It is expressed as:
[0058]
[0059] where s i Represents the historical performance score of node i:
[0060]
[0061] Indicates that node i passes Q Lth The number of times, Indicates that node i has not passed Q Lth The number of times, Indicates that node i passes Q Hth The number of times, Indicates that node i passes the quality threshold Q Lth But it does not pass the quality threshold Q Hth The number of times, k and g are and The weight parameter is used to increase the weight of nodes that do not pass Q Lth When the penalty is set, set k+g=1,k<g, and yes and The weight parameter is used to reflect the HthQ Lth For better performance, set So that the node passes Q Hth The more times you do it, the faster your reputation increases.
[0062] Reputation Plays a leading role in node reputation changes, taking into account historical performance reputation It can achieve an auxiliary balance. Because the historical performance reputation will be updated according to which set the node enters in each round, if node i performs poorly in this round, the updated will decrease if the historical performance reputation after the update is still positive, considering the excellent performance of node i before, plus the updated After that, it decreased Vice versa, if node j performs poorly in the previous rounds, but suddenly performs well in round t, its reputation value will increase. is a positive number, but still needs to be considered If after update If it is negative, then the reputation gained in this round will be reduced.
[0063] Furthermore, in step S42, the reputation of committee members needs to be obtained stably, and the benefits change accordingly according to the changes in their reputation. The reputation and benefits are as follows:
[0064]
[0065] N C is the number of committee members, N P is the number of participating nodes, x and y are the factors that change the reputation of the committee nodes. When a block is added to the chain, N C If the middle node passes the verification, its reputation will be increased, otherwise it will be reduced; when the block is not added to the chain, N C If the middle node passes the verification, its reputation will be reduced, otherwise it will be increased. In order to increase the penalty for wrong verification, x<y is set.
[0066] In our method, the committee plays an important role in aggregating the global model and verifying the legitimacy of new blocks. The selection of committee members should reduce the participation of malicious nodes. Therefore, the selection of committee members follows the following rules:
[0067] When a node joins the system, it needs to register information on the blockchain. Each node will receive an initial reputation value, which is randomly selected from N nodes in the initial stage. P Select N C , then from N C The node with the highest reputation is selected as the mining node to perform operations such as packaging new blocks in the aggregation model and leave its own signature. CThe node verifies whether the signature in the new block is from the mining node, whether it contains the hash value of the previous block, and randomly selects a transaction from the new block for verification to confirm whether the calculation is correct. When more than half of the nodes pass the verification, the new block is added to the blockchain. After each round, the mining node and the node with verification error exit the committee and reselect from the participating nodes. First, determine whether the node participated in the aggregation of the global model in this round, that is, determine the reputation value. Is it greater than 0? Then, among the nodes that meet the conditions, the reputation value The size of the weighted average is randomly selected.
[0068] The beneficial effects of the present invention are:
[0069] 1) The present invention proposes a privacy protection mechanism that integrates random response and differential privacy technology, which provides privacy protection for the model while reducing the impact of noise on model performance and reducing the computational time of local model training.
[0070] 2) The present invention designs an aggregation mechanism with dual quality thresholds to filter out malicious nodes and low-quality nodes, and assigns weights to nodes that pass different thresholds according to their contribution values when aggregating the global model, which can effectively improve model performance and resist poisoning attacks.
[0071] 3) The present invention designs a reputation incentive evaluation mechanism based on the historical performance of nodes, which can effectively incentivize nodes to participate honestly and punish nodes' malicious behavior.
[0072] Other advantages, objectives and features of the present invention will be described in the following description to some extent, and to some extent, will be obvious to those skilled in the art based on the following examination and study, or can be taught from the practice of the present invention. The objectives and other advantages of the present invention can be realized and obtained through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0073] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention will be described in detail below in conjunction with the accompanying drawings, wherein:
[0074] Figure 1 The working process of the medical blockchain data sharing method based on dual quality thresholds and federated learning of the present invention;
[0075] Figure 2 The system structure of the medical blockchain data sharing method based on dual quality thresholds and federated learning of the present invention; DETAILED DESCRIPTION
[0076] The following describes the embodiments of the present invention by specific examples, and those skilled in the art can easily understand other advantages and effects of the present invention from the contents disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the illustrations provided in the following embodiments only illustrate the basic concept of the present invention in a schematic manner, and the following embodiments and features in the embodiments can be combined with each other without conflict.
[0077] See also Figure 1 and Figure 2 ,The present invention provides a medical blockchain data sharing method based on dual quality thresholds and federated learning, such as Figure 1 As shown, the method mainly includes the following steps:
[0078] Step 1: The task publishing node publishes the federated learning training task and uploads the initial model to the blockchain. All nodes register their information on the blockchain and randomly select nodes to form a committee.
[0079] Step 2: Participating nodes download the latest global model parameters from the blockchain, use local data sets to train the model, and use random response and differential privacy techniques to protect the privacy of model gradients;
[0080] Step 3: The participating nodes that have completed the training upload the local model parameters to the blockchain in the form of transactions; the committee uses the quality threshold to determine which nodes can participate in the aggregation of global model parameters, and then calculates the participation of the participating nodes in the past. Finally, the contribution value is calculated based on the node training quality, local data set and participation, and the global model parameters are aggregated according to the weight of the contribution value;
[0081] Step 4: The mining node calculates the reputation value and reward and punishment incentives of each node based on its performance in this round of iterative training and its previous performance, and packages all the nodes’ reputation incentives and other information and global model parameters into a new block;
[0082] Step 5: After receiving the new block, other committee nodes will verify it. After more than half of the committee nodes pass the verification, the new block will be saved on the blockchain and the node composition of the committee will be adjusted.
[0083] Figure 2 The system structure diagram of the present invention is mainly composed of four parts: task publisher, blockchain, task participants, and committee. The following is an explanation with reference to the accompanying drawings, including the following steps:
[0084] 1) The task publisher publishes the federated learning training task and uploads the initial model and reward mechanism to the blockchain through transactions.
[0085] 2) All nodes first need to register their own information on the blockchain, including address, ID, initial reputation, initial account, etc., and then download the initial model from the blockchain.
[0086] 3) Participants use local datasets to train models. To prevent inference attacks, differential privacy is used to add Gaussian noise to the training data gradients.
[0087] 4) Task participants upload the trained local model parameters to the blockchain transaction pool in the form of transactions.
[0088] 5) The committee selects the node with the highest reputation as the mining node and downloads the local model parameters of all participants from the blockchain for aggregation.
[0089] 6) The mining node first calculates the training quality of each participant node through the local model parameters to evaluate whether the node can participate in the aggregation of global model parameters, then calculates the participation of each participant in the aggregation in the past, and finally calculates its contribution value through the node training quality, local data set and participation, and aggregates the global model parameters according to the weight of the contribution value.
[0090] 7) The mining node calculates the reputation value of each node through its performance in this round of iterative training and combined with its previous performance, and calculates rewards and punishments based on the changes in the reputation value in this round.
[0091] 8) The mining node packages all nodes’ reputation incentives and global model parameters into a new block and broadcasts it to other committee nodes for verification.
[0092] 9) After more than half of the verifications are passed, the new block will be added to the blockchain.
[0093] 10) The task publisher downloads the final global model of this round of training from the blockchain and uses it as the initial model for the next iteration of training
[0094] Optional, Figure 2 Step 3) is the local training process of random response differential privacy protection, which is as follows:
[0095]
[0096] Optional, Figure 2 Step 6) is to aggregate global model parameters based on the local model parameters sent by the participating nodes, specifically including:
[0097]
[0098]
[0099] Optional, Figure 2 Steps 7), 8), and 9) are reputation evaluation and incentive calculation, packaging new blocks, and consensus verification, specifically:
[0100] Reputation evaluation and incentive calculation: Comprehensively consider the performance reputation and historical performance reputation of the participating nodes in the tth round of model aggregation. The change in the income of the participating nodes in each round changes accordingly according to the change in their reputation value in each round. The final reputation of node i and the change in revenue per round Y i t It can be expressed as:
[0101]
[0102] in represents the historical cumulative reputation of node i, and Represents the change in reputation of node i after each round of iteration.
[0103] is the performance reputation of node i in round t, and is also the main influencing value of node reputation change. The node reputation is calculated based on whether the node participates in aggregation in each round as the judgment standard, and then:
[0104]
[0105] u i is the probability that node i successfully sends the local model parameters, and τ is the reputation compensation factor. The failure to receive the local model parameters may also be due to unstable network connection, which is not necessarily the responsibility of the participating nodes, so compensation is needed. ψ and ζ are the sensitivity of the quality contribution to the reputation of node i in round t, and in order to increase the penalty for malicious nodes and poorly performing nodes, the sensitivity is set to ψ < ζ.
[0106] is the historical performance reputation of node i, which will be updated according to the conditions such as whether the node enters different sets in each iteration and whether it participates in aggregation. It is expressed as:
[0107]
[0108] where s i Represents the historical performance score of node i:
[0109]
[0110] Indicates that node i passes Q Lth The number of times Indicates that node i has not passed Q Lth The number of times Indicates that node i passes Q Hth The number of times Indicates that node i passes the quality threshold Q Lth But it does not pass the quality threshold Q Hth The number of times k and g are and The weight parameter is used to increase the weight of nodes that do not pass Q Lth When the penalty is set, set k+g=1,k<g, and yes and The weight parameter is used to reflect the Hth Q Lth For better performance, set So that the node passes Q Hth The more times you do it, the faster your reputation increases.
[0111] The reputation of committee members needs to be steadily acquired, and the income changes accordingly. The reputation and income are as follows:
[0112]
[0113] N C is the number of committee members, N P is the number of participating nodes, x and y are the factors that change the reputation of the committee nodes. When a block is added to the chain, N C If the middle node passes the verification, its reputation will be increased, otherwise it will be reduced; when the block is not added to the chain, N C If the middle node passes the verification, its reputation will be reduced, otherwise it will be increased. In order to increase the penalty for wrong verification, x<y is set.
[0114] Packaging new blocks and consensus verification: In our method, the committee plays an important role in aggregating the global model and verifying the legitimacy of new blocks. The selection of committee members should reduce the participation of malicious nodes. Therefore, the selection of committee members follows the following rules:
[0115] When a node joins the system, it needs to register information on the blockchain. Each node will receive an initial reputation value, which is randomly selected from N nodes in the initial stage. P Select N C , then from N C The node with the highest reputation is selected as the mining node to perform operations such as packaging new blocks in the aggregation model and leave its own signature. CThe node verifies whether the signature in the new block is from the mining node, whether it contains the hash value of the previous block, and randomly selects a transaction from the new block for verification to confirm whether the calculation is correct. When more than half of the nodes pass the verification, the new block is added to the blockchain. After each round, the mining node and the node with verification error exit the committee and reselect from the participating nodes. First, determine whether the node participated in the aggregation of the global model in this round, that is, determine the reputation value. Is it greater than 0? Then, among the nodes that meet the conditions, the reputation value The size of the weighted average is randomly selected.
[0116] Finally, it should be noted that the above embodiments are only used to illustrate the technical solution of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solution of the present invention can be modified or replaced by equivalents without departing from the purpose and scope of the technical solution, which should be included in the scope of the claims of the present invention.
Claims
1. A medical blockchain data sharing method based on dual quality thresholds and federated learning, characterized by: The method comprises the following steps: S1: The task publishing node publishes the federated learning training task and uploads the initial model to the blockchain. All nodes register their information on the blockchain and randomly select nodes to form a committee. S2: Participating nodes download the latest global model parameters from the blockchain, use local data sets to train the model, and use random response and differential privacy techniques to protect the privacy of model gradients; S3: The participating nodes that have completed the training upload the local model parameters to the blockchain in the form of transactions; the committee selects the node with the highest reputation as the mining node and calculates the training quality of each participating node through the local model parameters to evaluate whether the node can participate in the aggregation of global model parameters, and then calculates the participation of the participating node in the aggregation in the past. Finally, the contribution value is calculated through the node training quality, local data set and participation, and the global model parameters are aggregated according to the weight of the contribution value; S4: The mining node calculates the reputation value and reward and punishment incentives of each node based on its performance in this round of iterative training and its previous performance, and packages all the reputation incentive information and global model parameters of all nodes into a new block; S5: After receiving the new block, other committee nodes will verify it. After more than half of the committee nodes pass the verification, the new block will be saved on the blockchain and the node composition of the committee will be adjusted.
2. According to the medical blockchain data sharing method based on dual quality thresholds and federated learning according to claim 1, it is characterized in that: The specific process of step S1 includes: the system model consists of four parts: task publishing node, blockchain, task participating node, and committee; the task publishing node is responsible for formulating federated learning tasks and reward mechanisms to attract node participation, and at the same time, the initial model needs to be uploaded to the blockchain through transactions; the blockchain replaces the original central server in federated learning, because it can provide data that cannot be tampered with and is open and transparent, ensuring that all key data such as task release, participant information, and model parameter reputation are securely stored, increasing the transparency and credibility of the system; the task participating node is an entity that provides data, responsible for obtaining tasks from the chain and performing model training locally; committee members are selected from task participants with good performance in each round, and are responsible for packaging blocks, aggregating models, and reaching consensus; after being selected into the committee, the node does not participate in the model training task. All nodes need to register information and create accounts on the blockchain when they first join the system, and the initial committee members are randomly selected from the nodes.
3. According to the medical blockchain data sharing method based on dual quality thresholds and federated learning according to claim 1, it is characterized in that: The specific process of step S2 includes: the publishing node will initially set the initial global model parameters of the task Upload to the blockchain; first, the participating nodes download the latest global model parameters of the e-th round iteration from the blockchain Perform local model training. Extract a small batch of data d for each iteration from the local dataset D and calculate each data d in d i Gradient In order to prevent the influence of individual gradients being too large on model updates, gradient clipping is required. The gradient after clipping is: Where C is the gradient clipping threshold. The above formula means that when the 2-norm of the gradient is less than the clipping threshold C, the gradient is retained as the original value. When the 2-norm of the gradient is greater than the clipping threshold C, the gradient is limited to C. Because if ||g i ||2, the sensitivity cannot be calculated, and DP noise cannot be added. According to the definition of sensitivity, the sensitivity is calculated as: Where D is the size of the local data set. The noise scale can be calculated by calculating the sensitivity. This paper uses Gaussian noise, and the noise scale is: where ε k =ε / E, where ε is the privacy budget, E is the number of local training rounds, and ε k It refers to the privacy budget consumed by exposing the current round of information of the node, and δ is the probability of information leakage. The random response is introduced into the differential privacy mechanism to interfere with the gradient. The gradient after interference is: Finally, update the parameters of the local training model for the kth round: In the first iteration To replace the initial global model parameters γ is the learning rate.
4. The medical blockchain data sharing method based on dual quality thresholds and federated learning according to claim 1 is characterized in that: In step S3, the committee selects the node with the highest reputation as the mining node to download all transactions from the blockchain for model aggregation, which specifically includes the following steps: S31: The mining node calculates the local model training quality of the participating nodes and divides the set according to different quality thresholds. S32: Calculate the participation of the participating nodes in the past in the aggregation, and then calculate their contribution values through the node training quality, local data set and participation. S33: Calculate the node aggregation weights of different sets according to the contribution values and aggregate the model parameters.
5. According to the medical blockchain data sharing method based on dual quality thresholds and federated learning according to claim 4, it is characterized in that: In step S31, after the participating nodes upload the local training model parameters to the blockchain, the committee members summarize them according to the proportion of local data sets of each participating node as weights, and calculate their accuracy ACC in the test set. all , then summarize the model parameters that do not include participating node i and calculate its accuracy ACC noi , then the model training quality of participating node i in the tth round of training can be expressed as: The model performance is expressed as accuracy. If the training quality of participating node i This means that the node's participation in aggregation improves the performance of the model. In order to resist poisoning attacks by malicious nodes, we set two quality thresholds Q Hth and Q Lth , and divide the nodes into three sets; below Q Lth The nodes are grouped into the set Set low , higher than Q Lth Lower than Q Hth The nodes are classified into the set, which is higher than Q Hth The nodes are grouped into a collection.
6. According to the medical blockchain data sharing method based on dual quality thresholds and federated learning according to claim 4, it is characterized in that: In step S32, in addition to considering the size of the local device data set of the participating node, we also consider the training quality of the participating node after each round of training and the participation of a node. The node contribution value is expressed as: in represents the contribution value of participating node i in round t, represents the proportion of the local device data set size of participating node i, represents the quality of the participating node i in the tth round of training, The node participation we introduced. We use the exponentially weighted moving average method (EWMA) to calculate, which is specifically expressed as: in represents the participation degree of node i in the previous round, θ represents the attenuation factor, which is generally taken as θ≥0.
9. and In the tth round of aggregation, whether the participating node i enters the set Set hight Or a Set accept Indicator function participating in aggregation, if the node enters the set Set hight Participate in aggregation, then If the node enters the set Set accept Participate in aggregation, then If the node does not participate in the aggregation, then α and β are expressed as indicator functions and The weight parameter, and α+β=2, α>β, indicates that the node enters the set Set hight Participate in aggregation and enter the set Set accept The degree of participation in aggregation is different. The more times a participating node enters the high-quality collection, the higher its participation degree.
7. The medical blockchain data sharing method based on dual quality thresholds and federated learning according to claim 4 is characterized in that: In step S33, the set Set hight and Set accept The aggregation weights of the participating nodes are expressed as: Aggregate the set Set according to the aggregation weight hight and Set accept Model parameters of midside nodes: in Represents the local model parameters uploaded by participating node i in round t, and finally aggregates the global model parameters: Where λ and η are sets Set hight Aggregated model parameters and Set accept Aggregated model parameters The weight of the global model parameter aggregation, in order to improve the performance of the model so that the high-quality set Set hight Aggregated model parameters carry more weight.
8. The medical blockchain data sharing method based on dual quality thresholds and federated learning according to claim 1 is characterized in that: In step S4, the mining node calculates the reputation value and reward and punishment incentives of each node based on its performance in this round of iterative training and in combination with its previous performance, and packages all the reputation incentive information and global model parameters of all nodes into a new block, which specifically includes the following steps: S41: Calculate the reputation and incentives of participating nodes; S42: Calculate the reputation and incentives of committee members. The mining node packages the global model parameters and node reputation incentives and other information into a new block for verification by other committee members.
9. The medical blockchain data sharing method based on dual quality thresholds and federated learning according to claim 8 is characterized in that: In step S41, in order to reasonably reflect the node performance reputation, we comprehensively consider the performance reputation and historical performance reputation of the participating nodes in the tth round of model aggregation. The change in the income of the participating nodes in each round changes accordingly according to the change value of their reputation in each round. The final reputation of node i and the change in revenue per round Y i t It can be expressed as: in represents the historical cumulative reputation of node i, and Represents the change in reputation of node i after each round of iteration. is the performance reputation of node i in round t, and is also the main influencing value of node reputation change. The node reputation is calculated based on whether the node participates in aggregation in each round as the judgment standard, and then: u i is the probability that node i successfully sends the local model parameters, and τ is the reputation compensation factor. The failure to receive the local model parameters may also be due to unstable network connection, which is not necessarily the responsibility of the participating nodes, so compensation is needed. ψ and ζ are the sensitivity of the quality contribution to the reputation of node i in round t, and in order to increase the penalty for malicious nodes and poorly performing nodes, the sensitivity is set to ψ < ζ. is the historical performance reputation of node i, which will be updated according to the conditions such as whether the node enters different sets in each iteration and whether it participates in aggregation. It is expressed as: where s i Represents the historical performance score of node i: Indicates that node i passes Q Lth The number of times, Indicates that node i has not passed Q Lth The number of times, Indicates that node i passes Q Hth The number of times, Indicates that node i passes the quality threshold Q Lth But it does not pass the quality threshold Q Hth The number of times, k and g are and The weight parameter is used to increase the weight of nodes that do not pass Q Lth When the penalty is set, set k+g=1,k<g, and yes and The weight parameter is used to reflect the Hth Q Lth For better performance, set So that the node passes Q Hth The more times you do it, the faster your reputation increases. Reputation Plays a leading role in node reputation changes, taking into account historical performance reputation It can achieve an auxiliary balance. Because the historical performance reputation will be updated according to which set the node enters in each round, if node i performs poorly in this round, the updated will decrease if the historical performance reputation after the update is still positive, considering the excellent performance of node i before, plus the updated After that, it decreased Vice versa, if node j performs poorly in the previous rounds, but suddenly performs well in round t, its reputation value will increase. is a positive number, but still needs to be considered If after update If it is negative, then the reputation gained in this round will be reduced.
10. The medical blockchain data sharing method based on dual quality thresholds and federated learning according to claim 8 is characterized in that: In step S42, the reputation of committee members needs to be obtained stably, and the income changes accordingly with the change of their reputation. The reputation and income are as follows: N C is the number of committee members, N P is the number of participating nodes, x and y are the factors that change the reputation of the committee nodes. When a block is added to the chain, N C If the middle node passes the verification, its reputation will be increased, otherwise it will be reduced; when the block is not added to the chain, N C If the middle node passes the verification, its reputation will be reduced, otherwise it will be increased. In order to increase the penalty for wrong verification, x<y is set. In our method, the committee plays an important role in aggregating the global model and verifying the legitimacy of new blocks. The selection of committee members should reduce the participation of malicious nodes. Therefore, the selection of committee members follows the following rules: When a node joins the system, it needs to register information on the blockchain. Each node will receive an initial reputation value, which is randomly selected from N nodes in the initial stage. P Select N C , then from N C The node with the highest reputation is selected as the mining node to perform operations such as packaging new blocks in the aggregation model and leave its own signature. C The node verifies whether the signature in the new block is from the mining node, whether it contains the hash value of the previous block, and randomly selects a transaction from the new block for verification to confirm whether the calculation is correct. When more than half of the nodes pass the verification, the new block is added to the blockchain. After each round, the mining node and the node with verification error exit the committee and reselect from the participating nodes. First, determine whether the node participated in the aggregation of the global model in this round, that is, determine the reputation value. Is it greater than 0? Then, among the nodes that meet the conditions, the reputation value The size of the weighted average is randomly selected.
Citation Information
Patent Citations
Fair privacy calculation method based on federated node contribution
CN116306910A
Trusted federal learning method based on alliance chain
CN116961939A
Hierarchical federal learning incentive method and system based on block chain
CN118052297A
Block chain data sharing method under assistance of asynchronous federated learning
CN118069607A
Block chain-based trusted federal learning system and method in edge computing scene
CN118350451A
Cited By
Decentralized federal continuous learning method based on task credibility evaluation
CN121835821A