Model distributed training method based on blockchain node trust management
By determining the trust value based on the local model training behavior data of the computing node in distributed training and aggregating the global model based on the trust management mechanism, the problem of insufficient reliability of distributed training global model in the existing technology is solved, and the security of the computing node and the training quality of the global model are improved.
Patent Information
- Application Number
- CN202411967319.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-26
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2044-12-26
AI Technical Summary
The existing distributed training methods for model are insufficient in aggregating global models, mainly due to the lack of dynamicity and transparency of the trust management mechanism, and it is impossible to effectively evaluate the trust value of nodes.
During the distributed training process, the trust value is determined based on the local model training behavior data of each computing node, and the global model is aggregated based on the trust management mechanism. The specific steps include: the calculation node trains the initial model based on the training data set, determines the local training behavior data, calculates the training trust value, and filters the nodes participating in the model training based on the trust value, and finally generates a global model by the main node aggregation.
The security and reliability of distributed training computing nodes are improved, thereby improving the training quality of the global model and enhancing the system's defense ability against malicious node attacks.
Smart Images

Figure CN120012168A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of blockchain technology, and in particular to a model distributed training method based on blockchain node trust management. Background Art
[0002] Distributed systems are widely used in data processing, collaborative computing, and intelligent training tasks, where multiple nodes jointly participate in task execution and achieve system goals by sharing computing results or data updates. With the rapid development of artificial intelligence and big data technologies, distributed learning has gradually become an important method for solving large-scale data modeling and training. In distributed learning, multiple nodes jointly train machine learning models through collaborative computing without sharing local raw data, thereby improving computing efficiency while protecting data privacy. The highly heterogeneous and open nature of the distributed learning environment also makes the system face the following major security and trust challenges, such as: 1) Insufficient security and uniqueness of node authentication; 2) Lack of dynamism and transparency in the trust assessment mechanism; 3) Limited data storage and tampering protection capabilities; 4) Destructive risks of malicious node behavior.
[0003] As a decentralized, tamper-proof and traceable data management technology, blockchain technology provides a new solution for node trust management in distributed learning.
[0004] The patent document with application publication number CN116776373A is cited here. CN116776373A ensures the credibility of users in the shared blockchain through a two-factor trusted model based on identity trusted authentication and behavior trusted evaluation, and uses a trusted federated average aggregation algorithm to increase the contribution of high-trust and high-quality model parameters in the aggregation model, thereby improving the accuracy of the medical data federated learning aggregation model. It can be seen that CN116776373A calculates the behavioral trust value of each data provider based on the retrieval mechanism and trust mechanism on the chain, and distributes and collaborates with members of the highly trusted federated learning training committee to train the same federated learning model.
[0005] However, existing model distributed training methods only consider the on-chain retrieval mechanism and trust mechanism to determine the members of the highly credible federated learning training committee, and the reliability of the global model generated by the aggregation is insufficient.
[0006] It should be noted that the information in the above background technology section is only used to enhance the understanding of the background technology of the present application, and therefore may include technical information that is not known or easily inferred by ordinary technicians in the field. Summary of the invention
[0007] In view of the above problems, the present application is proposed to provide a model distributed training method, device and system based on blockchain node trust management that overcomes the above problems or at least partially solves the above problems, including: A model distributed training method based on blockchain node trust management, the method involves a user terminal and a computing node; the computing node is used to store blockchain and train a distributed model; the computing node includes a master node and a slave node; the user terminal pre-stores an initial model and a training data set; the user terminal is used to send the initial model and the training data set to the computing node; The method comprises the following steps: The computing node trains an initial model based on the training data set and determines local training behavior data; The computing node determines the training trust value based on the local training behavior data; The computing node trains the initial model based on the training data set and the training trust value; wherein, when the initial model training is completed, the master node obtains the trained model from the slave node to aggregate and generate a global model, and sends the global model to the user end.
[0008] Furthermore, the local training behavior data includes the number of local iterations and the local iteration step; the step of determining the training trust value according to the local training behavior data by the computing node includes: The computing node determines the training trust value according to the local iteration number and the local iteration step size.
[0009] Furthermore, the step of determining the training trust value by the computing node according to the local iteration number and the local iteration step size includes: The computing node generates a local trust value according to the local iteration number and the local iteration step size; The computing node broadcasts the local trust value and the trained model to other nodes respectively; The computing node calculates a recommended trust value corresponding to the computing node performing broadcasting based on the training data set of the node, the local trust value received through broadcasting, and the trained model received through broadcasting, and broadcasts the recommended trust value to other nodes; The computing node generates the training trust value according to the local trust value and the recommended trust value of the node by other nodes.
[0010] Furthermore, the computing node pre-stores a deviation threshold, a trust value penalty term and a trust value threshold; the step of the computing node training the initial model according to the training data set and the training trust value includes: The computing node updates the training trust value based on the local trust value, the recommended trust value of other nodes for the node, the deviation threshold and the trust value penalty item; The computing nodes select computing nodes participating in model training according to the training trust value and the trust value threshold; The computing nodes participating in the model training train the initial model according to the training data set.
[0011] Furthermore, the step of updating the training trust value of the computing node according to the local trust value, the recommended trust value of the node by other nodes, the deviation threshold and the trust value penalty term includes: The computing node selects the computing node to be punished based on the local trust value, the recommended trust value of other nodes to the node and the deviation threshold; The computing node to be punished updates the training trust value according to the training trust value and the trust value penalty term.
[0012] Furthermore, the step of the computing node selecting the computing node to be punished according to the local trust value, the recommended trust value of other nodes to the node and the deviation threshold includes: The computing node calculates the trust deviation value based on the local trust value and the recommended trust value of other nodes for the node; The computing nodes select computing nodes to be punished according to the trust deviation value and the deviation threshold.
[0013] Furthermore, the step of calculating the trust deviation value according to the local trust value and the recommended trust value of the node by other nodes comprises: The computing node calculates the trust average value based on the local trust value and the trust value recommended by other nodes for the node; The computing node calculates the trust deviation value according to the local trust value and the trust average value.
[0014] Furthermore, the training trust value includes calculating the local trust value of node i ; The step of determining the training trust value according to the local iteration number and the local iteration step size by the computing node comprises: The computing node determines the number of local iterations according to the and the local iteration step size The local trust value of the computing node i is determined by the following formula: :
[0015] in, is the number of local iterations of the model in computing node i, is the local iteration step size of the model execution, and are parameters related to the machine model.
[0016] Furthermore, the step of generating the training trust value according to the local trust value and the recommended trust value of the node by other nodes comprises: The computing node is based on the local trust value , recommended trust value As well as the historical trust value, the training trust value of the computing node i in the tth round is calculated by the following formula: :
[0017] in, and is a weight factor, and the historical trust value includes the training trust value of the computing node when the round is less than t.
[0018] Furthermore, the step of updating the training trust value of the computing node to be punished according to the training trust value and the trust value penalty term includes: The computing node to be punished updates the training trust value according to the training trust value and the trust value penalty term through the following formula:
[0019] in, represents the training confidence value, Represents the trust value penalty item.
[0020] Furthermore, the step of calculating the trust deviation value according to the local trust value and the recommended trust value of the node by other nodes comprises: The computing node expresses the trust deviation value through the following polynomial based on the local trust value and the recommended trust value of other nodes to the node:
[0021] in, represents the local trust value, Indicates the recommendation trust value.
[0022] A model distributed training device based on blockchain node trust management, the device involves a user terminal and a computing node; the computing node is used to store blockchain and train distributed models; the computing node includes a master node and a slave node; the user terminal pre-stores an initial model and a training data set; the user terminal is used to send the initial model and the training data set to the computing node; The device comprises: A first training module is used to train an initial model based on a training data set and determine local training behavior data; A trust value generation module, used to determine a training trust value based on local training behavior data; The second training module is used to train the initial model based on the training data set and the training trust value; wherein, when the initial model training is completed, the master node obtains the trained model from the slave node to aggregate and generate a global model, and sends the global model to the user end.
[0023] A model distributed training system based on blockchain node trust management, the system comprising a user terminal and a computing node; the computing node is used to store blockchain and train distributed models; the computing node comprises a master node and a slave node; the user terminal pre-stores an initial model and a training data set; The system comprises: The user terminal is used to send the initial model and the training data set to the computing node; The computing node is used to train the initial model based on the training data set, and to determine the local training behavior data, and to determine the training trust value based on the local training behavior data, and to train the initial model based on the training data set and the training trust value; wherein, when the initial model training is completed, the master node is used to obtain the trained model from the slave node to aggregate and generate a global model, and to send the global model to the user end.
[0024] A computer device comprises a processor, a memory and a computer program stored in the memory and capable of running on the processor, wherein when the computer program is executed by the processor, the method described in any embodiment of the present application is implemented.
[0025] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method described in any embodiment of the present application is implemented.
[0026] This application has the following advantages: In the embodiments of the present application, in view of the insufficient reliability of the global model generated by the aggregation of the existing model distributed training method, the present application provides a solution of "using the training behavior data generated by the local model of each computing node in the distributed training process to determine the trust value of each computing node, and then aggregating the global model based on trust management", specifically: the computing node trains the initial model based on the training data set and determines the local training behavior data; the computing node determines the training trust value based on the local training behavior data; the computing node trains the initial model based on the training data set and the training trust value; wherein, when the initial model training is completed, the master node obtains the trained model from the slave node to aggregate and generate a global model, and sends the global model to the user end. By determining the training trust value based on the local training behavior data, the computing nodes are trained based on the trust management model, and then the models of the high trust value computing nodes are aggregated to generate a global model, thereby improving the security and reliability of the distributed training computing nodes, and then improving the training quality of the global model. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] In order to more clearly illustrate the technical solution of the present application, the drawings required for use in the description of the present application will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0028] Figure 1 This is a flowchart of the steps of a model distributed training method based on blockchain node trust management provided by an embodiment of the present application; Figure 2 It is a flowchart of the node trust management and model aggregation process of the blockchain provided by a specific embodiment of the present application; Figure 3 It is a schematic diagram of the framework of the node trust management and model aggregation process of the blockchain provided by a specific embodiment of the present application; Figure 4 It is a structural block diagram of a model distributed training device based on blockchain node trust management provided by an embodiment of the present application; Figure 5 It is a structural block diagram of a computer device provided in one embodiment of the present application. DETAILED DESCRIPTION
[0029] In order to make the objects, features and advantages of the present application more obvious and understandable, the present application is further described in detail below in conjunction with the accompanying drawings and specific implementation methods. Obviously, the described embodiments are part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in the field without creative work are within the scope of protection of the present application.
[0030] The inventors have found through analysis of the prior art that distributed learning systems face the following major security and trust challenges: 1) The security and uniqueness of node authentication are insufficient. The reason is that distributed learning systems are usually participated by multiple heterogeneous nodes (such as personal devices, edge computing nodes, and cloud servers). These nodes are distributed in different physical locations and controlled by different participants. Traditional identity authentication mechanisms are difficult to effectively ensure the uniqueness and authenticity of each node, thus providing opportunities for malicious nodes to disguise and infiltrate.
[0031] 2) The trust evaluation mechanism lacks dynamism and transparency. The reason is that in distributed learning, node behaviors such as the quality of calculation results, data integrity, and participation frequency will change over time. However, traditional trust evaluation usually adopts a centralized or predefined method, which cannot respond to changes in node behavior in real time and lacks a transparent trust management mechanism, which easily leads to the failure of the system trust mechanism.
[0032] 3) Data storage and tamper protection capabilities are limited. The reason is that distributed learning systems need to record the behavior of participating nodes, task completion status and trust assessment results, while traditional storage methods (such as centralized databases) are susceptible to tampering, forgery and single point failures, thereby reducing the security and audit capabilities of the system.
[0033] 4) The destructive risk of malicious node behavior is that nodes in distributed learning may be attacked maliciously or have malicious behaviors themselves, such as uploading false gradients, manipulating model updates, or refusing to perform computing tasks. Such behaviors will affect the performance of the global model and undermine the credibility of the system.
[0034] Based on the above analysis of existing technologies, blockchain technology, as a decentralized, tamper-proof and traceable data management technology, provides a new solution for node trust management in distributed learning. Generate a unique digital identity for each participating node, bind the physical entity of the node with the digital identity through encryption technology, and store it in the blockchain to ensure the authenticity and anti-counterfeiting of the identity. In addition, the behavior data of the node is recorded and evaluated in real time to build a node dynamic trust scoring mechanism. Key information such as the node's computing behavior data and trust evaluation results can be permanently stored in the blockchain, and the integrity and security of the data can be guaranteed through a distributed consensus mechanism. Combined with the historical behavior data recorded in the blockchain, abnormal behavior nodes can be detected through intelligent analysis methods, and isolation or removal mechanisms can be quickly triggered to ensure the overall reliability of distributed learning. It provides a high-security, high-reliability and high-transparency management method for distributed learning scenarios, which not only improves the training quality of the global model, but also enhances the system's defense capabilities against malicious node attacks.
[0035] Nevertheless, the fundamental reason why the global model generated by the aggregation of existing model distributed training methods is not reliable enough is that the existing trust management mechanism is limited to the trust evaluation method within the blockchain itself, and does not take into account the training behavior data generated by the local models distributed in each node during the iteration process. The training behavior data is used as the basis for evaluating the trust value of the node, and the models of high-trust value computing nodes are aggregated to generate a global model, thereby improving the security and reliability of distributed training computing nodes, and then improving the training quality of the global model.
[0036] Reference Figure 1 , shows a model distributed training method based on blockchain node trust management provided by an embodiment of the present application, the method involves a user terminal and a computing node; the computing node is used to store blockchain and train distributed models; the computing node includes a master node and a slave node; the user terminal pre-stores an initial model and a training data set; the user terminal is used to send the initial model and the training data set to the computing node; The method comprises the following steps: S110, the computing node trains an initial model according to the training data set, and determines local training behavior data; S120, the computing node determines a training trust value based on local training behavior data; S130, the computing node trains the initial model based on the training data set and the training trust value; wherein, when the initial model training is completed, the master node obtains the trained model from the slave node to aggregate and generate a global model, and sends the global model to the user end.
[0037] In the embodiments of the present application, in view of the insufficient reliability of the global model generated by the aggregation of the existing model distributed training method, the present application provides a solution of "using the training behavior data generated by the local model of each computing node in the distributed training process to determine the trust value of each computing node, and then aggregating the global model based on trust management", specifically: the computing node trains the initial model based on the training data set and determines the local training behavior data; the computing node determines the training trust value based on the local training behavior data; the computing node trains the initial model based on the training data set and the training trust value; wherein, when the initial model training is completed, the master node obtains the trained model from the slave node to aggregate and generate a global model, and sends the global model to the user end. By determining the training trust value based on the local training behavior data, the computing nodes are trained based on the trust management model, and then the models of the high trust value computing nodes are aggregated to generate a global model, thereby improving the security and reliability of the distributed training computing nodes, and then improving the training quality of the global model.
[0038] Below, a model distributed training method based on blockchain node trust management in this exemplary embodiment will be further described.
[0039] It should be noted that the user end can be a model user. The user end can collect and generate a large amount of data for the machine learning model, but it usually lacks sufficient computing resources and needs to rely on computing nodes with sufficient computing resources. Therefore, the user end needs to send the initial model and training data set to the computing node in advance; the computing node can perform machine learning computing tasks, but there is usually a risk of malicious behavior. The training data set can be multiple independent and identically distributed data sets, which are used to be sent to different computing nodes respectively.
[0040] The blockchain can store the calculation results, trust assessment data and related interaction information of the node, and use the distributed storage and network-wide consensus mechanism of the blockchain to ensure the integrity, immutability and transparency of the data. The permanence of block records ensures the reliability of trust management and provides a basis for subsequent node behavior audits. The blockchain can include a node identity blockchain and a node behavior blockchain, wherein the node identity blockchain is used to store static information of registered nodes, including digital identity and device type; the node behavior blockchain is used to store node calculation results and trust data during training. The blockchain can support traceable audit functions, and can use the trust assessment results recorded in the blockchain to trace the historical behavior of the node, and provide immutable proof of all relevant data in the event of a dispute.
[0041] Each computing node needs to generate a unique digital identity for the nodes participating in distributed training in advance to achieve the mapping between the node's physical entity and digital identity, so as to complete the node registration, ensure the authenticity and uniqueness of the node identity, and lay the foundation for trust management. The node registration steps can specifically include: using the device type and unique identifier (such as device ID) to generate a unique digital identity for the node through a hash algorithm, and then binding the digital identity information to the node's physical entity and storing it in the node identity blockchain. Node registration and identity authentication can be automatically triggered by distributed smart contracts.
[0042] In computing nodes, the master node is a relative concept of the slave node. It can be combined with the proof-of-stake consensus mechanism, with the node trust value as the stake. The node with the largest stake is elected as the master node. The master node collects the node trust evaluation results and related data to organize the writing of data blocks.
[0043] As described in step S110, the computing node trains the initial model according to the training data set and determines the local training behavior data.
[0044] It should be noted that the local training behavior data may represent data generated during the model training process and associated with the training process.
[0045] In one embodiment of the present application, the local training behavior data includes the number of local iterations and the local iteration step; the specific process of step S120 "the computing node determines the training trust value based on the local training behavior data" can be further explained in combination with the following description.
[0046] As described in the following steps, the computing node determines the training trust value according to the local iteration number and the local iteration step size.
[0047] It should be noted that the training trust value can be the local trust value of the node, or it can be a comprehensive trust value obtained by weighted calculation of the local trust value, recommended trust value and historical trust value of the node. It should be understood that the training trust value indicates that the trust value is associated with the training behavior of the node in trust management.
[0048] This application conducts trust evaluation on computing nodes by training trust values, and can dynamically call the trust evaluation algorithm and update the trust value of the node in real time through the distributed smart contract of the blockchain. The trust of the node can be dynamically evaluated by comprehensively analyzing the node's historical behavior, contribution, computing power and other information, and filtering based on the trust value threshold. The trust evaluation results and node behavior data can be stored in the blockchain, and blockchain technology can be used to ensure the integrity and security of the data. The trust evaluation mechanism can adopt a weighted scoring method, combined with the node's behavior pattern, to quantify the node's credibility and prevent malicious nodes from interfering with the training process.
[0049] In a specific implementation, other parameters related to the machine model can also be introduced when calculating the training trust value. The training data generated by the node training and the training trust value can establish a quantitative relationship, wherein the training trust value can be the local trust value of the calculation node i , can be calculated by the following formula:
[0050] in, is the number of local iterations of the model in computing node i, is the local iteration step size of the model execution, and are parameters related to the machine model.
[0051] In an embodiment of the present application, the specific process of "the computing node determining the training trust value based on the local iteration number and the local iteration step size" can be further explained in combination with the following description.
[0052] As described in the following steps, the computing node generates a local trust value according to the local iteration number and the local iteration step size; The computing node broadcasts the local trust value and the trained model to other nodes respectively; The computing node calculates a recommended trust value corresponding to the computing node performing broadcasting based on the training data set of the node, the local trust value received through broadcasting, and the trained model received through broadcasting, and broadcasts the recommended trust value to other nodes; The computing node generates the training trust value according to the local trust value and the recommended trust value of the node by other nodes.
[0053] It should be noted that the local trust value represents the trust value that the node evaluates itself, and the recommended trust value represents the trust value that other nodes evaluate the node. That is, the same computing node can have multiple trust values for its trust evaluation. The consistency between the calculation results provided by the node and the evaluation results of other nodes can be analyzed, thereby monitoring abnormal behavior patterns of the node.
[0054] In one specific implementation, computing node i will complete the local model trained locally And local trust value Broadcast to other nodes in the computing group, and other nodes (such as computing node j) use their own data sets to evaluate the received model. Test and calculate the model accuracy to get the recommended trust value (for example, the recommended trust value of computing node j for computing node i) ), and then broadcast the recommended trust value to other nodes in the group. Thus, all nodes in the computing group can obtain the recommended trust value of other nodes for this node.
[0055] In a specific implementation, after generating multiple training trust values, the training trust value generated this time can also be calculated in combination with the historically calculated training trust value, that is, the historical trust value. The computing node can calculate the training trust value based on the local trust value. , recommended trust value And the historical trust value, weighted calculation of the trust value of the computing node i in the tth round , the calculation formula can be expressed as:
[0056] in, and is the weight factor.
[0057] As described in step S130, the computing node trains the initial model based on the training data set and the training trust value; wherein, when the initial model training is completed, the master node obtains the trained model from the slave node to aggregate and generate a global model, and sends the global model to the user end.
[0058] It should be noted that the global model can be aggregated by calling the FedAvg method. Compared with the traditional federated learning framework, the user end of this application needs to first send the training data set and pre-trained model to the computing node for model training, and then the computing node returns the global model after training aggregation to the user end. Therefore, after obtaining the global model, the master node does not need to send the global model to other computing nodes. The master node can directly send the global model to the user end to enable related machine learning applications.
[0059] During the process of local training of the model by the computing node according to the preset number of iterations, the training trust value of each node, the local initial model of each node, and the computing nodes that will subsequently participate in model aggregation will undergo multiple rounds of update iterations. After the last round of trust management is completed, the master node will store the local model of each computing node in the local node in the node behavior blockchain. Therefore, before finally aggregating the local models to generate a global model, the present application performs multiple rounds of trust management on the computing nodes, thereby avoiding the influence of local models generated by nodes with lower trust values on the generation of the global model.
[0060] In one embodiment of the present application, the computing node pre-stores a deviation threshold, a trust value penalty item and a trust value threshold; the specific process of step S130 "the computing node trains the initial model based on the training data set and the training trust value" can be further explained in combination with the following description.
[0061] As described in the following steps, the computing node updates the training trust value based on the local trust value, the recommended trust value of other nodes for the node, the deviation threshold and the trust value penalty item; The computing nodes select computing nodes participating in model training according to the training trust value and the trust value threshold; The computing nodes participating in the model training train the initial model according to the training data set.
[0062] It should be noted that for nodes with large deviations between their own evaluation and other nodes' trust evaluation of themselves, the deviation threshold is used to check whether the deviation value reaches the preset threshold. If the deviation threshold is reached, the node will be regarded as having exaggerated trust or spreading false calculation results, that is, an abnormal behavior node, thereby triggering the trust value penalty mechanism. The training trust value of the node is adjusted through the trust value penalty item to downgrade the trust of the node, and then the trust value threshold is used to screen out computing nodes with higher updated trust values.
[0063] By successively introducing the deviation threshold and the trust value threshold, computing nodes with insufficient trust values can be excluded from the computing group participating in model training. The local models trained by these excluded computing nodes will not participate in the subsequent global model aggregation. That is, the computing nodes to which the local models participating in the global model aggregation belong all have a high degree of credibility, avoiding fraudulent behavior of computing nodes in the trust calculation process and preventing nodes with low trust values from remaining in the computing group, so as to further ensure the reliability of the global model.
[0064] In a specific implementation, for training trust value Nodes with a trust value less than the threshold , will be removed from the computing group, thereby screening out the computing nodes that will participate in model training and global aggregation in the future. The above removal mechanism can be implemented through the distributed smart contract of the blockchain. The triggering condition of the above removal mechanism can be expressed as:
[0065] In one embodiment of the present application, the specific process of "the computing node updates the training trust value based on the local trust value, the recommended trust value of other nodes for the node, the deviation threshold and the trust value penalty item" can be further explained in combination with the following description.
[0066] As described in the following steps, the computing node selects the computing node to be punished based on the local trust value, the recommended trust value of other nodes to the node and the deviation threshold; The computing node to be punished updates the training trust value according to the training trust value and the trust value penalty term.
[0067] In a specific implementation, for nodes that are considered to exaggerate trust or spread false calculation results, their trust values can be assigned corresponding trust value penalties , the calculation formula can be expressed as:
[0068] In one embodiment of the present application, the specific process of "the computing node screening out the computing nodes to be punished based on the local trust value, the recommended trust value of other nodes to the node and the deviation threshold" can be further explained in combination with the following description.
[0069] As described in the following steps, the computing node calculates the trust deviation value based on the local trust value and the recommended trust value of other nodes for the node; The computing nodes select computing nodes to be punished according to the trust deviation value and the deviation threshold.
[0070] In one embodiment of the present application, the specific process of "the computing node calculating the trust deviation value based on the local trust value and the recommended trust value of other nodes for the node" can be further explained in combination with the following description.
[0071] As described in the following steps, the computing node calculates the trust average value based on the local trust value and the recommended trust value of the node by other nodes; The computing node calculates the trust deviation value according to the local trust value and the trust average value.
[0072] In a specific implementation, for the node local trust value Recommend trust value with other nodes The average value is greater than the deviation threshold Nodes with will be considered to exaggerate trust or spread false calculation results. The calculation formula can be expressed as:
[0073] Where I represents the total number of computing nodes in the current computing group, and the polynomial on the left side of “>” represents the trust deviation value of computing node i. If the above formula is true, computing node i will be regarded as a node that exaggerates trust or spreads false computing results and will be a computing node to be punished.
[0074] Reference Figure 2-3 In a specific embodiment of the present application, the node trust management and model aggregation process of the blockchain of the present application may specifically include the following steps: Step 1: There is a set of model users with training data in the distributed system And computing nodes The master node is pre-assigned in the computing group composed of computing nodes, and the nodes participating in distributed training submit a unique identifier to the master node. (such as device ID) and related attribute information (such as device type).
[0075] Step 2: The master node generates a unique digital identity for the node through a hash algorithm , bind the node's digital identity to the physical entity, and store the registration information on the node identity blockchain to ensure the uniqueness and anti-counterfeiting of identity authentication.
[0076] Step 3: After completing the node identity information registration, the model user sends the pre-trained model and independent and identically distributed data set to the legally registered node, and the node performs local model training based on the received model and data set.
[0077] Step 4: After the pre-specified training time is completed, each node calculates the local trust value based on its own behavior data, including the quality of the calculation results, participation frequency, etc. To illustrate the quality of local model training, it can be expressed as:
[0078] in is the number of local iterations of computing node i, is the step size for performing local iterations, and are parameters related to the machine model.
[0079] Step 5: The node will complete the local model trained locally And local trust value Broadcast to other nodes in the group, and other nodes test the model in their own data sets, calculate the model accuracy, and thus obtain the recommended trust value of node j for node i , and broadcast it to other nodes in the group.
[0080] Step 6: Nodes based on local trust value , recommended trust value And the historical trust value, weighted calculation of the trust value of this round , which can be expressed as:
[0081] in and is the weight factor.
[0082] Step 7: Pass the node local trust value Recommend trust value with other nodes , analyze the consistency of the calculation results provided by the node and the evaluation results of other nodes, so as to monitor the abnormal behavior pattern of the node. Specifically, the deviation threshold is introduced , for the node local trust value Recommend trust value with other nodes Nodes whose average value is greater than the deviation threshold will be considered to exaggerate trust or spread false calculation results, which can be expressed as:
[0083] Step 8: For nodes that are considered to exaggerate trust or spread false calculation results, their trust values are assigned corresponding trust value penalties , which can be expressed as:
[0084] Step 9: Further introduce trust value threshold , for the node trust value Nodes with a trust value less than the threshold will be removed from the calculation group, which can be expressed as:
[0085] Step 10: Nodes remaining in the computing group are divided according to their trust values. size, re-elect the group master node, and the node with the highest trust value will be selected as the master node.
[0086] Step 11: Calculate the trust value of other nodes in the group And local training model Sent to the master node. The master node packages the received data to generate candidate blocks , and broadcast it to other nodes in the group for verification.
[0087] Step 12: Other nodes in the group check the candidate blocks After verification, the verification information is sent to the master node. After receiving the verification information of 2 / 3 of the group nodes, the master node completes the block verification and writes it into the node behavior blockchain.
[0088] Step 13: After the block is written, the nodes in the group download the local model completed in the previous round from the blockchain , and continue to train the local model. Repeat steps 4-12 until the predetermined round is reached, and the master node aggregates the local nodes stored in the node behavior blockchain after the last round of trust management method execution to obtain the global model , downloaded to the model user, completing the distributed training process.
[0089] It should be noted that in response to the security and reliability issues of node identity management and trust assessment in existing distributed training, this application combines the decentralization and immutability of blockchain technology to achieve dynamic trust management of distributed training nodes, solves the security and reliability issues of node trust management in distributed training, effectively realizes malicious node detection, data tampering protection and trust assessment quantification, and also builds a continuous supervision mechanism for nodes participating in the distributed training process, so that malicious nodes can be identified and eliminated more timely and efficiently, thereby further improving the stability and performance of the system. This application also provides a general solution for trusted computing in data-driven intelligent systems, which effectively guarantees the security and stable operation of distributed training.
[0090] The above is a description of the method embodiment of the present application. As for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0091] Reference Figure 4 , shows a model distributed training device based on blockchain node trust management provided by an embodiment of the present application, the device involves a user terminal and a computing node; the computing node is used to store blockchain and train distributed models; the computing node includes a master node and a slave node; the user terminal pre-stores an initial model and a training data set; the user terminal is used to send the initial model and the training data set to the computing node; The device comprises: A first training module 410 is used to train an initial model based on a training data set and determine local training behavior data; A trust value generating module 420, for determining a training trust value based on local training behavior data; The second training module 430 is used to train the initial model based on the training data set and the training trust value; wherein, when the initial model training is completed, the master node obtains the trained model from the slave node to aggregate and generate a global model, and sends the global model to the user end.
[0092] In an embodiment of the present application, the local training behavior data includes the number of local iterations and the local iteration step; the trust value generation module 420 includes: The first training trust value generating submodule is used to determine the training trust value according to the number of local iterations and the local iteration step size.
[0093] In one embodiment of the present application, the training trust value generation submodule includes: A local trust value generating unit, configured to generate a local trust value according to the local iteration number and the local iteration step size; A local trust value broadcasting unit, used to broadcast the local trust value and the trained model to other nodes; A recommended trust value broadcasting unit, configured to calculate a recommended trust value corresponding to a computing node performing broadcasting based on a training data set of the node, a local trust value received through broadcasting, and a trained model received through broadcasting, and broadcast the recommended trust value to other nodes; The training trust value generating unit is used to generate the training trust value according to the local trust value and the recommended trust value of other nodes for the node.
[0094] In an embodiment of the present application, the computing node pre-stores a deviation threshold, a trust value penalty item, and a trust value threshold; the second training module 430 includes: A training trust value updating unit, used to update the training trust value according to the local trust value, the recommended trust value of other nodes for the node, the deviation threshold and the trust value penalty item; A participating computing node screening unit is used to screen computing nodes participating in model training based on local trust values and trust value thresholds; The model training unit is used to train the initial model according to the training data set.
[0095] In one embodiment of the present application, the training trust value updating unit includes: A penalty computing node screening subunit is used to screen out computing nodes to be punished based on the local trust value, the recommended trust value of other nodes to this node, and the deviation threshold; The training trust value updating subunit is used to update the training trust value according to the training trust value and the trust value penalty item.
[0096] In one embodiment of the present application, the penalty calculation node screening subunit includes: A trust deviation value calculation module is used to calculate the trust deviation value based on the local trust value and the recommended trust value of other nodes for this node; The penalty computing node screening module is used to screen out computing nodes to be punished based on the trust deviation value and the deviation threshold.
[0097] In one embodiment of the present application, the trust deviation value calculation module includes: The trust average calculation submodule is used to calculate the trust average based on the local trust value and the recommended trust value of other nodes for this node; The trust deviation value calculation submodule is used to calculate the trust deviation value based on the local trust value and the trust average value.
[0098] As for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0099] An embodiment of the present application provides a model distributed training system based on blockchain node trust management, the system comprising a user terminal and a computing node; the computing node is used to store blockchain and train distributed models; the computing node comprises a master node and a slave node; the user terminal pre-stores an initial model and a training data set; The system comprises: The user terminal is used to send the initial model and the training data set to the computing node; The computing node is used to train the initial model based on the training data set, and to determine the local training behavior data, and to determine the training trust value based on the local training behavior data, and to train the initial model based on the training data set and the training trust value; wherein, when the initial model training is completed, the master node is used to obtain the trained model from the slave node to aggregate and generate a global model, and to send the global model to the user end.
[0100] Reference Figure 5 , shows a block diagram of a computer device provided in an embodiment of the present application. The computer device 12 is suitable for implementing the embodiment of the present invention, and may specifically include the following: The computer device 12 is in the form of a general-purpose computing device. The components of the computer device 12 may include, but are not limited to: one or more processing units 16, a system memory 28, and a bus 18 connecting different system components (including the system memory 28 and the processing unit 16). The computer device 12 may be a device attached to the bus.
[0101] Bus 18 represents one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, a processor or a local bus using any of a variety of bus architectures. By way of example, these architectures include, but are not limited to, an Industry Standard Architecture (ISA) bus, a Micro Channel Architecture (MCA) bus, an Enhanced ISA bus, a Video Electronics Standards Association (VESA) local bus, and a Peripheral Component Interconnect (PCI) bus.
[0102] The computer device 12 typically includes a variety of computer system readable media. These media can be any available media that can be accessed by the computer device 12, including volatile and non-volatile media, removable and non-removable media.
[0103] System memory 28 may include computer system readable media in the form of volatile memory, such as RAM 30 (random access memory) and / or cache 32. Computer device 12 may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, storage system 34 may be used to read and write to non-removable, non-volatile magnetic media (commonly referred to as a "hard drive"). Although Figure 5 Not shown, a disk drive for reading and writing to a removable non-volatile disk (e.g., a "floppy disk"), and an optical disk drive for reading and writing to a removable non-volatile optical disk (e.g., a CD-ROM, a DVD-ROM, or other optical media) may be provided. In these cases, each drive may be connected to the bus 18 via one or more data medium interfaces. The system memory 28 may include at least one program product having a set (e.g., at least one) of program modules that are configured to perform the functions of the various embodiments of the present invention.
[0104] A program / utility 40 having a set (at least one) of program modules 42 may be stored, for example, in system memory 28, such program modules 42 including, but not limited to, an operating system, one or more application programs, other program modules, and program data, each of which or some combination may include an implementation of a network environment. Program modules 42 generally perform the functions and / or methods of the embodiments described herein.
[0105] The computer device 12 may also communicate with one or more external devices 14 (e.g., keyboard, pointing device, display 24, etc.), one or more devices that enable a user to interact with the computer device 12, and / or any device that enables the computer device 12 to communicate with one or more other computing devices (e.g., network card, modem, etc.). Such communication may be performed through an I / O interface 22 (input / output interface). Furthermore, the computer device 12 may also communicate with one or more networks (e.g., local area network (LAN), wide area network (WAN) and / or public network (e.g., Internet)) through a network adapter 20. Figure 5 As shown, the network adapter 20 communicates with other modules of the computer device 12 via the bus 18. It should be understood that although Figure 5 Not shown, other hardware and / or software modules may be used in conjunction with computer device 12, including but not limited to microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.
[0106] The processing unit 16 executes various functional applications and data processing by running the programs stored in the system memory 28, such as implementing a model distributed training method based on blockchain node trust management provided by any embodiment of the present invention.
[0107] That is, when the program is executed by the processor, it is implemented as follows: the computing node trains the initial model based on the training data set and determines the local training behavior data; the computing node determines the training trust value based on the local training behavior data; the computing node trains the initial model based on the training data set and the training trust value; wherein, when the initial model training is completed, the master node obtains the trained model from the slave node to aggregate and generate a global model, and sends the global model to the user end.
[0108] The computer device 12 is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present invention.
[0109] An embodiment of the present application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements a model distributed training method based on blockchain node trust management as provided in any embodiment of the present application.
[0110] That is, when the program is executed by the processor, it is implemented as follows: the computing node trains the initial model based on the training data set and determines the local training behavior data; the computing node determines the training trust value based on the local training behavior data; the computing node trains the initial model based on the training data set and the training trust value; wherein, when the initial model training is completed, the master node obtains the trained model from the slave node to aggregate and generate a global model, and sends the global model to the user end.
[0111] Computer storage media can use any combination of one or more computer-readable media. Computer-readable media can be computer-readable signal media or computer-readable storage media. Computer-readable storage media can be, for example, but not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices or devices, or any combination of the above. More specific examples (non-exhaustive list) of computer-readable storage media include: an electrical connection with one or more wires, a portable computer disk, a hard disk, a RAM, a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable CD-ROM, an optical storage device, a magnetic storage device, or any suitable combination of the above. In this document, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, device or device.
[0112] Computer-readable signal media may include a data signal propagated in baseband or as part of a carrier wave, which carries a computer-readable program code. Such propagated data signals may take a variety of forms, including, but not limited to, electromagnetic signals, optical signals, or any suitable combination of the above. Computer-readable signal media may also be any computer-readable medium other than a computer-readable storage medium, which may send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device.
[0113] Program code embodied on a computer readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wireline, optical fiber cable, radio frequency (RF), etc., or any suitable combination of the foregoing.
[0114] Computer program code for performing the operation of the present invention may be written in one or more programming languages or a combination thereof, including object-oriented programming languages, such as Java, Smalltalk, C++, and conventional procedural programming languages, such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer via any type of network, including a LAN or WAN, or may be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0115] Although the preferred embodiments of the present application have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the embodiments of the present application.
[0116] Finally, it should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or terminal device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or terminal device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or terminal device including the elements.
[0117] The above is a detailed introduction to the model distributed training method, device and system based on blockchain node trust management provided by the present application. This article uses specific examples to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea; at the same time, for technical personnel in this field, according to the idea of the present application, there will be changes in the specific implementation method and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.
Claims
1. A model distributed training method based on blockchain node trust management, characterized in that: The method involves a user terminal and a computing node; the computing node is used to store blockchain and train a distributed model; the computing node includes a master node and a slave node; the user terminal pre-stores an initial model and a training data set; the user terminal is used to send the initial model and the training data set to the computing node; The method comprises the following steps: The computing node trains an initial model based on the training data set and determines local training behavior data; The computing node determines the training trust value based on the local training behavior data; The computing node trains the initial model based on the training data set and the training trust value; wherein, when the initial model training is completed, the master node obtains the trained model from the slave node to aggregate and generate a global model, and sends the global model to the user end.
2. The method according to claim 1, characterized in that The local training behavior data includes the number of local iterations and the local iteration step size; the step of determining the training trust value according to the local training behavior data by the computing node includes: The computing node determines the training trust value according to the local iteration number and the local iteration step size.
3. The method according to claim 2, characterized in that The step of determining the training trust value by the computing node according to the local iteration number and the local iteration step size comprises: The computing node generates a local trust value according to the local iteration number and the local iteration step size; The computing node broadcasts the local trust value and the trained model to other nodes respectively; The computing node calculates a recommended trust value corresponding to the computing node performing broadcasting based on the training data set of the node, the local trust value received through broadcasting, and the trained model received through broadcasting, and broadcasts the recommended trust value to other nodes; The computing node generates the training trust value according to the local trust value and the recommended trust value of the node by other nodes.
4. The method according to claim 3, characterized in that The computing node pre-stores a deviation threshold, a trust value penalty item, and a trust value threshold; The step of the computing node training the initial model according to the training data set and the training trust value includes: The computing node updates the training trust value based on the local trust value, the recommended trust value of other nodes for the node, the deviation threshold and the trust value penalty item; The computing nodes select computing nodes participating in model training according to the training trust value and the trust value threshold; The computing nodes participating in the model training train the initial model according to the training data set.
5. The method according to claim 4, characterized in that The step of updating the training trust value of the computing node according to the local trust value, the recommended trust value of the node by other nodes, the deviation threshold and the trust value penalty item includes: The computing node selects the computing node to be punished based on the local trust value, the recommended trust value of other nodes to the node and the deviation threshold; The computing node to be punished updates the training trust value according to the training trust value and the trust value penalty term.
6. The method according to claim 5, characterized in that The step of selecting the computing nodes to be punished according to the local trust value, the recommended trust value of the node by other nodes and the deviation threshold comprises: The computing node calculates the trust deviation value based on the local trust value and the recommended trust value of other nodes for the node; The computing nodes select computing nodes to be punished according to the trust deviation value and the deviation threshold.
7. The method according to claim 6, characterized in that The step of calculating the trust deviation value by the computing node according to the local trust value and the recommended trust value of the node by other nodes includes: The computing node calculates the trust average value based on the local trust value and the trust value recommended by other nodes for the node; The computing node calculates the trust deviation value according to the local trust value and the trust average value.
8. The method according to claim 2, characterized in that: The training trust value includes the local trust value of computing node i ; The step of determining the training trust value by the computing node according to the local iteration number and the local iteration step size comprises: The computing node determines the number of local iterations according to the and the local iteration step size The local trust value of the computing node i is determined by the following formula: : in, is the number of local iterations of the model in computing node i, is the local iteration step size of the model execution, and are parameters related to the machine model.
9. The method according to claim 3, characterized in that: The step of generating the training trust value according to the local trust value and the recommended trust value of the node by other nodes comprises: The computing node is based on the local trust value , recommended trust value As well as the historical trust value, the training trust value of the computing node i in the tth round is calculated by the following formula: : in, and is a weight factor, and the historical trust value includes the training trust value of the computing node when the round is less than t.
10. The method according to claim 5, characterized in that The step of updating the training trust value of the computing node to be punished according to the training trust value and the trust value penalty term includes: The computing node to be punished updates the training trust value according to the training trust value and the trust value penalty term through the following formula: in, represents the training confidence value, Represents the trust value penalty item.
11. The method according to claim 6, characterized in that The step of calculating the trust deviation value by the computing node according to the local trust value and the recommended trust value of the node by other nodes includes: The computing node expresses the trust deviation value through the following polynomial based on the local trust value and the recommended trust value of other nodes to the node: in, represents the local trust value, Indicates the recommendation trust value.
12. A model distributed training device based on blockchain node trust management, characterized in that: The device involves a user terminal and a computing node; the computing node is used to store blockchain and train a distributed model; the computing node includes a master node and a slave node; The user terminal pre-stores an initial model and a training data set; The user terminal is used to send the initial model and the training data set to the computing node; The device comprises: A first training module is used to train an initial model based on a training data set and determine local training behavior data; A trust value generation module, used to determine a training trust value based on local training behavior data; The second training module is used to train the initial model based on the training data set and the training trust value; wherein, when the initial model training is completed, the master node obtains the trained model from the slave node to aggregate and generate a global model, and sends the global model to the user end.
13. A model distributed training system based on blockchain node trust management, characterized in that: The system includes a user terminal and a computing node; the computing node is used to store blockchain and train distributed models; The computing nodes include a master node and a slave node; The user terminal pre-stores an initial model and a training data set; The system comprises: The user terminal is used to send the initial model and the training data set to the computing node; The computing node is used to train the initial model based on the training data set, and to determine the local training behavior data, and to determine the training trust value based on the local training behavior data, and to train the initial model based on the training data set and the training trust value; wherein, when the initial model training is completed, the master node is used to obtain the trained model from the slave node to aggregate and generate a global model, and to send the global model to the user end.
14. A computer device, characterized in that: The method comprises a processor, a memory, and a computer program stored in the memory and capable of running on the processor, wherein the computer program implements the method according to any one of claims 1 to 11 when executed by the processor.
15. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 11 is implemented.
Citation Information
Patent Citations
Medical data trusted sharing method based on block chain and federal learning
CN116776373A
Marine Internet of Things data security sharing method under edge computing framework based on federated learning and block chain technology
CN112348204A
Multi-domain DDoS attack detection method and device based on trusted federated learning
CN115102763A
Decentralized federated learning training behavior supervision method based on digital watermarking technology
CN115713126A
Industrial Internet of Things security data sharing method based on block chain and federal learning
CN117421779A
Cited By
Distributed training method and device based on block chain, equipment and medium
CN120811917A
Blockchain-based distributed training method, apparatus, device, and medium
CN120811917B
Data decentration trust distribution method and system based on block chain
CN121000719A