A blockchain-based federated learning processing method and related device
By enabling computing nodes in a blockchain network to compete for and share training data, the problem of limited computing resources for IoT devices is solved, data utilization and model performance are improved, and the security of data interaction is ensured.
Patent Information
- Application Number
- CN202210726244.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-24
- Publication Date
- 2026-08-25
- Estimated Expiration
- 2042-06-24
AI Technical Summary
In IoT scenarios, some IoT devices have limited computing resources, resulting in low data utilization and affecting model performance.
By enabling computing nodes in a blockchain network to compete for and share training data, and utilizing target computing nodes with computing capabilities for model training, data utilization can be improved.
It improves data utilization in the federated learning process, enhances model performance, and ensures the security of data interaction.
Smart Images

Figure CN117332871B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular to a blockchain-based federated learning processing method and related equipment. Background Technology
[0002] As a collaborative machine learning paradigm, federated learning has attracted significant attention from industry and academia in recent years. The typical federated learning process involves training local models on edge devices, then aggregating the updates on all local models at a central server and averaging them as the global model update. Each edge device then retrieves the updated global model from the central server and continues training it using its own data until global model training is complete.
[0003] With the widespread adoption of smart IoT devices, federated learning is increasingly being applied to IoT scenarios. It's understandable that the more data used in a federated learning algorithm for model training, the better the model's performance. However, in IoT scenarios, some IoT devices have limited computing resources. Although these devices possess the training data required by the model, they cannot participate in training. Therefore, in IoT scenarios, due to limited device computing power, data utilization is low, resulting in insufficient performance of the ultimately trained model. Summary of the Invention
[0004] This application provides a blockchain-based federated learning processing method and related equipment, which can improve the data utilization rate of federated learning and thus improve model performance.
[0005] This application provides a blockchain-based federated learning processing method, characterized in that the method is executed by a first computing node in a blockchain network, which further includes a task initiating node and M second computing nodes, where M is a positive integer; the method includes:
[0006] The first computing node receives the federated learning task generated by the task initiating node through the task smart contract, and obtains the training data associated with the federated learning task as the first training data.
[0007] If the first computing node does not have the computing capability for federated learning of the first training data, the first training competition request for the federated learning task will be broadcast to M second computing nodes. Among the M second computing nodes, the second computing node that wins the competition and meets the node credibility condition will be selected as the target computing node. The target computing node has the computing capability for federated learning of the first training data.
[0008] If a first data sharing request is received from the target computing node, the first training data is sent to the target computing node so that the target computing node can train the initial model associated with the federated learning task based on the first training data and the stored training data associated with the federated learning task to obtain the first branch training update parameters; the task initiating node is also used to globally update the initial model through N branch training update parameters; the N branch training update parameters include the first branch training update parameters, where N is a positive integer.
[0009] This application provides a blockchain-based federated learning processing method, characterized in that the method is executed by a task-initiating node in a blockchain network, which further includes W computing nodes, where W is a positive integer; the method includes:
[0010] The task initiating node generates a federated learning task through a task smart contract and broadcasts the task to W computing nodes. If there are computationally restricted nodes among the W computing nodes, these nodes broadcast training competition requests for shared training data to the corresponding communicable computing nodes. The restricted nodes also determine the target communicable computing node that successfully competes for the shared training data and meets the node trustworthiness criteria. The shared training data is the training data associated with the federated learning task already stored in the restricted nodes. The restricted nodes do not possess the computational capability for federated learning of the shared training data.
[0011] Receive federated learning task response requests from Y capability computing nodes, where Y is a positive integer; the Y capability computing nodes include target communicable computing nodes and do not include computing-restricted nodes.
[0012] Based on the Y federated learning task response requests, X training computing nodes are determined from the Y capability computing nodes. A federated learning task issuance instruction carrying the initial model associated with the federated learning task is sent to the X training computing nodes, so that the X training computing nodes train the initial model according to the available training data to obtain branch training update parameters; X is a positive integer; if the X training computing nodes include the target communicable computing node, the available training data corresponding to the target communicable computing node includes shared training data and training data already stored in the target communicable computing node;
[0013] The initial model is globally updated based on the aggregated training update parameters until a target model that meets the training conditions indicated by the federated learning task is obtained; the aggregated training update parameters are generated based on the training update parameters of X branches.
[0014] This application provides a blockchain-based federated learning processing device, characterized in that the federated learning processing device is applied to a first computing node in a blockchain network, the blockchain network further including a task initiating node and M second computing nodes, where M is a positive integer; the federated learning processing device includes:
[0015] The task receiving module is used to receive the federated learning task generated by the task initiating node through the task smart contract, and obtain the training data associated with the federated learning task as the first training data.
[0016] The competition broadcast module is used to broadcast the first training competition request for the federated learning task to M second computing nodes if the first computing node does not have the computing capability for federated learning of the first training data.
[0017] The node determination module is used to determine the second computing node that successfully competes and meets the node credibility condition from among M second computing nodes, and to serve as the target computing node; the target computing node has the federated learning computing capability for the first training data.
[0018] The data sharing module is used to send the first training data to the target computing node if it receives a first data sharing request from the target computing node, so that the target computing node can train the initial model associated with the federated learning task based on the first training data and the stored training data associated with the federated learning task to obtain the first branch training update parameters; the task initiating node is also used to globally update the initial model through N branch training update parameters; the N branch training update parameters include the first branch training update parameters, where N is a positive integer.
[0019] The node determination module includes:
[0020] The competition information receiving unit is used to receive first training competition response information sent by L competing nodes during the competition period; each first training competition response information contains a number of competing digital resources; the L competing nodes are nodes among M second computing nodes that have federated learning computing capabilities for the first training data; L is a positive integer less than or equal to M.
[0021] The trusted node filtering unit is used to remove competing nodes that do not meet the node trustworthiness conditions from L competing nodes, and obtain S trusted competing nodes; S is a positive integer less than or equal to L;
[0022] The node pre-selection unit is used to select the trusted competing node with the highest number of competing digital resources from S trusted competing nodes as the pre-selected computing node;
[0023] The competition confirmation unit is used to send the first training competition confirmation request to the pre-selected computing node;
[0024] The competition confirmation unit is further configured to determine the pre-selected computing node as the target computing node if it receives the first training competition confirmation response information sent by the pre-selected computing node in accordance with the first training competition confirmation request within the confirmation time period.
[0025] Among the L competing nodes, competing node Z is included. i , where i is a positive integer less than or equal to L;
[0026] The aforementioned federated learning processing unit also includes:
[0027] The trusted node determination module is used to obtain a node trusted probability table;
[0028] The trusted node determination module is also used to query the competing node Z from the node trusted probability table. i The probability of credibility;
[0029] The trusted node determination module is also used to determine the trusted node based on the comparison with competing node Z. i Connections between them and competing nodes Z i The reliable probability determines the competing node Z i Credibility;
[0030] The trusted node determination module is also used when there is contention for node Z. i If the credibility of a node is less than the credibility threshold, then the competing node Z is determined. i The node trustworthiness condition is not met;
[0031] The trusted node determination module is also used when there is contention for node Z. i If the credibility of a node is greater than or equal to the credibility threshold, then the competing node Z is determined. i The node trustworthiness condition is met.
[0032] The federated learning task includes the task deadline and initial model information;
[0033] The aforementioned federated learning processing device also includes:
[0034] The computational capacity determination module is used to determine the primary computational resources required for the federated learning task based on the initial model information.
[0035] The computing power determination module is also used to determine that the first computing node does not have the computing power for federated learning of the first training data if the first required computing resources are greater than the available computing resources; available computing resources refer to the idle computing resources of the first computing node.
[0036] The computing power determination module is also used to determine the first training duration corresponding to the federated learning computation for the first training data based on the amount of data in the first training data if the first required computing resources are less than or equal to the available computing resources.
[0037] The computing power determination module is also used to determine that the first computing node has federated learning computing capabilities for the first training data if the first training duration is less than or equal to the task deadline.
[0038] The computing power determination module is also used to determine that if the first training duration is longer than the task deadline, the first computing node does not have the computing power for federated learning of the first training data.
[0039] The blockchain network also includes P third-party computing nodes, where P is a positive integer;
[0040] The aforementioned federated learning processing device also includes:
[0041] The contention request receiving module is used to receive second training contention requests for the federated learning task sent by P third computing nodes respectively; the second training contention request sent by the j-th third computing node includes the data volume of the j-th second training data; the j-th second training data is the training data associated with the federated learning task that has been stored in the j-th third computing node; the j-th third computing node does not have the federated learning computing capability for the j-th second training data; j is a positive integer less than or equal to P;
[0042] The node determination module is used to determine the third computing node that successfully competes among P third computing nodes as the target acquisition node if the first computing node has federated learning computing capabilities for the first training data; the first computing node also has federated learning capabilities for the target training data; the target training data is the second training data corresponding to the target acquisition node;
[0043] The task response module is used to send a federated learning task response request carrying training data information to the task initiating node; the training data information is generated based on the first training data and the target training data;
[0044] The data acquisition module is used to send a second data sharing request to the target acquisition node if it receives a federated learning task issuance instruction from the task initiating node; the federated learning task issuance instruction contains the initial model associated with the federated learning task.
[0045] The model training module is used to receive target training data sent by the target acquisition node according to the second data sharing request, train the initial model according to the first training data and the target training data, and obtain the second branch training update parameters; the task initiation node is also used to perform global updates on the initial model through H branch training update parameters; the H branch training update parameters include the second branch training update parameters, where H is a positive integer.
[0046] The node determination module includes:
[0047] The node selection unit is used to traverse the second training competition requests for the federated learning task sent by P third computing nodes if the first computing node has the federated learning computing capability for the first training data.
[0048] The node selection unit is also used to obtain the second training data G from the second training competition request sent by the kth third computing node. k The amount of data, based on the second training data G k The amount of data is determined for the second training data G. k The second training duration T of the data volume k k is a positive integer less than or equal to P;
[0049] The node selection unit is also used to select the second training duration T. k Add the total training time T to the first training session duration. k总 ;
[0050] The node selection unit is also used if the total training time T k总 If the timeout is less than or equal to the task deadline, then the first computing node is determined to have the capability for the second training data G. k The federated learning computing capability sends the second training competition response information R to the k-th third computing node. k ;
[0051] The node confirmation unit is used to receive second training competition confirmation requests sent by Q third computing nodes respectively during the competition waiting period; the Q third computing nodes are nodes that received the second training competition response information sent by the first computing node, and Q is a positive integer; a second training competition confirmation request contains a target number of digital resources to be competed for;
[0052] The node confirmation unit is used to determine the third computing node corresponding to the second training competition confirmation request with the highest number of target competitive digital resources as the third computing node that has successfully competed, send the second training competition confirmation response information to the third computing node that has successfully competed, and designate the third computing node that has successfully competed as the target acquisition node.
[0053] The aforementioned federated learning processing device also includes:
[0054] The encryption module is used to encrypt the second branch training update parameters based on the received random differential privacy noise when it receives the federated learning task issuance instruction from the task initiating node, thereby obtaining the target encrypted branch training update parameters, and then sending the target encrypted branch training update parameters to the task initiating node. The task initiating node is also used to aggregate the information of the S encrypted branch training update parameters to obtain encrypted aggregated training update parameters, where S is a positive integer. The task initiating node is also used to add the encrypted aggregated training update parameters to the differential privacy key to obtain aggregated training update parameters, and then perform a global update of the initial model based on the aggregated training update parameters. The S encrypted branch training update parameters include the target encrypted branch training update parameters.
[0055] The aforementioned federated learning processing device also includes:
[0056] The data sharing record module is used to obtain the agreed-upon amount of digital resources with the target computing node;
[0057] The data sharing record module is also used to package the agreed number of digital resources, the amount of data in the first training data, and the data sharing relationship with the target computing node into a data sharing record transaction, and cache the data sharing record transaction in the record transaction pool;
[0058] The data sharing record module is also used to reach consensus and upload data sharing record transactions to the blockchain at the same time as receiving the first data sharing request sent by the target computing node;
[0059] The data sharing record module is also used to receive digital resources sent by the target computing node that correspond to the agreed number of digital resources.
[0060] This application provides a blockchain-based federated learning processing device, characterized in that the federated learning processing device is applied to a task initiating node in a blockchain network, the blockchain network further comprising W computing nodes, where W is a positive integer; the federated learning processing device includes:
[0061] The task broadcasting module is used to generate federated learning tasks through task smart contracts and broadcast these tasks to W computing nodes. If there are computationally restricted nodes among the W computing nodes, these restricted nodes broadcast training competition requests for shared training data to the corresponding communicable computing nodes. The restricted nodes also determine the target communicable computing node that successfully competes for the shared training data and meets the node trustworthiness criteria. The shared training data is the training data associated with the federated learning tasks already stored in the restricted nodes. The restricted nodes do not have the computational capability for federated learning of the shared training data.
[0062] The response receiving module is used to receive federated learning task response requests sent by Y capability computing nodes respectively; Y is a positive integer; the Y capability computing nodes include target communicable computing nodes and do not include computing-restricted nodes;
[0063] The training node determination module is used to determine X training computing nodes from Y capability computing nodes based on Y federated learning task response requests, where X is a positive integer.
[0064] The task distribution module is used to send federated learning task distribution instructions to X training computing nodes, carrying the initial model associated with the federated learning task, so that the X training computing nodes can train the initial model according to the available training data and obtain branch training update parameters. If the X training computing nodes include a target communicable computing node, the available training data corresponding to the target communicable computing node includes shared training data and training data already stored in the target communicable computing node.
[0065] The model update module is used to globally update the initial model based on the aggregated training update parameters until a target model that meets the training conditions indicated by the federated learning task is obtained; the aggregated training update parameters are generated based on the training update parameters of X branches.
[0066] One federated learning task response request includes one unit of data resource consumption and one total amount of training data;
[0067] The training node determination module includes:
[0068] The consumption acquisition unit is used to acquire the predicted consumption of digital resources corresponding to the federated learning task.
[0069] The ascending sorting unit is used to sort the data resource consumption of Y units in ascending order according to the response requests of Y federated learning tasks, and obtain the sorted data resource consumption of Y units.
[0070] The training node selection unit is used to traverse the sorted Y units of data resource consumption, sequentially obtain the e-th unit of data resource consumption, and determine the predicted digital resource consumption of the e-th node based on the e-th unit of data resource consumption and the total amount of training data corresponding to the e-th unit of data resource consumption; e is a positive integer less than or equal to Y.
[0071] The training node selection unit is also used to add the capability computing node corresponding to the e-th unit of data resource consumption to the node pre-selection queue.
[0072] The training node selection unit is also used to subtract the current remaining digital resource prediction consumption from the e-th node's digital resource prediction consumption if the predicted consumption of digital resources at the e-th node is less than the current remaining digital resource prediction consumption, and then continue to traverse to obtain the e+1-th unit's data resource consumption.
[0073] The training node selection unit is also used to stop traversing if the predicted consumption of digital resources of the e-th node is greater than or equal to the current predicted consumption of remaining digital resources, or if e equals Y, and to use all X capability calculation nodes in the node pre-selection queue as training calculation nodes.
[0074] The aforementioned federated learning processing device also includes:
[0075] The digital resource recording module is used to sequentially obtain X-1 training computing nodes from the node pre-selection queue;
[0076] The digital resource recording module is also used to use the predicted consumption of node digital resources corresponding to X-1 training computing nodes as the transaction digital resource consumption corresponding to X-1 training computing nodes.
[0077] The digital resource recording module is also used to retrieve the Xth training computing node from the node pre-selection queue.
[0078] The digital resource recording module is also used to take the predicted consumption of the node digital resources corresponding to the Xth training computing node as the transaction digital resource consumption of the Xth training computing node if the predicted consumption of the node digital resources corresponding to the Xth training computing node is less than the target remaining predicted consumption of the node digital resources. The target remaining predicted consumption of the node digital resources refers to the value remaining after subtracting the predicted consumption of the node digital resources corresponding to the X-1 training computing nodes from the predicted consumption of the node digital resources.
[0079] The digital resource recording module is also used to take the target remaining digital resource prediction consumption as the transaction digital resource consumption of the Xth training computing node if the node digital resource prediction consumption is greater than or equal to the target remaining digital resource prediction consumption.
[0080] The digital resource recording module is also used to associate and package each training computing node and the corresponding transaction digital resource consumption into a digital resource recording transaction, and cache the digital resource recording transaction in the recording transaction pool.
[0081] The digital resource recording module is also used to send the digital resources corresponding to the associated transaction digital resource consumption to each training computing node when the target model that meets the training conditions indicated by the federated learning task is obtained, and to perform consensus on the digital resource recording transaction on the blockchain.
[0082] The aforementioned federated learning processing device also includes:
[0083] The noise processing module is used to generate a set of random differential privacy noise; the set of differential privacy noise contains random differential privacy noise corresponding to each training computation node.
[0084] The noise processing module is also used to add X random differential privacy noises to obtain the total random differential privacy noise;
[0085] The noise processing module is also used to take the negative of the total random differential privacy noise as the differential privacy key;
[0086] The noise transmission module is used to send the federated learning task issuance instruction carrying the initial model associated with the federated learning task to X training computing nodes, and at the same time send the corresponding random differential privacy noise to each of the X training computing nodes, so that the X training computing nodes can encrypt the branch training update parameters according to the received random differential privacy noise to obtain encrypted branch training update parameters.
[0087] The parameter aggregation module is used to aggregate the encrypted branch training update parameters returned by X training computing nodes to obtain encrypted aggregated training update parameters.
[0088] The parameter decryption module adds the encrypted aggregate training update parameters to the differential privacy cipher to obtain the aggregate training update parameters.
[0089] The parameter aggregation module includes:
[0090] The accuracy verification unit is used to acquire the training and test sets;
[0091] The accuracy verification unit is also used to verify the encrypted branch training update parameters returned by X training computing nodes according to the training test set, and obtain X update accuracies.
[0092] The target parameter aggregation unit is used to take the encrypted branch training update parameters returned by the training computation node whose update accuracy meets the training accuracy condition as the target encrypted branch training update parameters.
[0093] The target parameter aggregation unit is also used to perform information aggregation processing on the target encrypted branch training update parameters to obtain encrypted aggregated training update parameters.
[0094] The aforementioned federated learning processing device also includes:
[0095] The precision recording module is used to package X training computing nodes and X update precisions into an update precision recording transaction, and initiate consensus processing for the update precision recording transaction to the blockchain network. This enables each node in the blockchain network to update the stored precision recording table according to the X update precisions when the consensus of the update precision recording transaction is passed. The precision recording table is used to record each node in the blockchain network and the current update precision corresponding to each node.
[0096] The aforementioned federated learning processing device also includes:
[0097] The block production right determination module is used to obtain the precision record table based on the vote confirmation request;
[0098] The block production right determination module is also used to determine the contribution of each node based on the current update precision of each node in the precision record table.
[0099] The block production right determination module is also used to determine the resource rights of each node based on the contribution weight, the contribution of each node, and the duration of digital resource ownership for each node.
[0100] The block production right determination module is also used to send a ballot to each node according to the ratio between the resource rights of each node, so that each node can vote on the received ballot;
[0101] The block production right determination module is also used to, if there are computationally restricted nodes in each node, have the computationally restricted nodes vote for the target communicable computing node with the received votes.
[0102] The block production right determination module is also used to receive the number of votes from each node after voting, and determine the block production right acquisition ratio of each node based on the number of votes from each node.
[0103] The block production right determination module is also used to randomly determine the block production node from each node according to the block production right acquisition ratio; the block production node has the right to produce new blocks;
[0104] The block production right determination module is also used to send reward digital resources to the block producing node so that the block producing node can allocate the reward digital resources according to the voting composition ratio; the voting composition ratio refers to the ratio between the number of votes sent by the task initiating node and the number of votes cast by the target computationally restricted node.
[0105] One embodiment of this application provides a computer device, including: a processor, a memory, and a network interface;
[0106] The processor is connected to the memory and the network interface. The network interface is used to provide a data communication network element, the memory is used to store a computer program, and the processor is used to call the computer program to execute the method in the embodiments of this application.
[0107] One aspect of this application provides a computer-readable storage medium storing a computer program adapted for loading by a processor and executing the methods described in this application.
[0108] One aspect of this application provides a computer program product or computer program that includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the method described in this application.
[0109] In this embodiment, when a first computing node in the blockchain network receives a federated learning task generated by a task initiating node through a task smart contract, it obtains training data associated with the federated learning task as first training data. If the first computing node lacks the federated learning computing capability for the first training data, it can broadcast a first training competition request for the federated learning task to M second computing nodes. Among the M second computing nodes, the second computing node that successfully competes and meets the node trustworthiness condition is selected as the target computing node. If a first data sharing request is received from the target computing node, the first training data is sent to the target computing node. The target computing node can then train the initial model associated with the federated learning task based on the first training data and the stored training data associated with the federated learning task to obtain the first branch training update parameters. Here, M is a positive integer, and the target computing node possesses the federated learning computing capability for the first training data. The task initiating node is also used to globally update the initial model using N branch training update parameters. The N branch training update parameters include the first branch training update parameters, where N is a positive integer. Using the method provided in this application embodiment, if the first computing node is computationally limited, it can share its stored first training data with a target computing node that has federated learning computing capabilities for the first training data and meets the node trustworthiness conditions. The target computing node can then use both its own stored training data and the first training data to train the model, avoiding the waste of training data in computationally limited computing nodes, improving the data utilization rate in the federated learning process, and allowing more training data to participate in the training, resulting in a higher performance model. Moreover, since both the first computing node and the target computing node are nodes in the blockchain network and meet the trustworthiness conditions between them, the data interaction will be more secure. Attached Figure Description
[0110] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0111] Figure 1 This is a schematic diagram of the structure of a blockchain network provided in an embodiment of this application;
[0112] Figure 2a This is a schematic diagram of a scenario for broadcasting a federated learning task provided in an embodiment of this application;
[0113] Figure 2b This is a schematic diagram of a scenario for a federated learning processing method provided in an embodiment of this application;
[0114] Figure 2c This is a schematic diagram of a federated learning processing scenario provided in an embodiment of this application;
[0115] Figure 3 This is a flowchart illustrating a blockchain-based federated learning processing method provided in an embodiment of this application.
[0116] Figure 4 This is a schematic diagram of a data structure for recording transactions provided in an embodiment of this application;
[0117] Figure 5 This is a flowchart illustrating a blockchain-based federated learning processing method provided in an embodiment of this application.
[0118] Figure 6 This is a flowchart illustrating a blockchain-based federated learning processing method provided in an embodiment of this application.
[0119] Figure 7 This is a flowchart illustrating a blockchain-based federated learning processing method provided in an embodiment of this application.
[0120] Figure 8 This is a flowchart of a data competition method provided in an embodiment of this application;
[0121] Figure 9 This is a flowchart of a task allocation method provided in an embodiment of this application;
[0122] Figure 10 This is a flowchart of an encrypted model training method provided in an embodiment of this application;
[0123] Figure 11 This is a schematic diagram of the structure of a blockchain-based federated learning processing device provided in an embodiment of this application;
[0124] Figure 12 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application;
[0125] Figure 13 This is a schematic diagram of another blockchain-based federated learning processing device provided in an embodiment of this application;
[0126] Figure 14 This is a schematic diagram of the structure of another computer device provided in an embodiment of this application. Detailed Implementation
[0127] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.
[0128] For ease of understanding, the relevant concepts used in this application will be explained below:
[0129] Federated machine learning (also known as federated learning, consortium learning, or federated learning) is a machine learning framework that effectively helps multiple organizations use data and perform machine learning modeling while meeting user privacy, data security, and government regulations. The typical federated learning process involves training local models on edge devices, then aggregating the updates from all local models at a central server and averaging them as the global model update. Each edge device then retrieves the updated global model from the central server and continues training it using its own data until the global model training is complete.
[0130] Social Internet of Things (SIoT): As the application of the Internet of Things (IoT) advances, IoT technology is combining with social networks. The IoT not only includes connections between things and between things and people, but also incorporates relationships between people, thus better depicting a world where everything is interconnected. Generally, social networks are networks of relationships between people, with people as nodes in the network, organized together by friendships and connections.
[0131] Blockchain-enhanced Federated Learning Market (BFL Market): A distributed social IoT market with fair competition and data trading, where IoT devices in the BFL Market can earn benefits by completing published federated learning tasks or trading data.
[0132] Blockchain is the carrier and organizational method for running blockchain technology. Blockchain technology (BT), also known as distributed ledger technology, is an internet database technology characterized by decentralization and transparency, allowing everyone to participate in database recording. Blockchain technology utilizes a block-chain data structure to verify and store data, distributed node consensus algorithms to generate and update data, cryptography to ensure the security of data transmission and access, and smart contracts composed of automated script code to program and manipulate data—a distributed infrastructure and computing method.
[0133] Blockchain is a novel application model of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanisms, and encryption algorithms. It is primarily used to organize data chronologically and encrypt it into a ledger, making it tamper-proof and forgery-proof. It also allows for data verification, storage, and updating. Essentially, a blockchain is a decentralized database where each node stores an identical copy. The blockchain network can distinguish between consensus nodes and business nodes, with the consensus node responsible for achieving consensus across the entire network. The process of writing transaction data into the ledger in a blockchain network can be summarized as follows: a client sends transaction data to a business node, which then relays this data among the business nodes in the blockchain network until the consensus node receives it. The consensus node then packages the transaction data into a block and reaches a consensus with other consensus nodes. Once consensus is reached, the block containing the transaction data is written into the ledger.
[0134] A block is a data packet that carries transaction data (i.e., transaction business) on a blockchain network. It is a data structure that is marked with a timestamp and the hash value of the previous block. The transactions in the block are verified and confirmed by the network's consensus mechanism.
[0135] Hash value: Also known as information feature value or characteristic value, a hash value is generated by converting input data of arbitrary length into cryptographic data and producing a fixed output through a hash algorithm. The original input data cannot be retrieved by decrypting the hash value; it is a one-way cryptographic function. In a blockchain, each block (except the initial block) contains the hash value of its predecessor block, which is called the parent block of the current block. The hash value is a core and crucial aspect of blockchain technology, preserving the authenticity of recorded and viewed data, as well as the integrity of the blockchain as a whole.
[0136] Transaction: A transaction sent by a blockchain account contains a transaction hash as a unique identifier and includes an account address to identify the blockchain account that sent the transaction.
[0137] Smart contracts: A smart contract can be code that can be understood and executed by all nodes in a blockchain (including consensus nodes), capable of executing arbitrary logic and producing results. It should be understood that a blockchain can include one or more smart contracts, which can be distinguished by an identity document (ID) or name. Transaction requests can carry the identity document or name of the smart contract to specify the smart contract that the blockchain needs to run.
[0138] Computing resources: Computing resources refer to the hardware or network resources required by a device to perform computing tasks. They typically include CPU (central processing unit) resources, GPU (graphics processing unit) resources, memory resources, network bandwidth resources, and disk resources.
[0139] Please see Figure 1 , Figure 1 This is a schematic diagram of the structure of a blockchain network provided in an embodiment of this application. Figure 1 The blockchain network shown may include, but is not limited to, the blockchain network corresponding to a consortium blockchain. This blockchain network may include multiple blockchain nodes, specifically blockchain node 10a, blockchain node 10b, blockchain node 10c, blockchain node 10d, ..., blockchain node 10n. Each blockchain node, during normal operation, can receive data sent from the outside world and perform block-up processing based on the received data, and can also send data to the outside world. To ensure data interoperability between the various blockchain nodes, data connections may exist between each blockchain node; for example, there is a data connection between blockchain node 10a and blockchain node 10b, a data connection between blockchain node 10a and blockchain node 10c, and a data connection between blockchain node 10b and blockchain node 10c.
[0140] It is understandable that blockchain nodes can transmit data or blocks through the aforementioned data connections. The blockchain network can establish data connections between blockchain nodes based on node identifiers. Each blockchain node in the network has a corresponding node identifier, and each blockchain node can store the node identifiers of other blockchain nodes that are connected to it. This allows it to broadcast acquired data or generated blocks to other blockchain nodes based on their node identifiers. For example, blockchain node 10a can maintain a node identifier list as shown in Table 1, which stores the node names and node identifiers of other nodes.
[0141] Table 1
[0142] Node 10a AAA.AAA.AAA.AAA Node 10b BBB.BBB.BBB.BBB Node 10c CCC.CCC.CCC.CCC Node 10d DDD.DDD.DDD.DDD … … Node 10n EEE.EEE.EEE.EEE
[0143] The node identifier can be an Internet Protocol (IP) address or any other information that can be used to identify a node in a blockchain network. Table 1 only uses IP addresses as an example. For instance, blockchain node 10a can send information (e.g., a block) to blockchain node 10b using the node identifier BBB.BBB.BBB.BBB, and blockchain node 10b can determine that the information was sent by blockchain node 10a using the node identifier AAA.AAA.AAA.AAA.
[0144] In blockchain, before a block is added to the chain, it must pass consensus among the consensus nodes in the blockchain network. Only after consensus is reached can the block be added to the blockchain. Understandably, when blockchain is used in scenarios involving government or commercial institutions, not all participating nodes in the blockchain (i.e., the blockchain nodes in the aforementioned blockchain node system) have sufficient resources and the necessity to become consensus nodes. For example, in... Figure 1 In the blockchain network shown, blockchain nodes 10a, 10b, 10c, and 10d can be considered as consensus nodes. Consensus nodes participate in consensus, which means reaching an agreement on blocks (containing a batch of transactions), including generating blocks and voting on them. Non-consensus nodes do not participate in consensus but help propagate block and voting messages, and synchronize states with each other.
[0145] In such Figure 1 In the blockchain network shown, some blockchain nodes (whether or not they are consensus nodes) can be IoT devices. These IoT device blockchain nodes can form a blockchain-enhanced social IoT federated learning market, for example, in... Figure 1 In the blockchain network shown, blockchain nodes 10a, 10b, 10c, and 10d are all IoT devices. Therefore, blockchain nodes 10a, 10b, 10c, and 10d can constitute a blockchain-enhanced social IoT federated learning market. In this market, any IoT device can act as a task initiating node, generating a federated learning task through a task smart contract on the blockchain. This task is then broadcast to the computing node cluster (i.e., other IoT devices in the federated learning market). Each computing node in the cluster can determine whether to respond to the federated learning task based on its own computing resources, or share its stored training data with other computing nodes that have abundant computing resources.
[0146] Taking the first computing node in the computing node cluster as an example, the first computing node receives the federated learning task generated by the task initiating node through the task smart contract, and can obtain the training data associated with the federated learning task as the first training data. If the first computing node does not have the federated learning computing capability for the first training data, that is, the computing resources of the first computing node are insufficient to complete the federated learning task, it can broadcast the first training competition request for the federated learning task to M second computing nodes (which can be computing nodes within the communication range of the first computing node). Then, the first computing node can determine the second computing node that wins the competition and meets the node trustworthiness condition from the M second computing nodes as the target computing node. It should be noted that the target computing node has the federated learning computing capability for the first training data. Then, if the first computing node receives the first data sharing request sent by the target computing node, it sends the first training data to the target computing node. Then, the target computing node can train the initial model associated with the federated learning task based on the first training data and the stored training data associated with the federated learning task, obtain the first branch training update parameters, and then send them to the task initiating node. In addition to receiving the first branch training update parameters returned by the target computing node, the task initiating node also receives branch training update parameters returned by other computing nodes. Finally, the task initiating node performs a global update on the initial model corresponding to the federated learning task using the received N branch training update parameters. Here, M and N are both positive integers.
[0147] It is understood that the above-mentioned data connection is not limited to the connection method. It can be connected directly or indirectly through wired communication, or directly or indirectly through wireless communication, or through other connection methods. This application does not impose any restrictions on this.
[0148] It is understood that the federated learning processing method provided in this application embodiment can be executed by IoT devices, including but not limited to servers and terminal devices. The aforementioned servers can be independent physical servers, server clusters or distributed systems composed of multiple physical servers, or cloud servers providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The aforementioned terminals can be smartphones, tablets, laptops, desktop computers, smart speakers, smartwatches, etc., but are not limited to these.
[0149] It is understood that the embodiments of this application can be applied to various scenarios, including but not limited to smart IoT, artificial intelligence and other scenarios.
[0150] It is understood that, in the specific embodiments of this application, the training data and other related data involved need to obtain user permission or consent when the above embodiments of this application are applied to specific products or technologies, and the collection, use and processing of related data need to comply with the relevant laws, regulations and standards of relevant countries and regions.
[0151] To better understand the architecture of the blockchain-enhanced social IoT federated learning market and the federated learning process described above, please refer to [link to relevant documentation]. Figure 2a , Figure 2a This is a schematic diagram illustrating a scenario of federated learning task broadcasting provided in an embodiment of this application. Figure 2a As shown, Federated Learning Market 2000 is a blockchain-enhanced social IoT federated learning market. Federated Learning Market 2000 includes task initiating node 200, computing node 201, computing node 202, and computing node 203. Among them, task initiating node 200, computing node 201, computing node 202, and computing node 203 are all IoT devices, and each node can be one of the aforementioned... Figure 1 Any blockchain node in the blockchain network shown, for example, task initiation node 200, can be one of the above-mentioned nodes. Figure 1 The blockchain node 10b shown can be compute node 201 as described above. Figure 1 The blockchain node 10d shown can be compute node 202 as described above. Figure 1 The blockchain node 10c shown can be compute node 203 as described above. Figure 1 The blockchain node 10a is shown. In other words, the Federated Learning Market 2000 can be implemented on, for example... Figure 1 In the blockchain network shown, all nodes in the Federated Learning Market 2000 are blockchain nodes in the blockchain network. Therefore, each node can store the same blockchain 2001. The interaction process between nodes in the Federated Learning Market 2000 can be carried out in the form of transactions. After consensus is reached by the consensus nodes in the blockchain network, the transactions are recorded in blockchain 2001.
[0152] like Figure 2aAs shown, in the Federated Learning Market 2000, if a task initiating node 200 needs to train an initial model A using federated learning, it can invoke the task smart contract to generate a federated learning task B for that initial model A. Then, federated learning task B is broadcast to all computing nodes in the Federated Learning Market 2000. The task smart contract, belonging to the smart contract of the aforementioned blockchain network, is a piece of code that the task initiating node 200 can understand and execute. It should be noted that in the Federated Learning Market 2000, any node can be a task initiating node, and other nodes can serve as the corresponding computing nodes. Furthermore, the same task initiating node can initiate different federated learning tasks simultaneously, and different task initiating nodes can initiate different federated learning tasks simultaneously. Within the same federated learning market, the processing of different federated learning tasks does not affect each other.
[0153] Upon receiving federated learning task B, each computing node can determine whether to respond to task B based on its stored data and computing resources. For ease of understanding, computing node 201, the first computing node mentioned above, will be used as an example. Please refer to [link to relevant documentation]. Figure 2b , Figure 2b This is a schematic diagram illustrating a scenario of a federated learning processing method provided in an embodiment of this application. For example... Figure 2bAs shown, after receiving the federated learning task B generated by the task initiating node 200 through the task smart contract, computing node 201 can obtain the training data C associated with the federated learning task B stored locally. It then determines whether it has the computational capability for federated learning of the training data C. Simply put, computing node 201 predicts the required computational resources for training the initial model A associated with the federated learning task B, and then queries its available computational resources. If the available computational resources are less than the required computational resources, computing node 201 determines that it does not have the capability for federated learning of the training data C. In this case, computing node 201 can be called a computationally constrained node and cannot directly respond to the federated learning task B. However, to improve the utilization rate of data in the federated learning market 2000, computing node 201 can share the training data C with computing nodes that have abundant computational resources, allowing these nodes to have more training data to complete the federated learning task B. Correspondingly, computing nodes that obtain the training data C from computing node 201 need to provide computing node 201 with corresponding digital resources. To obtain more digital resources, computing node 201 can generate a training contention request D and then broadcast it to other computing nodes, such as computing nodes 202 and 203. The training contention request D can be considered an auction request; upon receiving it, other computing nodes can choose whether to compete for the right to share the training data C based on the availability of their own computing resources. It can be understood that when the number of computing nodes in the federated learning market 2000 is small, the training contention request from a computationally limited node can be broadcast to the remaining computing nodes; when the number of computing nodes in the federated learning market 2000 is large, the training contention request from a computationally limited node can be broadcast only to a subset of nearby computing nodes.
[0154] like Figure 2bAs shown, after receiving the training competition request D, if computing node 202 determines that it has federated computing capabilities for both locally stored training data and training data C, it can generate a training competition response message E and return it to computing node 201. The training competition response message E contains the number of competing digital resources e, informing computing node 201 that if computing node 201 is willing to share training data C, it can send the digital resources corresponding to the number of competing digital resources e to computing node 201. Similarly, computing node 203 can send a training competition response message F to computing node 201 carrying the number of competing digital resources f. Computing node 201 will intelligently select the computing node that successfully competes and meets the node trustworthiness condition from computing nodes 202 and 203, and use it as the target computing node. The node trustworthiness condition refers to the requirement that computing node 201's trustworthiness towards other computing nodes must exceed a trustworthiness threshold. As mentioned above, the Federated Learning Market 2000 belongs to a social IoT scenario; therefore, the relationships between computing nodes are similar to those between people, with varying degrees of closeness or distance. That is, for computing node 201 in the Federated Learning Market 2000, its trustworthiness towards other computing nodes is not uniform. For example, computing node 201 may have high trustworthiness towards some computing nodes with whom it frequently interacts with data, while its trustworthiness towards some computing nodes with whom it has never interacted may be low. The computing node that wins the competition is usually the one that offers the largest number of competing digital resources.
[0155] Assuming that the target computing node determined by computing node 201 is computing node 202, the available training data for computing node 202 for the federated learning task includes the locally stored training data and training data C. Figure 2b As can be seen, when a computing node in the Federated Learning Market 2000 receives Federated Learning Task B, if its computing resources are limited, it can share its stored training data with a trusted computing node that has abundant computing resources; if its computing resources are abundant, it can compete for the training data of other computing nodes with limited computing resources. After the competition for training data among the computing nodes is completed, the computing node with abundant computing resources and available training data can act as a capable computing node and send a Federated Learning Task response request to the task initiating node 200. For ease of understanding, please refer to [link to relevant documentation]. Figure 2c , Figure 2c This is a schematic diagram illustrating a scenario of federated learning processing provided in an embodiment of this application. For example... Figure 2cAs shown, assuming that both computing nodes 202 and 203 are capability computing nodes, they can each send a federated learning task response request to the task initiating node 200. Upon receiving the federated learning task response request, the task initiating node 200 can select all or a subset of capability computing nodes from the responding capability computing nodes as training computing nodes. Assuming the task initiating node 200 selects computing nodes 202 and 203 as training computing nodes, it will issue federated learning task issuance instructions carrying the initial model A to both computing nodes 202 and 203 respectively. Figure 2c As shown, because compute node 202 successfully acquired the training data C stored by compute node 201, upon receiving the instruction from the federated learning task, compute node 202 first generates a data sharing request and then sends this request to compute node 201. Upon receiving the data sharing request, compute node 201 sends the training data C to compute node 202. After receiving the training data C, compute node 202 also retrieves the locally stored training data H, and then trains the initial model A using both training data C and training data H. Finally, it obtains the branch training update parameters H and then returns to the task initiating node 200. Figure 2c As shown, since computing node 203 did not compete for training data from other computing nodes, it can directly train the initial model A based on the locally stored training data I, ultimately obtaining the branch training update parameters J, and then returning them to the task initiating node 200. The task initiating node 200 can then perform a global update of the initial model A based on all the received branch training update parameters.
[0156] Further, please see Figure 3 , Figure 3 This is a flowchart illustrating a blockchain-based federated learning processing method provided in this application embodiment. The method can be executed by a first computing node (which can be any blockchain node in the aforementioned blockchain network). The blockchain network also includes a task initiating node (which can be any blockchain node other than the first computing node) and M second computing nodes (which can be any M blockchain nodes other than the first computing node and the task initiating node), where M is a positive integer. The following description uses the execution of this method by the first computing node as an example. This blockchain-based federated learning processing method can include at least the following steps S101-S103:
[0157] Step S101: The first computing node receives the federated learning task generated by the task initiating node through the task smart contract, and obtains the training data associated with the federated learning task as the first training data.
[0158] Specifically, a federated learning task can include information such as initial model information, task deadline, and the type of training data required. The first computing node can first acquire the training data associated with the federated learning task based on the required training data type, using this as the initial training data. Then, the first computing node can determine whether it possesses the federated learning computation capability for the initial training data based on the initial model information, task deadline, and other information. Having the federated learning computation capability for the initial training data means that the first computing node can use the initial training data to complete the training of the initial model associated with the federated learning task within the task deadline.
[0159] Optionally, a feasible implementation process for the first computing node to determine whether it possesses the computational capability for federated learning of the first training data can be as follows: The first computing node can determine the first required computing resources corresponding to the federated learning task based on the initial model information. If the first required computing resources are greater than the available computing resources, then the first computing node is determined not to possess the computational capability for federated learning of the first training data. Here, available computing resources refer to the idle computing resources of the first computing node. If the first required computing resources are less than or equal to the available computing resources, then the first computing node can determine the first training duration corresponding to the federated learning computation of the first training data based on the amount of data in the first training data. If the first training duration is less than or equal to the task deadline, then the first computing node is determined to possess the computational capability for federated learning of the first training data; if the first training duration is greater than the task deadline, then the first computing node is determined not to possess the computational capability for federated learning of the first training data.
[0160] Step S102: If the first computing node does not have the federated learning computing capability for the first training data, the first training competition request for the federated learning task is broadcast to the M second computing nodes. Among the M second computing nodes, the second computing node that successfully competes and meets the node credibility condition is determined as the target computing node; the target computing node has the federated learning computing capability for the first training data.
[0161] Specifically, lacking the computational capability for federated learning of the initial training data means that although the first computing node possesses the initial training data associated with the federated learning task, it lacks the computational resources to train the initial model, or its resources are insufficient to guarantee completion of training on the initial training data within the task deadline. In this case, the first computing node can choose to share the initial training data through a secure and reliable data sharing strategy to gain some benefit. Therefore, the first computing node can broadcast its initial training competition request for the federated learning task to M second computing nodes in the blockchain network. These M second computing nodes can be computing nodes within the communication range of the first computing node in the blockchain network. The initial training competition request can be understood as an auction request, informing the M second computing nodes that they can choose to compete for the initial training data using a certain amount of digital resources. These digital resources refer to virtual digital assets that can be traded within the blockchain network.
[0162] Specifically, a feasible implementation process for the first computing node to select a successful second computing node from among M second computing nodes that meets the node trustworthiness condition as the target computing node can be as follows: During the competition period, receive first training competition response information sent by L competing nodes respectively; each first training competition response information contains a number of competing digital resources; the L competing nodes are nodes among the M second computing nodes that have federated learning computing capabilities for the first training data; L is a positive integer less than or equal to M; from the L competing nodes, eliminate competing nodes that do not meet the node trustworthiness condition to obtain S trustworthy competing nodes; S is a positive integer less than or equal to L; from the S trustworthy competing nodes, select the trustworthy competing node with the highest number of competing digital resources as the pre-selected computing node; send a first training competition confirmation request to the pre-selected computing node; if a first training competition confirmation response information sent by the pre-selected computing node according to the first training competition confirmation request is received within the confirmation period, then the pre-selected computing node is determined as the target computing node.
[0163] It's understandable that among the M second computing nodes, not all nodes necessarily possess sufficient computing resources to complete the training of the first training data. Therefore, the first computing node may only receive first training competition response information from L competing nodes with federated learning computing capabilities for the first training data. This first training competition response information informs the first computing node of the amount of digital resources it is willing to pay for the first training data. The first computing node will naturally prioritize the competing node with the most competing digital resources. However, it should be noted that the first computing node does not consider the trustworthiness of the second computing nodes when broadcasting the first training competition request. In reality, in social IoT scenarios, trust issues exist between nodes; untrusted nodes often do not interact with each other due to potential data security problems. Therefore, among the L competing nodes, there may be untrusted competing nodes. To ensure data security, the first computing node should first eliminate untrusted competing nodes to obtain trusted competing nodes, and then select the trusted competing node with the highest competing resources.
[0164] Specifically, whether a competing node is trustworthy can be determined by the node trustworthiness condition. A competing node that meets the node trustworthiness condition is trustworthy. The node trustworthiness condition can be that the first computing node's trustworthiness of the competing node is greater than the trustworthiness threshold.
[0165] Optionally, assume that among the L competing nodes is competing node Z. i Where i is a positive integer less than or equal to L, and the first compute node determines Z. i A feasible implementation process for determining whether the node trustworthiness condition is met can be as follows: obtain the node trustworthiness probability table; query the competing node Z from the node trustworthiness probability table. i The reliability probability; based on the competition node Z i Connections between them and competing nodes Z i The reliable probability determines the competing node Z i The credibility of the competing node Z; i If the credibility of a node is less than the credibility threshold, then the competing node Z is determined. i The node trustworthiness condition is not met; if the competing node Z... i If the credibility of a node is greater than or equal to the credibility threshold, then the competing node Z is determined. i The node trustworthiness condition is met. The node trustworthiness probability table contains the trustworthiness probabilities of the first computing node towards its directly adjacent nodes in the blockchain network. These probabilities are divided into 11 levels from 0 to 1, with increments of 0.1. The first computing node can select the trustworthiness probability of a node based on the social trust relationship with it, and then, based on the competing node Z... i Connections between them and competing nodes Z iThe reliable probability determines the competing node Z i Credibility.
[0166] Specifically, the first computing node is determined based on the competing node Z. i Connections between them and competing nodes Z i The reliable probability determines the competing node Z i A feasible implementation process for ensuring credibility can be: if competing node Z i If directly connected to the first computing node, the first computing node can use an entropy-based method to compute the value of the competing node Z. i The credibility between them is calculated using the following formula:
[0167]
[0168] Where p refers to the competing node Z i The reliable probability of the competing node Z when p is between 0 and 0.5. i The credibility is When p is between 0.5 and 1, the competing node Z i The credibility is
[0169] If the competing node Z i It is not directly connected to the first computing node, and the first computing node did not find the competing node Z in the confidence probability table. i When calculating the credible probability, the first computing node can do so by competing with node Z. i The trustworthiness of nodes relied upon in indirect connections is calculated using cascading and multi-path propagation methods to compare with competing nodes Z. i The reliability of the calculation is determined by multiplication rules in cascading and weighted average rules in multipath propagation. For ease of understanding, assume the first computing node corresponds to device a, and the competing nodes are z. i For device d, device a is directly connected to devices b and c, device d is directly connected to devices b and c, and device b is not connected to device c. The formula for calculating the reliability between device a and device d in this case is as follows:
[0170]
[0171] Among them, L ad That is, the trustworthiness between device a and device d, L ab That is, the trustworthiness between device a and device b, because device a and device b are directly connected, therefore L ab Based on the above formula (1), we can obtain L. ac That is, the trustworthiness between device a and device c, L bd That is, the trustworthiness between device b and device d, Lcd That is, the reliability between device c and device d can be determined based on the above formula (1).
[0172] Specifically, after acquiring a trusted competing node with the highest number of competing digital resources, the first computing node will only initially select it as a pre-selected computing node. This is because the same competing node can simultaneously compete for training data corresponding to multiple computationally limited nodes. When the first computing node selects a competing node, the competing node can also determine whether to agree to the first computing node's selection. Therefore, the first computing node also needs to send a first training competition confirmation request to the pre-selected computing node. This first training competition confirmation request can carry an agreed-upon number of digital resources. If the first computing node receives a first training competition confirmation response from the pre-selected computing node within the confirmation period, the first computing node will designate the pre-selected computing node as the target computing node. This means that after the first computing node shares the first training data with the target computing node, the target computing node needs to provide the first computing node with digital resources corresponding to the agreed-upon number of digital resources. It can be understood that the agreed-upon number of digital resources can be less than or equal to the number of competing digital resources initially corresponding to the target computing node. The specific implementation can be determined according to different situations. For example, the agreed-upon number of digital resources can be equal to the second largest number of competing digital resources received by the first computing node, which can better determine the data sharing relationship with the target computing node.
[0173] Step S103: If a first data sharing request is received from the target computing node, the first training data is sent to the target computing node so that the target computing node trains the initial model associated with the federated learning task based on the first training data and the stored training data associated with the federated learning task to obtain the first branch training update parameters; the task initiating node is also used to globally update the initial model through N branch training update parameters; the N branch training update parameters include the first branch training update parameters, where N is a positive integer.
[0174] Specifically, after identifying the target computing node, the first computing node establishes a data-sharing relationship with it. However, at this point, the first computing node does not yet send the first training data to the target computing node. Instead, it waits for the target computing node to send a first data-sharing request before sending the first training data. This is because the task initiating node has limited budgeted digital resources for a federated learning task, and there are usually multiple computable nodes in the blockchain network that possess the training data associated with the federated learning task and are capable of completing it. The target computing node is one of these computable nodes. Therefore, the task initiating node typically needs to allocate tasks, acquiring as much training data as possible to complete the federated learning task without exceeding the budgeted digital resources. If the target computing node is among the computable nodes selected by the task initiating node, then the target computing node will send a first data-sharing request to the first computing node. After acquiring the first training data, the target computing node can train the initial model associated with the federated learning task based on the first training data and the stored training data associated with the federated learning task, obtaining the first branch training update parameters, and then returning them to the task initiating node. Similarly, other computable nodes in the computable nodes will also return branch training update parameters to the task initiating node. In other words, the task initiating node can eventually receive N branch training update parameters, and the task initiating node will perform a global update on the initial model based on these N branch training update parameters.
[0175] Optionally, when the first computing node determines the target computing node, it acquires the agreed-upon amount of digital resources with the target computing node. Then, it packages the agreed-upon amount of digital resources, the amount of the first training data, and the data sharing relationship with the target computing node into a data sharing record transaction, which is then cached in the record transaction pool. Upon receiving the first data sharing request from the target computing node, the first computing node can perform consensus and upload the data sharing record transaction to the blockchain; then, it receives the digital resources corresponding to the agreed-upon amount of digital resources from the target computing node. In other words, the data sharing between the first computing node and the target computing node is stored in the blockchain through a data sharing record transaction for easy verification later. Please also refer to... Figure 4 , Figure 4 This is a schematic diagram of a data structure for recording transactions provided in an embodiment of this application. For example... Figure 4As shown, a transaction record includes a transaction identifier, transaction type, message recipient, message sender, payload data, and digital signature. The transaction identifier uniquely identifies a transaction and is determined by the transaction generation rules. The transaction type includes data sharing record transactions, specifically those generated for the first computing node. In this case, the message recipient is the target computing node, and the message sender is the first computing node. The payload data includes the agreed-upon number of digital resources, the amount of the first training data, and the data sharing relationship between the first and target computing nodes. The digital signature is generated by the first computing node using its private key.
[0176] Using the method provided in this application embodiment, if the first computing node is computationally limited, it can share its stored first training data with a target computing node that has federated learning computing capabilities for the first training data and meets the node trustworthiness conditions. The target computing node can then use both its own stored training data and the first training data to train the model, avoiding the waste of training data in computationally limited computing nodes, improving the data utilization rate in the federated learning process, and allowing more training data to participate in the training, resulting in a higher performance model. Moreover, since both the first computing node and the target computing node are nodes in the blockchain network and meet the trustworthiness conditions between them, the data interaction will be more secure.
[0177] Further, please see Figure 5 , Figure 5 This is a flowchart illustrating a blockchain-based federated learning processing method provided in this application embodiment. The method can be executed by a first computing node in a blockchain network (which can be any blockchain node in the aforementioned blockchain network). The blockchain network also includes a task initiating node (which can be any blockchain node other than the first computing node) and M second computing nodes (which can be any M blockchain nodes other than the first computing node and the task initiating node), where M is a positive integer. The following description uses the execution of this method by the first computing node as an example. This blockchain-based federated learning processing method can include at least the following steps S201-S206:
[0178] Step S201: The first computing node receives the federated learning task generated by the task initiating node through the task smart contract, and obtains the training data associated with the federated learning task as the first training data.
[0179] Specifically, the implementation of step S201 can be found in step S101 above, and will not be repeated here.
[0180] Step S202: Receive the second training competition requests for the federated learning task sent by the P third computing nodes respectively.
[0181] Specifically, among the P third computing nodes, the second training competition request sent by the j-th third computing node includes the amount of data for the j-th second training data; the j-th second training data is the training data associated with the federated learning task that has been stored in the j-th third computing node; the j-th third computing node does not have the federated learning computing capability for the j-th second training data; j is a positive integer less than or equal to P.
[0182] Specifically, the federated learning task initiated by the task-initiating node is broadcast in the federated learning market within the blockchain network. Therefore, multiple computing nodes will receive the task. Among these nodes, there may be several computationally limited nodes that possess training data associated with the task but lack the computational capability to complete it. These computationally limited nodes will send training competition requests to computing nodes within their communication range. For the first computing node, the third computing node refers to a computationally limited node within its communication range that possesses the second training data but lacks the computational capability for federated learning based on that second training data.
[0183] Step S203: If the first computing node has the federated learning computing capability for the first training data, then the third computing node that successfully competed among the P third computing nodes is determined as the target acquisition node; the first computing node also has the federated learning capability for the target training data; the target training data is the second training data corresponding to the target acquisition node.
[0184] Specifically, before responding to P second training competition requests, the first computing node needs to determine whether it possesses the federated learning computational capability for the first training data. If the first computing node cannot even train its own first training data, it is considered a computationally constrained node and therefore cannot compete for the second training data of other computationally constrained nodes, i.e., the third computing node. The process for the first computing node to determine its federated learning computational capability for the first training data can be found above. Figure 3 The optional implementation process of step S101 in the corresponding embodiment will not be described in detail here.
[0185] Specifically, if the first computing node possesses the federated learning computing capability for the first training data, then a feasible implementation process for determining the third computing node that successfully competes among the P third computing nodes as the target acquisition node can be as follows: If the first computing node possesses the federated learning computing capability for the first training data, then iterate through the second training competition requests for the federated learning task sent by the P third computing nodes respectively; obtain the second training data G from the second training competition request sent by the kth third computing node. k The amount of data, based on the second training data G k The amount of data is determined for the second training data G. k The second training duration T of the data volume k ; k is a positive integer less than or equal to P; the second training duration T k Add the total training time T to the first training session duration. k总 If the total training time T k总 If the timeout is less than or equal to the task deadline, then the first computing node is determined to have the capability for the second training data G. k The federated learning computing capability sends the second training competition response information R to the k-th third computing node. kDuring the competition waiting period, the system receives second training competition confirmation requests from Q third computing nodes, where Q is a positive integer less than or equal to P. Each second training competition confirmation request contains a target number of competing digital resources. The system determines the third computing node corresponding to the second training competition confirmation request with the lowest unit data resource consumption as the successful third computing node, sends a second training competition confirmation response to the successful third computing node, and designates the successful third computing node as the target acquisition node. The unit data resource consumption of a third computing node is determined based on the target number of competing digital resources corresponding to the third computing node and the amount of second training data. In short, the first computing node determines the third computing nodes that can complete the training of its stored second training data, and then responds to the second training competition requests of these third computing nodes by sending second training competition response information to them respectively. As can be seen from step S102 above, a third computing node typically receives other training competition response information besides the second training competition response information. Therefore, a third computing node that receives the second training competition response information does not necessarily send a second training competition confirmation request to the first computing node. It will only send a second training competition confirmation request to the first computing node if it has selected the first computing node as a pre-selected computing node. Therefore, the first computing node can ultimately receive second training competition confirmation requests from Q third computing nodes, where Q is a positive integer less than or equal to P. The first computing node will naturally select the third computing node corresponding to the second training competition confirmation request with the lowest unit data resource consumption as the successful third computing node. Here, unit data resource consumption refers to the amount of digital resources required to obtain the second training data corresponding to a unit amount of data. It can be understood that if the first computing node has abundant computing resources, it can also select multiple third computing nodes from the Q third computing nodes as successful third computing nodes, as long as it can complete the training of the second training data corresponding to all the selected successful third computing nodes and its own stored first training data within the task deadline.
[0186] Step S204: Send a federated learning task response request carrying training data information to the task initiating node; the training data information is generated based on the first training data and the target training data.
[0187] Specifically, the training data information includes the total amount of training data and the unit data resource consumption. The total amount of training data is the sum of the data volume corresponding to the first training data and the data volume corresponding to the target training data, which can be calculated using the following formula (3):
[0188]
[0189] Among them, g i,j This indicates that for the task initiating node r i First computing node e j The amount of data possessed, γ i,k Indicates the first computing node e j From the third computing node e that won the competition k The amount of data obtained is μ, which represents the number of third computing nodes that successfully competed for the data.
[0190] The unit data resource consumption can be calculated using the following formula (4):
[0191]
[0192] in, For the first computing node e j The corresponding total amount of training data; v i,j Indicates the first computing node e j Initial unit data resource consumption, c k,j Indicates the first computing node e j The third computing node e that wins the competition needs to be selected. k The amount of data resources consumed per unit of payment. Indicates the first computing node e j The final corresponding unit data resource consumption.
[0193] Step S205: If a federated learning task issuance instruction is received from the task initiating node, a second data sharing request is sent to the target acquisition node; the federated learning task issuance instruction includes the initial model associated with the federated learning task.
[0194] Specifically, in addition to the first computing node, the task initiating node may also receive federated learning task response requests from other computable nodes. The task initiating node will select some computable nodes from the computable nodes corresponding to the received federated learning task response requests and send federated learning task issuance instructions.
[0195] Step S206: Receive the target training data sent by the target acquisition node according to the second data sharing request; train the initial model according to the first training data and the target training data to obtain the second branch training update parameters; the task initiating node is further configured to perform a global update of the initial model using H branch training update parameters; the H branch training update parameters include the second branch training update parameters, where H is a positive integer.
[0196] Optionally, if random differential privacy noise is received simultaneously with the federated learning task issuance instruction sent by the task initiating node, the training update parameters of the second branch are encrypted based on the received random differential privacy noise to obtain the target encrypted branch training update parameters, which are then sent to the task initiating node. The task initiating node is also used to aggregate the information of the S encrypted branch training update parameters to obtain encrypted aggregated training update parameters; S is a positive integer. The task initiating node is further used to add the encrypted aggregated training update parameters to the differential privacy key to obtain aggregated training update parameters, and to globally update the initial model based on the aggregated training update parameters. The S encrypted branch training update parameters include the target encrypted branch training update parameters. The encryption formula used for encrypting the second branch training update parameters is as follows:
[0197]
[0198] Among them, W i This represents the initial parameters corresponding to the initial model. This represents the parameters updated by the second branch after training, i.e., the parameters corresponding to the initial model after training, δ. j This represents the random differential privacy noise received by the second computing node.
[0199] Using the method provided in this application embodiment, if the first computing node has sufficient computing resources, it can compete for the second training data of the third computing node. In this way, when completing the federated learning task corresponding to the task initiating node, more training data can be used, the trained model performance is higher, and the waste of training data in computing nodes with limited computing resources is avoided.
[0200] Further, please see Figure 6 , Figure 6 This is a flowchart illustrating a blockchain-based federated learning processing method provided in this application embodiment. The method can be executed by a task initiating node (which can be any blockchain node in the aforementioned blockchain network), which also includes W computing nodes (which can be any W blockchain nodes other than the task initiating node), where W is a positive integer. The following description uses the execution of this method by a task initiating node as an example. This blockchain-based federated learning processing method can include at least the following steps S301-S304:
[0201] Step S301: The task initiating node generates a federated learning task through the task smart contract and broadcasts the federated learning task to the W computing nodes.
[0202] Specifically, if there are computationally restricted nodes among the W computing nodes, the computationally restricted nodes broadcast training competition requests for shared training data to the corresponding communicable computing nodes. The computationally restricted nodes also determine, among the communicable computing nodes, the target communicable computing node that successfully competes and meets the node trustworthiness condition. The shared training data is the training data associated with the federated learning task already stored in the computationally restricted nodes. The computationally restricted nodes do not possess the federated learning computing capabilities for the shared training data. The specific implementation process for the computationally restricted nodes to determine the target communicable computing node can be found above. Figure 3 The process of the first computing node determining the target computing node in step S102 of the corresponding embodiment will not be described again here.
[0203] Step S302: Receive federated learning task response requests sent by Y capability computing nodes respectively; Y is a positive integer; the Y capability computing nodes include the target communicable computing node, and the Y capability computing nodes do not include the computing-restricted node.
[0204] Specifically, a federated learning task response request includes a unit data resource consumption and a total training data volume. The process by which each capability computing node determines its corresponding unit data resource consumption and total training data volume can be found above. Figure 5 The description of the first computing node determining the total amount of training data and the unit data resource consumption in step S204 of the corresponding embodiment.
[0205] Step S303: Based on the Y federated learning task response requests, determine X training computing nodes from the Y capability computing nodes, and send federated learning task issuance instructions carrying the initial model associated with the federated learning task to the X training computing nodes, so that the X training computing nodes train the initial model according to the available training data to obtain branch training update parameters; X is a positive integer; if the X training computing nodes include the target communicable computing node, then the available training data corresponding to the target communicable computing node includes the shared training data and the training data already stored in the target communicable computing node.
[0206] Specifically, a feasible implementation process for determining X training computing nodes from Y capability computing nodes based on Y federated learning task response requests can be as follows: Obtain the predicted digital resource consumption corresponding to the federated learning tasks; sort the Y unit data resource consumptions in ascending order according to the Y federated learning task response requests to obtain the sorted Y unit data resource consumptions; iterate through the sorted Y unit data resource consumptions, sequentially obtaining the e-th unit data resource consumption; and determine the predicted digital resource consumption of the e-th node based on the e-th unit data resource consumption and the total amount of training data corresponding to the e-th unit data resource consumption; where e is less than or equal to 0. A positive integer equal to Y; add the capability calculation node corresponding to the e-th unit of data resource consumption to the node pre-selection queue; if the predicted digital resource consumption of the e-th node is less than the current predicted remaining digital resource consumption, subtract the current predicted remaining digital resource consumption from the predicted digital resource consumption of the e-th node to obtain the updated predicted remaining digital resource consumption, and continue traversing to obtain the (e+1)-th unit of data resource consumption; if the predicted digital resource consumption of the e-th node is greater than or equal to the current predicted remaining digital resource consumption, or e equals Y, stop traversing and use all X capability calculation nodes in the node pre-selection queue as training calculation nodes. For ease of understanding, assume the predicted digital resource consumption of the task initiating node is 100, the total training data for capability calculation node A is 20 with a unit data resource consumption of 2, the total training data for capability calculation node B is 40 with a unit data resource consumption of 1, and the total training data for capability calculation node C is 50 with a unit data resource consumption of 1.5. The three unit data resource consumptions after ascending sorting are 1, 1.5, and 2. The task initiating node will first select capability calculation node B, determining its predicted digital resource consumption to be 1 * 40 = 40, and then add it to the node pre-selection queue. Since 40 is less than 100, the task initiating node subtracts 40 from 100 to obtain a new remaining predicted digital resource consumption of 60. Then, the task initiating node will continue to select capability calculation node C, determining its predicted digital resource consumption to be 50 * 1.5 = 75, and then add capability calculation node C to the node pre-selection queue. Since 75 is greater than 60, the task initiating node selection process ends. At this point, capability computing nodes B and C in the node pre-selection queue are the training computing nodes. It should be noted that for capability computing node B, the task initiating node can allocate digital resources corresponding to its predicted digital resource consumption, as the task initiating node has sufficient budgeted digital resources. However, for capability computing node C, the task initiating node can only allocate digital resources corresponding to the remaining predicted digital resource consumption of 60. Therefore, when capability computing node C completes the federated learning task issuance instruction issued by the task initiating node, it will only acquire training data corresponding to the remaining predicted digital resource consumption of 60 for training.
[0207] Optionally, X-1 training computing nodes are sequentially selected from the node pre-selection queue; the predicted digital resource consumption of each of the X-1 training computing nodes is used as the transaction digital resource consumption of each of the X-1 training computing nodes; if the predicted digital resource consumption of the Xth training computing node is less than the target remaining predicted digital resource consumption, then the predicted digital resource consumption of the Xth training computing node is used as the transaction digital resource consumption of the Xth training computing node; the target remaining predicted digital resource consumption refers to the predicted digital resource consumption minus the predicted digital resource consumption of each of the X-1 training computing nodes. The remaining value after the source predicted consumption; if the node digital resource predicted consumption corresponding to the Xth training computing node is greater than or equal to the target remaining digital resource predicted consumption, then the target remaining digital resource predicted consumption is used as the transaction digital resource consumption corresponding to the Xth training computing node; each training computing node and its corresponding transaction digital resource consumption are associated and packaged into a digital resource record transaction, and the digital resource record transaction is cached in the record transaction pool; when the target model that meets the training conditions indicated by the federated learning task is obtained, the digital resources corresponding to the associated transaction digital resource consumption are sent to each training computing node, and consensus is achieved on the blockchain for the digital resource record transaction. Please see also Figure 4 , Figure 4 The transaction types shown also include digital resource record transactions. The task initiating node will generate a digital resource record transaction for each training computing node. In this case, in the data structure of the digital resource record transaction, the message receiver is the training computing node, the message sender is the task initiating node, the payload data is the training computing node and the digital resource consumption of the corresponding training computing node, and the digital signature is generated by the task initiating node using its private key.
[0208] Step S304: Globally update the initial model according to the aggregated training update parameters until a target model that meets the training conditions indicated by the federated learning task is obtained; the aggregated training update parameters are generated based on the X branch training update parameters.
[0209] Specifically, as mentioned above, after the task initiating node receives X branch training update parameters, it performs information aggregation processing on the X branch training update parameters to obtain the aggregated training update parameters.
[0210] Specifically, the training conditions indicated by the federated learning task usually include reaching the target number of iterations or the target accuracy of the parameters. Therefore, if the parameter accuracy of the aggregated training to update the parameters does not reach the target accuracy, the task initiating node can distribute the globally updated initial model to X training computing nodes, so that the X training computing nodes can conduct a new round of training on the globally updated initial model based on the available training data.
[0211] Optionally, for data transmission security, the task initiating node can generate a set of random differential privacy noise. The set of differential privacy noise contains random differential privacy noise corresponding to each training computing node. The X random differential privacy noises are added together to obtain the total random differential privacy noise. The negative of the total random differential privacy noise is used as the differential privacy key. Then, while sending the federated learning task issuance instruction carrying the initial model associated with the federated learning task to the X training computing nodes, the corresponding random differential privacy noise is sent to each of the X training computing nodes, so that the X training computing nodes encrypt the branch training update parameters according to the received random differential privacy noise to obtain encrypted branch training update parameters.
[0212] Optionally, the task initiating node determines the aggregated training update parameters at this point. This can be achieved by aggregating the encrypted branch training update parameters returned by X training computation nodes to obtain encrypted aggregated training update parameters, and then adding these encrypted aggregated training update parameters to a differential privacy cipher to obtain the aggregated training update parameters. A feasible implementation of aggregating the encrypted branch training update parameters returned by X training computation nodes to obtain the encrypted aggregated training update parameters can be as follows: obtain a training test set; verify the encrypted branch training update parameters returned by the X training computation nodes based on the training test set to obtain X update accuracies; use the encrypted branch training update parameters returned by the training computation nodes whose update accuracies meet the training accuracy conditions as the target encrypted branch training update parameters; and perform information aggregation processing on the target encrypted branch training update parameters to obtain the encrypted aggregated training update parameters.
[0213] Optionally, X training computing nodes and X update precisions are packaged into an update precision record transaction, and a consensus process for the update precision record transaction is initiated to the blockchain network. This ensures that each node in the blockchain network updates its stored precision record table based on the X update precisions when the consensus for the update precision record transaction is passed. The precision record table records each node in the blockchain network and its corresponding current update precision. The data structure of the update precision record transaction can also be found above. Figure 4 The data structure for recording transactions in the corresponding embodiments will not be described in detail here.
[0214] Optionally, the consistency of the distributed ledger can be guaranteed based on a contribution-driven consensus mechanism. When the task initiating node, acting as the execution node for a round, determines which node should produce a new block, it can obtain a precision record table based on the vote confirmation request; determine the contribution of each node based on the current update precision of each node in the precision record table; determine the resource rights of each node based on the contribution weight, the contribution of each node, and the duration of digital resource ownership for each node; send votes to each node according to the ratio between the resource rights of each node, so that each node can vote on the received votes; if there are computationally restricted nodes among the nodes, the computationally restricted nodes are used to vote for the target communicable computing node; after receiving the votes, the number of votes of each node determines the block production right acquisition ratio of each node; randomly select a block production node from each node according to the block production right acquisition ratio; the block production node has the right to produce the new block; send reward digital resources to the block production node, so that the block production node can allocate the reward digital resources according to the vote composition ratio; the vote composition ratio refers to the ratio between the number of votes received by the block production node from the task initiating node and the number of votes cast by the target computationally restricted node.
[0215] The contribution of each node, determined based on the current update precision of each node in the precision record table, can be calculated using the following formula:
[0216] ψ(ε)=-εlog2(ε)-(1-ε)log2(1-ε) Formula (6)
[0217] Where ε is the current update precision of the node. When ε is between 0 and 0.5, the contribution of the node in this update is 1-ψ(ε). When ε is between 0.5 and 1, the contribution of the node in this update is ψ(ε)-1.
[0218] The resource rights of each node are determined based on its contribution weight, the contribution level of each node, and the duration of digital resource ownership for each node. The formula for calculating the rights is as follows:
[0219]
[0220] Where coinage represents the duration of ownership of digital resources corresponding to a node, which is the quantity of digital resources held by the device multiplied by the holding time. i,j Represents node e j In the task initiator r i The contribution gained in the federal learning task.
[0221] The method provided in this application uses blockchain to record relevant transactions and smart contracts to regulate market order. An improved cancelable differential privacy noise encryption algorithm is used to eliminate the impact of noise on model accuracy, thereby effectively preventing attacks from malicious devices and ensuring that all federated learning tasks can be safely trained.
[0222] Further, please see Figure 7 , Figure 7 This is a flowchart illustrating a blockchain-based federated learning processing method provided in this application embodiment. The method can be jointly implemented by IoT devices in a federated learning marketplace (BFL marketplace) within a blockchain network. Any IoT device in the BFL marketplace can act as a task initiating node, while other IoT devices can be considered as computing nodes. The following description will use the example of this method being jointly executed by a task initiating node and computing nodes. This blockchain-based federated learning processing method can include at least the following steps S401-S404:
[0223] Step S401: The task initiating node generates a federated learning task and broadcasts the federated learning task in the BFL market.
[0224] Specifically, the task initiating node requests to initiate federated learning by calling a smart contract to create a federated learning task. The task initiating node can then broadcast the completed federated learning task in the BFL market, specifying the type of training data required, the task deadline, and the global model size.
[0225] In step S402, the computing node receives the federated learning task sent by the task initiating node and determines whether it has the training data required for the federated learning task and whether it has the ability to complete the federated learning task. The computing-limited node that has the training data but does not have the ability to complete the federated learning task can initiate a data competition within its communication range and share the data with trusted devices in the vicinity. The computing nodes that have the ability to complete the federated learning task participate in the competition for data resources. The computing nodes that win the competition will obtain the training data of the computing-limited node when participating in the federated learning task.
[0226] Specifically, it can be understood that both computationally restricted nodes and computable nodes belong to the category of computing nodes.
[0227] Specifically, data contention initiated by computationally constrained nodes can be accomplished using a trust-enhanced collaborative learning strategy (TCL) based on data sharing. This TCL algorithm is described above. Figure 3 This is a specific embodiment of step S102 in the corresponding embodiment. For ease of understanding, please also refer to... Figure 8, Figure 8 This is a flowchart of a data competition method provided in an embodiment of this application. Figure 8 As shown, data competition includes the following steps:
[0228] Step S501: Calculate the data contention initiated by the restricted node.
[0229] For details, please refer to the above. Figure 3 The implementation process described in the corresponding embodiment is as follows: the first computing node, which does not have the ability to perform federated learning computation on the first training data, broadcasts the first training competition request to M second computing nodes. The computing-limited node can broadcast the data competition request it initiates within its communication range, including the amount of training data it owns and the task deadline.
[0230] Step S502: Calculate the number of nodes participating in the competition and send the number of digital resources to be competed for.
[0231] Specifically, after receiving a data competition request, the computable node determines whether it can complete the training of the competing training data within the specified time based on the amount of training data to be competed for and the task deadline. If it can complete the training, the computable node bids for the computationally limited node, that is, it sends a training competition response message carrying the number of competing digital resources.
[0232] Step S503: Calculate the set of competing digital resources obtained by the restricted nodes, and set n=1.
[0233] Specifically, after broadcasting a data competition request, computationally restricted nodes can receive bids within a fixed time period. After the fixed time period ends, they no longer accept new bids. The computationally restricted node needs to remove the number of competing digital resources from the received bidders whose credibility is below the credibility threshold. Then, from the remaining number of competing digital resources, it selects the computable node with the highest number of competing digital resources as the potential winner of this data competition (i.e., the pre-selected node). Based on the second-highest number of competing digital resources, it calculates the agreed-upon number of digital resources for this data competition for the potential winner and sends a confirmation request to that node. First, the computationally restricted node puts the number of competing digital resources received within the fixed time period into a set of competing digital resources, and then sets n=1.
[0234] Step S504: Calculate the number of competitive digital resources that the restricted node takes from the set of competitive digital resource quantities.
[0235] Step S505: Calculate the confidence level of the restricted node to determine whether it is greater than the confidence level threshold. If the confidence level of the nth computable node is greater than the confidence level threshold, then proceed to step S506; otherwise, increment n and proceed to step S504.
[0236] Specifically, the reliability of determining the nth computable node by calculating the restricted nodes can be referenced above. Figure 3 In the corresponding embodiment, in step S102, the first computing node determines the competing node Z. i The process of determining credibility will not be elaborated here. Furthermore, n++ means incrementing n by one.
[0237] Step S506: Calculate whether the number of the nth competing digital resources in the restricted node is greater than the current highest competing digital resource number. If the number of the nth competing digital resources is greater than the current highest competing digital resource number, then proceed to step S507; otherwise, proceed to step S508.
[0238] Step S507: Calculate the highest and second-highest number of digital resources updated by the restricted node.
[0239] Specifically, the computationally constrained node will first update the second-highest number of digital resources to the current highest number of competing digital resources, and then update the highest number of digital resources to the nth number of competing digital resources.
[0240] Step S508: Calculate whether the restricted node determines whether n is equal to the total number of competing digital resources in the set of competing digital resources. If it is equal, proceed to step S509; otherwise, increment n and proceed to step S504.
[0241] Specifically, n equals the total number of competing digital resources in the set of competing digital resources, indicating that the computationally constrained node has completed its traversal of the set of competing digital resources.
[0242] Step S509: The computationally restricted node sends an acknowledgment request to the computable node corresponding to the highest number of competing digital resources.
[0243] Specifically, the confirmation request is actually the above. Figure 3 The first training competition confirmation request described in the corresponding embodiment carries a predetermined number of digital resources, which is usually generated based on the second-highest number of digital resources.
[0244] Step S510: Calculate whether the restricted node has received an acknowledgment response from the computable node corresponding to the highest number of competing digital resources. If no acknowledgment is received, proceed to step S501 or end the process; if acknowledgment is received, proceed to step S511.
[0245] Specifically, the confirmation response is actually as described above. Figure 3 The first training competition confirmation response information described in the corresponding embodiment.
[0246] Step S511: Compute the restricted node to store the data competition transaction record into the blockchain (not yet effective).
[0247] Specifically, a feasible process for computationally constrained nodes to store data competition transaction records in the blockchain can be found in the above. Figure 3 The corresponding embodiment describes the generation of the data sharing record transaction in step S103. At this time, the data sharing record transaction can be cached in the record transaction pool and has not been put on the blockchain for consensus.
[0248] In step S403, the task initiating node assigns the federated learning task to the appropriate training computing node.
[0249] Specifically, training compute nodes are computeable nodes selected by the task initiating node. When assigning nodes, the task initiating node must ensure that it can obtain as much data as possible within a given budget, and the training compute nodes must have the training data for the assigned federated learning task.
[0250] Specifically, when allocating federated learning tasks, the task initiating node can employ a quality-oriented task allocation algorithm (QTA) that combines a greedy strategy. This QTA algorithm is described above. Figure 6 This is a specific embodiment of step S303 in the corresponding embodiment. For ease of understanding, please also refer to... Figure 9 , Figure 9 This is a flowchart of a task allocation method provided in an embodiment of this application. For example... Figure 9 As shown, the task allocation method includes the following steps:
[0251] Step S601 can calculate the amount of data held by the node and the amount of data resources consumed per unit.
[0252] Specifically, the process by which a compute node determines its data holdings and unit data resource consumption can be referenced above. Figure 5 The description of the first computing node determining the total amount of training data and the unit data resource consumption in step S204 of the corresponding embodiment.
[0253] Step S602: The task initiating node obtains the set of unit data resource consumption, sorts the set of unit data resource consumption in ascending order, and sets m=1.
[0254] Specifically, in order to save on budget, the task initiation node should start by selecting the computable node with the lowest unit data resource consumption. Therefore, it is necessary to sort the unit data resource consumption from low to high.
[0255] Step S603: The task initiating node takes the m-th computable node and adds it to the node pre-selection queue.
[0256] Step S604: The task initiating node determines the node digital resource consumption corresponding to the m-th computable node, and uses the node digital resource consumption corresponding to it as the transaction digital resource consumption.
[0257] Specifically, node digital resource consumption = data holdings * unit data resource consumption. Transaction digital resource consumption refers to the amount of digital resources that the task-initiating node will ultimately send to that computable node.
[0258] Step S605: The task initiating node updates the predicted consumption of remaining digital resources.
[0259] Step S606: The task initiating node determines whether the predicted consumption of remaining digital resources is greater than the digital resource consumption of the node corresponding to the (m+1)th computable node; if yes, then m++, and execute step S603; if no, then execute step S607.
[0260] In step S607, the task initiating node uses the remaining digital resource prediction consumption as the transaction digital resource consumption corresponding to the (m+1)th computable node, and adds the (m+1)th computable node to the node pre-selection queue.
[0261] Step S608: The task initiating node stores the task allocation transaction record into the blockchain (already effective).
[0262] Specifically, the process by which the task initiating node stores the task allocation transaction record into the blockchain can be referred to the above. Figure 6 In the corresponding embodiment, step S303 is the process from the generation of digital resource record transactions to consensus on the blockchain. When the consensus on the digital resource record transaction is successfully uploaded to the blockchain, it means that the task allocation transaction record is stored in the blockchain and becomes effective.
[0263] Step S609: The data competition transaction records corresponding to the computable nodes in the node pre-selection queue take effect.
[0264] Step S610: The task initiating node determines the set of training computing nodes.
[0265] Specifically, the computable nodes in the node pre-selection queue are the training computation nodes that are ultimately determined by the task initiation node.
[0266] In step S404, after the federated learning task is assigned, the task initiating node sends the initial model parameters and model training hyperparameters to the training computing node. The training computing node completes the training of the initial model locally and sends the local update information to the task initiating node. After the task initiating node aggregates the local update information, it performs a global update on the initial model and iterates repeatedly until the accuracy requirements are met or the initially set number of iterations is reached.
[0267] Specifically, considering data security during transmission, a simple, noise-cancellable encrypted model training scheme (EMT) based on differential privacy encryption algorithms can be proposed to achieve secure data transmission. This EMT architecture is as described above. Figure 6 This is a specific embodiment of step S304 in the corresponding embodiment. For ease of understanding, please also refer to... Figure 10 , Figure 10 This is a flowchart illustrating the training of an encrypted model according to an embodiment of this application. Figure 10 As shown, the steps for training this encryption model include:
[0268] Step S701: The task initiating node sends the federated learning task distribution instruction and random differential privacy noise to the training computing node.
[0269] Specifically, the task initiating node will generate a random differential privacy noise for each training computing node in the training computing node set. The instructions issued for the federated learning task will include the model parameters and model training hyperparameters corresponding to the initial model.
[0270] Step S702: The training computing node downloads the initial model and trains it using available training data.
[0271] Specifically, the available training data package contains shared training data obtained from computationally limited nodes as well as training data stored locally.
[0272] Step S703: Train the computing nodes to obtain local update information and encrypt it using random differential privacy noise to obtain encrypted local update information.
[0273] The local update information is the branch training update parameter mentioned above, and the encrypted local update information is the encrypted branch training update parameter mentioned above. The encryption process can be found in the above formula (5).
[0274] Step S704: The task initiating node receives and verifies the encrypted local update information.
[0275] Specifically, after receiving all the encrypted local update information, the task initiating node will use a random test set to verify all the encrypted local update information in turn, and record the accuracy of each encrypted local update information on the random test set, i.e. the update precision, into the corresponding transaction as a measure of the local update quality of each training computing node.
[0276] In step S705, the task initiating node determines whether the local update accuracy meets the standard; if it does, proceed to step S706; if it does not, proceed to step S704.
[0277] Step S706: The task initiating node adds the qualified training computing nodes to the candidate aggregation queue.
[0278] In step S707, the task initiating node determines whether the task deadline has been reached. If yes, proceed to step S708; otherwise, proceed to step S704.
[0279] In step S708, the task initiating node aggregates local update information and completes the global update of the initial model.
[0280] In step S709, the task initiating node determines whether the initial model after global update meets the accuracy requirements or reaches the initially set number of iterations. If neither is met, step S701 is executed; if one or both are met, step S710 is executed.
[0281] In step S710, the task initiating node stores the transaction record of this federated learning task into the blockchain.
[0282] Step S405: After the federated learning task is completed, the task initiating node pays the corresponding reward to the training computing node that successfully participated in the local model training, and notifies the smart contract to revoke its federated learning task.
[0283] Specifically, after the federated learning task is completed, the task initiator pays reward digital resources to the training computing nodes that submitted local updates before the deadline based on the corresponding transaction records, while no payment is given to the training computing nodes that fail to complete the training on time.
[0284] The method provided in this application organizes the BFL market using blockchain technology. Each step employs smart contracts to enforce compliance among participants, ensuring fairness in data transactions and security in federated learning. Furthermore, the use of a collaborative trusted learning strategy based on data sharing (TCL) and a task allocation algorithm oriented towards training quality (QTA) significantly improves data utilization in the BFL market. This is because, with the TCL, computationally limited nodes, unable to undertake federated learning tasks, can send data to nearby trusted computationally available nodes, protecting data privacy and effectively preventing data waste. Moreover, due to market competition, computationally limited nodes often offer lower prices than the market average to maximize the value of their data. In QTA, task-initiating nodes prioritize low-bid computationally available nodes for training. Therefore, TCL and QTA enable task-initiating nodes to obtain more data within a fixed budget. Simultaneously, an EMT encryption architecture is used for model training, ensuring secure data transmission.
[0285] Please see Figure 11 , Figure 11This is a schematic diagram of the structure of a blockchain-based federated learning processing device provided in an embodiment of this application. The federated learning processing device can be a computer program (including program code) running on a computer device; for example, the federated learning processing device is an application software. This device can be used to execute corresponding steps in the federated learning processing method provided in the embodiments of this application. Figure 11 As shown, the federated learning processing device 1 may include: a task receiving module 101, a contention broadcasting module 102, a node determination module 103, and a data sharing module 104.
[0286] The task receiving module 101 is used to receive the federated learning task generated by the task initiating node through the task smart contract, and obtain the training data associated with the federated learning task as the first training data.
[0287] The competition broadcast module 102 is used to broadcast the first training competition request for the federated learning task to M second computing nodes if the first computing node does not have the computing capability for federated learning of the first training data.
[0288] The node determination module 103 is used to determine the second computing node that has successfully competed and meets the node credibility condition among the M second computing nodes, and to select the target computing node; the target computing node has the federated learning computing capability for the first training data.
[0289] The data sharing module 104 is used to send the first training data to the target computing node if it receives the first data sharing request sent by the target computing node, so that the target computing node can train the initial model associated with the federated learning task based on the first training data and the stored training data associated with the federated learning task to obtain the first branch training update parameters; the task initiating node is also used to globally update the initial model through N branch training update parameters; the N branch training update parameters include the first branch training update parameters, where N is a positive integer.
[0290] The specific implementation methods of the task receiving module 101, the contention broadcasting module 102, the node determination module 103, and the data sharing module 104 can be found in the above description. Figure 3 The specific descriptions of steps S101-S103 in the corresponding embodiments will not be repeated here.
[0291] The node determination module 103 includes: a competition information receiving unit 1031, a trusted node filtering unit 1032, a node pre-selection unit 1033, and a competition confirmation unit 1034.
[0292] The competition information receiving unit 1031 is used to receive first training competition response information sent by L competing nodes during the competition period; each first training competition response information contains a number of competing digital resources; the L competing nodes are nodes among M second computing nodes that have federated learning computing capabilities for the first training data; L is a positive integer less than or equal to M.
[0293] The trusted node filtering unit 1032 is used to remove competitive nodes that do not meet the node trustworthiness conditions from L competitive nodes, and obtain S trusted competitive nodes; S is a positive integer less than or equal to L.
[0294] The node pre-selection unit 1033 is used to select the trusted competing node with the highest number of competing digital resources from S trusted competing nodes as the pre-selected computing node;
[0295] The competition confirmation unit 1034 is used to send a first training competition confirmation request to the pre-selected computing node;
[0296] The competition confirmation unit 1034 is further configured to determine the pre-selected computing node as the target computing node if it receives the first training competition confirmation response information sent by the pre-selected computing node according to the first training competition confirmation request within the confirmation time period.
[0297] The specific implementation methods of the competition information receiving unit 1031, the trusted node filtering unit 1032, the node pre-selection unit 1033, and the competition confirmation unit 1034 can be found in the above description. Figure 3 The specific description of step S102 in the corresponding embodiment will not be repeated here.
[0298] Among the L competing nodes, competing node Z is included. i , where i is a positive integer less than or equal to L;
[0299] The aforementioned federated learning processing device 1 also includes: a trusted node determination module 105.
[0300] Trusted node determination module 105 is used to obtain a node trust probability table;
[0301] The trusted node determination module 105 is also used to query the competing node Z from the node trusted probability table. i The probability of credibility;
[0302] The trusted node determination module 105 is also used to determine the trusted node based on the competition node Z. i Connections between them and competing nodes Z i The reliable probability determines the competing node Z i Credibility;
[0303] The trusted node determination module 105 is also used if there is contention for node Z i If the credibility of a node is less than the credibility threshold, then the competing node Z is determined. i The node trustworthiness condition is not met;
[0304] The trusted node determination module 105 is also used if there is contention for node Z i If the credibility of a node is greater than or equal to the credibility threshold, then the competing node Z is determined. i The node trustworthiness condition is met.
[0305] The specific implementation of the trusted node determination module 105 can be found in the above description. Figure 3 The optional description of step S102 in the corresponding embodiment will not be repeated here.
[0306] The federated learning task includes the task deadline and initial model information;
[0307] The aforementioned federated learning processing device 1 also includes: a computing power determination module 106.
[0308] The computing power determination module 106 is used to determine the first required computing resources for the federated learning task based on the initial model information.
[0309] The computing power determination module 106 is further configured to determine that the first computing node does not have the computing power for federated learning of the first training data if the first required computing resources are greater than the available computing resources; available computing resources refer to the idle computing resources of the first computing node.
[0310] The computing power determination module 106 is further configured to determine the first training duration corresponding to the federated learning computation for the first training data based on the amount of data of the first training data if the first required computing resources are less than or equal to the available computing resources.
[0311] The computing power determination module 106 is also used to determine that the first computing node has federated learning computing power for the first training data if the first training duration is less than or equal to the task deadline.
[0312] The computing power determination module 106 is also used to determine that the first computing node does not have the computing power for federated learning of the first training data if the first training duration is longer than the task deadline.
[0313] For specific implementation details, please refer to the above. Figure 3 The optional description of step S101 in the corresponding embodiment will not be repeated here.
[0314] The blockchain network also includes P third-party computing nodes, where P is a positive integer;
[0315] The aforementioned federated learning processing device 1 further includes: a contention request receiving module 107, a node acquisition determination module 108, a task response module 109, a data acquisition module 110, and a model training module 111.
[0316] The competition request receiving module 107 is used to receive second training competition requests for the federated learning task sent by P third computing nodes respectively; the second training competition request sent by the j-th third computing node includes the data volume of the j-th second training data; the j-th second training data is the training data associated with the federated learning task that has been stored in the j-th third computing node; the j-th third computing node does not have the federated learning computing capability for the j-th second training data; j is a positive integer less than or equal to P;
[0317] The node determination module 108 is used to determine the third computing node that has successfully competed among P third computing nodes as the target acquisition node if the first computing node has federated learning computing capabilities for the first training data; the first computing node also has federated learning capabilities for the target training data; the target training data is the second training data corresponding to the target acquisition node.
[0318] Task response module 109 is used to send a federated learning task response request carrying training data information to the task initiating node; the training data information is generated based on the first training data and the target training data;
[0319] The data acquisition module 110 is used to send a second data sharing request to the target acquisition node if it receives a federated learning task issuance instruction sent by the task initiating node; the federated learning task issuance instruction contains the initial model associated with the federated learning task.
[0320] The model training module 111 is used to receive target training data sent by the target acquisition node according to the second data sharing request, train the initial model according to the first training data and the target training data, and obtain the second branch training update parameters; the task initiation node is also used to perform global updates on the initial model through H branch training update parameters; the H branch training update parameters include the second branch training update parameters, where H is a positive integer.
[0321] The specific implementation methods of the contention request receiving module 107, the node acquisition and determination module 108, the task response module 109, the data acquisition module 110, and the model training module 111 can be found above. Figure 5 The specific descriptions of steps S201-S206 in the corresponding embodiments will not be repeated here.
[0322] The node determination module 108 includes a node selection unit 1081 and a node confirmation unit 1082.
[0323] The node selection unit 1081 is used to traverse the second training competition requests for the federated learning task sent by P third computing nodes respectively if the first computing node has the federated learning computing capability for the first training data.
[0324] The node selection unit 1081 is also used to obtain the second training data G from the second training competition request sent by the kth third computing node. k The amount of data, based on the second training data G k The amount of data is determined for the second training data G. k The second training duration T of the data volume k k is a positive integer less than or equal to P;
[0325] The node selection unit 1081 is also used to select the second training duration T. k Add the total training time T to the first training session duration. k总 ;
[0326] Node selection unit 1081 is also used if the total training time T k总 If the timeout is less than or equal to the task deadline, then the first computing node is determined to have the capability for the second training data G. k The federated learning computing capability sends the second training competition response information R to the k-th third computing node. k ;
[0327] The node confirmation unit 1082 is used to receive second training competition confirmation requests sent by Q third computing nodes respectively during the competition waiting period; the Q third computing nodes are nodes that received the second training competition response information sent by the first computing node, and Q is a positive integer; a second training competition confirmation request contains a target number of digital resources to be competed for;
[0328] The node confirmation unit 1082 is used to determine the third computing node corresponding to the second training competition confirmation request with the highest number of target competitive digital resources as the third computing node that has successfully competed, send the second training competition confirmation response information to the third computing node that has successfully competed, and regard the third computing node that has successfully competed as the target acquisition node.
[0329] For specific implementation details, please refer to the above. Figure 5 The specific description of step S203 in the corresponding embodiment will not be repeated here.
[0330] The aforementioned federated learning processing device 1 also includes an encryption module 112.
[0331] The encryption module 112 is used to encrypt the second branch training update parameters according to the received random differential privacy noise when it receives the federated learning task issuance instruction sent by the task initiating node, so as to obtain the target encrypted branch training update parameters, and send the target encrypted branch training update parameters to the task initiating node; the task initiating node is also used to perform information aggregation processing on S encrypted branch training update parameters to obtain encrypted aggregated training update parameters; S is a positive integer; the task initiating node is also used to add the encrypted aggregated training update parameters to the differential privacy key to obtain aggregated training update parameters, and perform global update of the initial model according to the aggregated training update parameters; the S encrypted branch training update parameters include the target encrypted branch training update parameters.
[0332] The specific implementation of the encryption module 112 can be found in the above description. Figure 5 The optional descriptions of the steps in the corresponding embodiments will not be repeated here.
[0333] The aforementioned federated learning processing device 1 also includes a data sharing record module 113.
[0334] The data sharing record module 113 is used to obtain the agreed-upon number of digital resources with the target computing node;
[0335] The data sharing record module 113 is also used to package the agreed number of digital resources, the amount of data of the first training data, and the data sharing relationship with the target computing node into a data sharing record transaction, and cache the data sharing record transaction in the record transaction pool;
[0336] The data sharing record module 113 is also used to perform consensus and on-chain recording of data sharing record transactions at the same time as receiving the first data sharing request sent by the target computing node;
[0337] The data sharing record module 113 is also used to receive digital resources sent by the target computing node that correspond to the agreed number of digital resources.
[0338] The specific implementation of the data sharing record module 113 can be found in the above description. Figure 3 The optional descriptions in the corresponding embodiments will not be repeated here.
[0339] Please see Figure 12 , Figure 12 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application. Figure 12 As shown above, Figure 11The federated learning processing device 1 in the corresponding embodiment can be applied to a computer device 1000, which may include a processor 1001, a network interface 1004, and a memory 1005. Furthermore, the computer device 1000 may also include a user interface 1003 and at least one communication bus 1002. The communication bus 1002 is used to implement communication between these components. The user interface 1003 may include a display screen and a keyboard; optionally, the user interface 1003 may also include a standard wired interface or a wireless interface. The network interface 1004 may optionally include a standard wired interface or a wireless interface (such as a Wi-Fi interface). The memory 1005 may be high-speed RAM or non-volatile memory, such as at least one disk storage device. Optionally, the memory 1005 may also be at least one storage device located remotely from the processor 1001. Figure 12 As shown, the memory 1005, which is a computer-readable storage medium, may include an operating system, a network communication module, a user interface module, and a device control application.
[0340] In such Figure 12 In the computer device 1000 shown, the network interface 1004 provides network communication elements; the user interface 1003 is mainly used to provide an input interface for users; and the processor 1001 can be used to call the device control application stored in the memory 1005. This computer device 1000 can serve as a first computing node to achieve:
[0341] The receiving task initiating node obtains the training data associated with the federated learning task generated by the task smart contract, and uses it as the first training data.
[0342] If the first computing node does not have the computing capability for federated learning of the first training data, the first training competition request for the federated learning task will be broadcast to M second computing nodes. Among the M second computing nodes, the second computing node that wins the competition and meets the node credibility condition will be selected as the target computing node. The target computing node has the computing capability for federated learning of the first training data.
[0343] If a first data sharing request is received from the target computing node, the first training data is sent to the target computing node so that the target computing node can train the initial model associated with the federated learning task based on the first training data and the stored training data associated with the federated learning task to obtain the first branch training update parameters; the task initiating node is also used to globally update the initial model through N branch training update parameters; the N branch training update parameters include the first branch training update parameters, where N is a positive integer.
[0344] It should be understood that the computer device 1000 described in the embodiments of this application can execute the foregoing text. Figure 3 , Figure 5 The description of the federated learning processing method in any corresponding embodiment will not be repeated here. Furthermore, the beneficial effects of using the same method will also not be repeated.
[0345] Furthermore, it should be noted that this application embodiment also provides a computer-readable storage medium, which stores a computer program executed by the aforementioned federated learning processing device 1. The computer program includes program instructions, and when the processor executes the program instructions, it can execute the aforementioned... Figure 3 , Figure 5 The description of the federated learning processing method in any corresponding embodiment is already provided, and therefore will not be repeated here. Furthermore, the beneficial effects of using the same method will also not be repeated. For technical details not disclosed in the computer-readable storage medium embodiments related to this application, please refer to the description of the method embodiments of this application.
[0346] Further, please see Figure 13 , Figure 13 This is a schematic diagram of another blockchain-based federated learning processing device provided in this application embodiment. The aforementioned federated learning processing device 2 can be a computer program (including program code) running on a computer device; for example, the federated learning processing device 2 is an application software. This device can be used to execute the corresponding steps in the method provided in this application embodiment. Figure 13 As shown, the federated learning processing device 2 may include: a task broadcasting module 21, a response receiving module 22, a training node determination module 23, a task distribution module 24, and a model update module 25.
[0347] The task broadcasting module 21 is used to generate federated learning tasks through task smart contracts and broadcast the federated learning tasks to W computing nodes. If there are computationally restricted nodes among the W computing nodes, the computationally restricted nodes are used to broadcast training competition requests for shared training data to the corresponding communicable computing nodes. The computationally restricted nodes are also used to determine the target communicable computing nodes that have successfully competed and meet the node trustworthiness conditions among the communicable computing nodes. The shared training data is the training data associated with the federated learning tasks that has been stored in the computationally restricted nodes. The computationally restricted nodes do not have the computational capability for federated learning of the shared training data.
[0348] The response receiving module 22 is used to receive federated learning task response requests sent by Y capability computing nodes respectively; Y is a positive integer; the Y capability computing nodes include target communicable computing nodes and do not include computing-restricted nodes.
[0349] The training node determination module 23 is used to determine X training computing nodes from Y capability computing nodes based on Y federated learning task response requests, where X is a positive integer.
[0350] The task distribution module 24 is used to send federated learning task distribution instructions carrying the initial model associated with the federated learning task to X training computing nodes, so that the X training computing nodes can train the initial model according to the available training data and obtain branch training update parameters; if the X training computing nodes include the target communicable computing node, the available training data corresponding to the target communicable computing node includes shared training data and training data already stored in the target communicable computing node.
[0351] The model update module 25 is used to globally update the initial model based on the aggregated training update parameters until a target model that meets the training conditions indicated by the federated learning task is obtained; the aggregated training update parameters are generated based on the X branch training update parameters.
[0352] The specific implementation methods of the task broadcasting module 21, response receiving module 22, training node determination module 23, task distribution module 24, and model update module 25 can be found above. Figure 6 The specific descriptions of steps S301-S304 in the corresponding embodiments will not be repeated here.
[0353] One federated learning task response request includes one unit of data resource consumption and one total amount of training data;
[0354] The training node determination module 23 includes: a consumption acquisition unit 231, an ascending sorting unit 232, and a training node selection unit 233.
[0355] Consumption acquisition unit 231 is used to acquire the predicted consumption of digital resources corresponding to the federated learning task.
[0356] The ascending sorting unit 232 is used to sort the Y units of data resource consumption in ascending order according to the Y federated learning task response requests, and obtain the sorted Y units of data resource consumption.
[0357] The training node selection unit 233 is used to traverse the sorted Y units of data resource consumption, sequentially obtain the e-th unit of data resource consumption, and determine the predicted digital resource consumption of the e-th node based on the e-th unit of data resource consumption and the total amount of training data corresponding to the e-th unit of data resource consumption; e is a positive integer less than or equal to Y.
[0358] The training node selection unit 233 is also used to add the capability calculation node corresponding to the e-th unit of data resource consumption to the node pre-selection queue.
[0359] The training node selection unit 233 is also used to subtract the current remaining digital resource prediction consumption from the e-th node's digital resource prediction consumption if the predicted consumption of digital resources at the e-th node is less than the current remaining digital resource prediction consumption, and then continue to traverse to obtain the e+1-th unit's data resource consumption.
[0360] The training node selection unit 233 is also used to stop traversing if the predicted consumption of digital resources of the e-th node is greater than or equal to the current predicted consumption of remaining digital resources, or if e equals Y, and to use all X capability calculation nodes in the node pre-selection queue as training calculation nodes.
[0361] The specific implementation methods of the consumption acquisition unit 231, the ascending sorting unit 232, and the training node selection unit 233 can be found above. Figure 6 The specific description of step S303 in the corresponding embodiment will not be repeated here.
[0362] The aforementioned federated learning processing device 2 also includes a digital resource recording module 26.
[0363] The digital resource recording module 26 is used to sequentially obtain X-1 training computing nodes from the node pre-selection queue;
[0364] The digital resource recording module 26 is also used to use the predicted consumption of node digital resources corresponding to X-1 training computing nodes as the transaction digital resource consumption corresponding to X-1 training computing nodes.
[0365] The digital resource recording module 26 is also used to obtain the Xth training computing node from the node pre-selection queue.
[0366] The digital resource recording module 26 is further configured to use the predicted consumption of the node digital resources corresponding to the Xth training computing node as the transaction digital resource consumption of the Xth training computing node if the predicted consumption of the node digital resources corresponding to the Xth training computing node is less than the target remaining predicted consumption of the node digital resources. The target remaining predicted consumption of the node digital resources refers to the value remaining after subtracting the predicted consumption of the node digital resources corresponding to the X-1 training computing nodes from the predicted consumption of the digital resources.
[0367] The digital resource recording module 26 is also used to take the target remaining digital resource prediction consumption as the transaction digital resource consumption of the Xth training computing node if the node digital resource prediction consumption corresponding to the Xth training computing node is greater than or equal to the target remaining digital resource prediction consumption.
[0368] The digital resource recording module 26 is also used to associate and package each training computing node and the corresponding transaction digital resource consumption of each training computing node into a digital resource recording transaction, and cache the digital resource recording transaction in the recording transaction pool;
[0369] The digital resource recording module 26 is also used to send the digital resources corresponding to the associated transaction digital resource consumption to each training computing node when the target model that meets the training conditions indicated by the federated learning task is obtained, and to perform consensus on the digital resource recording transaction on the blockchain.
[0370] The specific implementation of the digital resource recording module 26 can be found in the above description. Figure 6 The optional description of step S303 in the corresponding embodiment will not be repeated here.
[0371] The aforementioned federated learning processing device 2 further includes: a noise processing module 27, a noise transmission module 28, a parameter aggregation module 29, and a parameter decryption module 210.
[0372] Noise processing module 27 is used to generate a set of random differential privacy noise; the set of differential privacy noise contains random differential privacy noise corresponding to each training computing node;
[0373] The noise processing module 27 is also used to add up X random differential privacy noises to obtain the total random differential privacy noise;
[0374] The noise processing module 27 is also used to take the negative of the total random differential privacy noise as the differential privacy key;
[0375] The noise transmission module 28 is used to send the federated learning task issuance instruction carrying the initial model associated with the federated learning task to X training computing nodes, and send the corresponding random differential privacy noise to each of the X training computing nodes, so that the X training computing nodes can encrypt the branch training update parameters according to the received random differential privacy noise to obtain encrypted branch training update parameters.
[0376] The parameter aggregation module 29 is used to perform information aggregation processing on the encrypted branch training update parameters returned by X training computing nodes to obtain encrypted aggregated training update parameters.
[0377] The parameter decryption module 210 adds the encrypted aggregate training update parameters to the differential privacy cipher to obtain the aggregate training update parameters.
[0378] The specific implementation methods of the noise processing module 27, the noise transmission module 28, the parameter aggregation module 29, and the parameter decryption module 210 can be found in the above description. Figure 6 The optional description of step S304 in the corresponding embodiment will not be repeated here.
[0379] The parameter aggregation module 29 includes: an accuracy verification unit 291 and a target parameter aggregation unit 292.
[0380] Precision verification unit 291 is used to acquire training and test sets;
[0381] The accuracy verification unit 291 is also used to verify the encrypted branch training update parameters returned by X training computing nodes according to the training test set, and obtain X update accuracies.
[0382] The target parameter aggregation unit 292 is used to take the encrypted branch training update parameters returned by the training computation node whose update accuracy meets the training accuracy condition as the target encrypted branch training update parameters.
[0383] The target parameter aggregation unit 292 is also used to perform information aggregation processing on the target encrypted branch training update parameters to obtain encrypted aggregated training update parameters.
[0384] The specific implementation methods of the accuracy verification unit 291 and the target parameter aggregation unit 292 can be found in the above description. Figure 6 The optional description of step S304 in the corresponding embodiment will not be repeated here.
[0385] The aforementioned federated learning processing device 2 also includes: a precision recording module 211.
[0386] The precision recording module 211 is used to package X training computing nodes and X update precisions into an update precision recording transaction, and initiate consensus processing for the update precision recording transaction to the blockchain network, so that each node in the blockchain network updates the stored precision recording table according to the X update precisions when the consensus of the update precision recording transaction is passed; the precision recording table is used to record each node in the blockchain network and the current update precision corresponding to each node.
[0387] The specific implementation of the precision recording module 211 can be found in the above description. Figure 6 The optional description of step S304 in the corresponding embodiment will not be repeated here.
[0388] The aforementioned federated learning processing device 2 also includes:
[0389] The block production right determination module 212 is used to obtain the precision record table based on the vote confirmation request;
[0390] The block production right determination module 212 is also used to determine the contribution of each node based on the current update precision of each node in the precision record table;
[0391] The block production right determination module 212 is also used to determine the resource rights of each node based on the contribution weight, the contribution of each node, and the duration of digital resource ownership of each node.
[0392] The block production right determination module 212 is also used to send a ballot to each node according to the ratio between the resource rights of each node, so that each node can vote on the received ballot;
[0393] The block production right determination module 212 is also used to, if there are computationally restricted nodes in each node, have the computationally restricted nodes vote the received votes to the target communicable computing node.
[0394] The block production right determination module 212 is also used to receive the number of votes of each node after voting, and determine the block production right acquisition ratio of each node according to the number of votes of each node;
[0395] The block production right determination module 212 is also used to randomly determine the block production node from each node according to the block production right acquisition ratio; the block production node has the right to produce new blocks;
[0396] The block production right determination module 212 is also used to send reward digital resources to the block production node so that the block production node can allocate the reward digital resources according to the vote composition ratio; the vote composition ratio refers to the ratio between the number of votes sent by the task initiating node and the number of votes cast by the target computationally restricted node.
[0397] The specific implementation of the block production right determination module 212 can be found in the above description. Figure 6 The optional description of step S304 in the corresponding embodiment will not be repeated here.
[0398] Further, please see Figure 14 , Figure 14 This is a schematic diagram of the structure of another computer device provided in an embodiment of this application. For example... Figure 14 As shown above, Figure 13 The federated learning processing device 2 in the corresponding embodiment can be applied to a computer device 2000, which may include a processor 2001, a network interface 2004, and a memory 2005. Furthermore, the computer device 2000 also includes a user interface 2003 and at least one communication bus 2002. The communication bus 2002 is used to implement communication between these components. The user interface 2003 may include a display screen and a keyboard; optionally, the user interface 2003 may also include a standard wired interface or a wireless interface. The network interface 2004 may optionally include a standard wired interface or a wireless interface (such as a Wi-Fi interface). The memory 2005 may be high-speed RAM or non-volatile memory, such as at least one disk storage device. Optionally, the memory 2005 may also be at least one storage device located remotely from the aforementioned processor 2001. Figure 14 As shown, the memory 2005, which is a computer-readable storage medium, may include an operating system, a network communication module, a user interface module, and a device control application.
[0399] exist Figure 14 In the computer device 2000 shown, the network interface 2004 provides network communication functionality; the user interface 2003 is mainly used to provide an input interface for the user; and the processor 2001 can be used to call the device control application stored in the memory 2005. This computer device can act as a task initiation node to achieve:
[0400] A federated learning task is generated through a task smart contract and broadcast to W computing nodes. If there are computationally restricted nodes among the W computing nodes, the computationally restricted nodes are used to broadcast training competition requests for shared training data to the corresponding communicable computing nodes. The computationally restricted nodes are also used to determine the target communicable computing node that has successfully competed and meets the node trustworthiness condition among the communicable computing nodes. The shared training data is the training data associated with the federated learning task that has been stored in the computationally restricted nodes. The computationally restricted nodes do not have the computational capability for federated learning of the shared training data.
[0401] Receive federated learning task response requests from Y capability computing nodes, where Y is a positive integer; the Y capability computing nodes include target communicable computing nodes and do not include computing-restricted nodes.
[0402] Based on the Y federated learning task response requests, X training computing nodes are determined from the Y capability computing nodes. A federated learning task issuance instruction carrying the initial model associated with the federated learning task is sent to the X training computing nodes, so that the X training computing nodes train the initial model according to the available training data to obtain branch training update parameters; X is a positive integer; if the X training computing nodes include the target communicable computing node, the available training data corresponding to the target communicable computing node includes shared training data and training data already stored in the target communicable computing node;
[0403] The initial model is globally updated based on the aggregated training update parameters until a target model that meets the training conditions indicated by the federated learning task is obtained; the aggregated training update parameters are generated based on the training update parameters of X branches.
[0404] It should be understood that the computer device 2000 described in the embodiments of this application can execute the access control method described in the preceding embodiments, and can also execute the methods described in the preceding embodiments. Figure 13 The description of the federated learning processing device 2 in the corresponding embodiments will not be repeated here. Furthermore, the beneficial effects of using the same method will also not be repeated.
[0405] Furthermore, it should be noted that this application also provides a computer-readable storage medium storing the computer program executed by the aforementioned federated learning processing device 2. When the processor loads and executes the computer program, it can perform the access control method described in any of the preceding embodiments; therefore, it will not be repeated here. Additionally, the beneficial effects of using the same method will not be repeated here either. For technical details not disclosed in the embodiments of the computer-readable storage medium involved in this application, please refer to the description of the method embodiments of this application.
[0406] The aforementioned computer-readable storage medium can be an internal storage unit of the federated learning processing apparatus provided in any of the foregoing embodiments or the aforementioned computer device, such as a hard disk or memory of the computer device. The computer-readable storage medium can also be an external storage device of the computer device, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., provided on the computer device. Furthermore, the computer-readable storage medium can include both internal storage units and external storage devices of the computer device. The computer-readable storage medium is used to store the computer program and other programs and data required by the computer device. The computer-readable storage medium can also be used to temporarily store data that has been output or will be output.
[0407] Furthermore, it should be noted that this application also provides a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. The processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the method provided in any of the preceding corresponding embodiments.
[0408] The terms "first," "second," etc., in the specification, claims, and drawings of this application are used to distinguish different objects, not to describe a specific order. Furthermore, the term "comprising," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, apparatus, product, or device that includes a series of steps or units is not limited to the listed steps or modules, but may optionally include steps or modules not listed, or may optionally include other step units inherent to these processes, methods, apparatuses, products, or devices.
[0409] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in the foregoing description as a network element. Whether these network elements are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can implement the described network elements using different methods for each specific application, but such implementation should not be considered beyond the scope of this application.
[0410] The above-disclosed embodiments are merely preferred embodiments of this application and should not be construed as limiting the scope of this application. Therefore, any equivalent variations made in accordance with the claims of this application shall still fall within the scope of this application.
Claims
1. A blockchain-based federated learning processing method, characterized in that, The method is executed by a first computing node in a blockchain network, which further includes a task initiating node and M second computing nodes, where M is a positive integer; the method includes: The first computing node receives the federated learning task generated by the task initiating node through the task smart contract, and obtains the training data associated with the federated learning task as the first training data. The federated learning task includes the task deadline and initial model information. Based on the initial model information, a first required computing resource is determined for the federated learning task. If the first required computing resource is greater than the available computing resource, it is determined that the first computing node does not have the capability for federated learning computing on the first training data. The available computing resource refers to the idle computing resource of the first computing node. If the first required computing resource is less than or equal to the available computing resource, a first training duration corresponding to the federated learning computing on the first training data is determined based on the amount of data in the first training data. If the first training duration is greater than the task deadline, it is determined that the first computing node does not have the capability for federated learning computing on the first training data. If the first computing node does not have the computing capability for federated learning of the first training data, the first training competition request for the federated learning task is broadcast to the M second computing nodes. Among the M second computing nodes, the second computing node that wins the competition and meets the node credibility condition is determined as the target computing node; the target computing node has the computing capability for federated learning of the first training data. If a first data sharing request is received from the target computing node, the first training data is sent to the target computing node so that the target computing node can train the initial model associated with the federated learning task based on the first training data and the stored training data associated with the federated learning task to obtain the first branch training update parameters; the task initiating node is also used to globally update the initial model through N branch training update parameters; the N branch training update parameters include the first branch training update parameters, where N is a positive integer.
2. The method according to claim 1, characterized in that, The process of determining the second computing node that successfully competes among the M second computing nodes and meets the node credibility condition as the target computing node includes: During the competition period, the system receives first training competition response information sent by L competing nodes respectively; each first training competition response information contains a number of competing digital resources; the L competing nodes are nodes among the M second computing nodes that have federated learning computing capabilities for the first training data; L is a positive integer less than or equal to M. From the L competing nodes, those that do not meet the node trustworthiness condition are removed, resulting in S trustworthy competing nodes; S is a positive integer less than or equal to L. From the S trusted competing nodes, select the trusted competing node with the highest number of competing digital resources as the pre-selected computing node; Send a first training competition confirmation request to the pre-selected computing node; If the first training competition confirmation response information sent by the pre-selected computing node according to the first training competition confirmation request is received within the confirmation time period, then the pre-selected computing node is determined to be the target computing node.
3. The method according to claim 2, characterized in that, The L competing nodes include competing node Z. i , where i is a positive integer less than or equal to L; the method further includes: Obtain the node reliability probability table; Query the competing node Z from the node reliability probability table. i The probability of credibility; Based on the competing node Z i The connection relationship between them and the competing node Z i The reliable probability determines the competing node Z i Credibility; If the competing node Z i If the credibility of a node is less than the credibility threshold, then the competing node Z is determined. i The node trustworthiness condition is not met; If the competing node Z i If the credibility of a node is greater than or equal to the credibility threshold, then the competing node Z is determined. i The node trustworthiness condition is met.
4. The method according to claim 1, characterized in that, The method further includes: If the first training duration is less than or equal to the task deadline, then the first computing node is determined to have federated learning computing capabilities for the first training data.
5. The method according to claim 4, characterized in that, The blockchain network also includes P third-party computing nodes, where P is a positive integer; The method further includes: The system receives second training competition requests for the federated learning task from each of the P third computing nodes; the second training competition request sent by the j-th third computing node includes the amount of data for the j-th second training data; the j-th second training data is the training data associated with the federated learning task already stored in the j-th third computing node; the j-th third computing node does not have the federated learning computing capability for the j-th second training data; j is a positive integer less than or equal to P. If the first computing node has the federated learning computing capability for the first training data, then the third computing node that successfully competes among the P third computing nodes is determined as the target acquisition node; the first computing node also has the federated learning capability for the target training data; the target training data is the second training data corresponding to the target acquisition node; Send a federated learning task response request carrying training data information to the task initiating node; the training data information is generated based on the first training data and the target training data. If a federated learning task issuance instruction is received from the task initiating node, a second data sharing request is sent to the target acquisition node; the federated learning task issuance instruction includes the initial model associated with the federated learning task. The task initiation node receives the target training data sent by the target acquisition node according to the second data sharing request, trains the initial model according to the first training data and the target training data, and obtains the second branch training update parameters; the task initiation node is also used to perform a global update of the initial model through H branch training update parameters; the H branch training update parameters include the second branch training update parameters, where H is a positive integer.
6. The method according to claim 5, characterized in that, If the first computing node possesses the federated learning computing capability for the first training data, then among the P third computing nodes, the third computing node that successfully competed for the data is determined as the target acquisition node, including: If the first computing node has the federated learning computing capability for the first training data, then iterate through the second training competition requests sent by the P third computing nodes for the federated learning task. Obtain the second training data G from the second training contention request sent by the kth third computing node. k The amount of data, based on the second training data G k The amount of data is determined for the second training data G. k The second training duration T of the data volume k k is a positive integer less than or equal to P; The second training duration T k Add the first training duration to obtain the total training duration T. k总 ; If the total training time T k总 If the task deadline is less than or equal to the specified deadline, then the first computing node is determined to have the capability to process the second training data G. k The federated learning computing capability sends the second training competition response information R to the k-th third computing node. k ; During the competition waiting period, receive second training competition confirmation requests sent by Q third computing nodes respectively; Q is a positive integer less than or equal to Q; each second training competition confirmation request contains a target number of digital resources to compete for. The third computing node corresponding to the second training competition confirmation request with the lowest unit data resource consumption is determined as the successful third computing node. The second training competition confirmation response information is sent to the successful third computing node, and the successful third computing node is designated as the target acquisition node. The unit data resource consumption of a third computing node is determined based on the number of target competition digital resources and the amount of second training data corresponding to a third computing node.
7. The method according to claim 5, characterized in that, Also includes: If, at the same time as receiving the federated learning task issuance instruction sent by the task initiating node, random differential privacy noise is received, then the second branch training update parameters are encrypted according to the received random differential privacy noise to obtain target encrypted branch training update parameters, and the target encrypted branch training update parameters are sent to the task initiating node; the task initiating node is also used to perform information aggregation processing on S encrypted branch training update parameters to obtain encrypted aggregated training update parameters; S is a positive integer; the task initiating node is also used to add the encrypted aggregated training update parameters to the differential privacy key to obtain aggregated training update parameters, and perform global updates on the initial model according to the aggregated training update parameters; the S encrypted branch training update parameters include the target encrypted branch training update parameters.
8. The method according to claim 1, characterized in that, Also includes: Obtain the agreed-upon number of digital resources for the target computing node; The agreed-upon number of digital resources, the amount of data in the first training data, and the data sharing relationship with the target computing node are packaged into a data sharing record transaction, and the data sharing record transaction is cached in the record transaction pool. Upon receiving the first data sharing request sent by the target computing node, consensus is reached and the data sharing record transaction is uploaded to the blockchain. Receive digital resources sent by the target computing node that correspond to the agreed-upon number of digital resources.
9. A blockchain-based federated learning processing method, characterized in that, The method is executed by a task-initiating node in a blockchain network, which also includes W computing nodes, where W is a positive integer; the method includes: The task initiating node generates a federated learning task through a task smart contract and broadcasts the federated learning task to the W computing nodes. If there is a computationally restricted node among the W computing nodes, the computationally restricted node, if it is determined that it does not have the federated learning computing capability for the shared training data associated with the federated learning task, broadcasts a training competition request for the shared training data to the communicable computing node corresponding to the computationally restricted node. The computationally restricted node is also used to determine a target communicable computing node among the communicable computing nodes that has successfully competed and meets the node trustworthiness condition. The shared training data is the training data associated with the federated learning task that has been stored in the computationally restricted node. The computationally restricted node does not have the federated learning computing capability for the shared training data. Receive federated learning task response requests from Y capability computing nodes respectively; Y is a positive integer; the Y capability computing nodes include the target communicable computing node, and the Y capability computing nodes do not include the computationally restricted node; Based on the Y federated learning task response requests, X training computing nodes are determined from the Y capability computing nodes. A federated learning task issuance instruction carrying the initial model associated with the federated learning task is sent to the X training computing nodes, so that the X training computing nodes train the initial model according to the available training data to obtain branch training update parameters; X is a positive integer; if the X training computing nodes include the target communicable computing node, then the available training data corresponding to the target communicable computing node includes the shared training data and the training data already stored in the target communicable computing node; The initial model is globally updated based on the aggregated training update parameters until a target model that meets the training conditions indicated by the federated learning task is obtained; the aggregated training update parameters are generated based on the X branch training update parameters.
10. The method according to claim 9, characterized in that, A federated learning task response request includes a unit data resource consumption and a total amount of training data; The step of determining X training computing nodes from the Y capability computing nodes based on the Y federated learning task response requests includes: Obtain the predicted digital resource consumption corresponding to the federated learning task; Based on the Y federated learning task response requests, sort the Y units of data resource consumption in ascending order to obtain the sorted Y units of data resource consumption. Traverse the sorted Y units of data resource consumption, sequentially obtain the e-th unit of data resource consumption, and determine the predicted digital resource consumption of the e-th node based on the e-th unit of data resource consumption and the total amount of training data corresponding to the e-th unit of data resource consumption; e is a positive integer less than or equal to Y. Add the capability calculation node corresponding to the e-th unit data resource consumption to the node pre-selection queue; If the predicted consumption of digital resources at the e-th node is less than the current predicted consumption of remaining digital resources, then the current predicted consumption of remaining digital resources is subtracted from the predicted consumption of digital resources at the e-th node to obtain the updated predicted consumption of remaining digital resources, and the process continues to iterate to obtain the consumption of data resources at the (e+1)-th unit. If the predicted consumption of digital resources at the e-th node is greater than or equal to the current predicted consumption of remaining digital resources, or if e equals Y, then stop traversing and use all X capability calculation nodes in the node pre-selection queue as training calculation nodes.
11. The method according to claim 10, characterized in that, Also includes: X-1 training computation nodes are sequentially obtained from the node pre-selection queue; The predicted consumption of node digital resources corresponding to each of the X-1 training computing nodes is taken as the transaction digital resource consumption corresponding to each of the X-1 training computing nodes. Obtain the Xth training computation node from the node pre-selection queue. If the predicted consumption of node digital resources corresponding to the Xth training computing node is less than the target remaining predicted consumption of digital resources, then the predicted consumption of node digital resources corresponding to the Xth training computing node is taken as the transaction digital resource consumption of the Xth training computing node; the target remaining predicted consumption of digital resources refers to the value remaining after subtracting the predicted consumption of node digital resources corresponding to the X-1 training computing nodes from the predicted consumption of digital resources. If the predicted consumption of node digital resources corresponding to the Xth training computing node is greater than or equal to the predicted consumption of target remaining digital resources, then the predicted consumption of target remaining digital resources is taken as the transaction digital resource consumption corresponding to the Xth training computing node. Each training computing node and the corresponding digital resource consumption of each training computing node are associated and packaged into a digital resource record transaction, and the digital resource record transaction is cached in the record transaction pool. When a target model that meets the training conditions indicated by the federated learning task is obtained, digital resources corresponding to the associated transaction digital resource consumption are sent to each training computing node, and the transactions recorded by the digital resources are consensus-based and uploaded to the blockchain.
12. The method according to claim 9, characterized in that, Also includes: Generate a set of random differential privacy noise; the set of differential privacy noise contains random differential privacy noise corresponding to each training computation node; The total random differential privacy noise is obtained by summing the X random differential privacy noises. Use the negative of the total random differential privacy noise as the differential privacy key; While sending the federated learning task issuance instruction carrying the initial model associated with the federated learning task to the X training computing nodes, the corresponding random differential privacy noise is sent to each of the X training computing nodes, so that the X training computing nodes encrypt the branch training update parameters according to the received random differential privacy noise to obtain encrypted branch training update parameters. The encrypted branch training update parameters returned by the X training computing nodes are aggregated to obtain encrypted aggregated training update parameters. The encrypted aggregated training update parameters are then added to the differential privacy cipher to obtain aggregated training update parameters.
13. The method according to claim 12, characterized in that, The process of aggregating the encrypted branch training update parameters returned by the X training computing nodes to obtain encrypted aggregated training update parameters includes: Obtain the training and test sets; Based on the training and test set, the encrypted branch training update parameters returned by the X training computing nodes are verified respectively to obtain X update accuracies; The encrypted branch training update parameters returned by the training computation node whose update accuracy meets the training accuracy condition are used as the target encrypted branch training update parameters. The target encrypted branch training update parameters are processed by information aggregation to obtain encrypted aggregated training update parameters.
14. The method according to claim 13, characterized in that, Also includes: The X training computing nodes and the X update precisions are packaged into an update precision record transaction, and a consensus process for the update precision record transaction is initiated to the blockchain network. This allows each node in the blockchain network to update its stored precision record table according to the X update precisions when the consensus for the update precision record transaction is passed. The precision record table is used to record each node in the blockchain network and the current update precision corresponding to each node.
15. The method according to claim 14, characterized in that, Also includes: Obtain the accuracy record table based on the ballot confirmation request; The contribution of each node is determined based on the current update precision of each node in the precision record table. The resource rights of each node are determined based on the contribution weight, the contribution of each node, and the duration of digital resource ownership for each node. A ballot is sent to each node according to the proportion of resource rights corresponding to each node, so that each node can vote on the received ballot; If any of the nodes contains a computationally restricted node, the computationally restricted node is used to vote the received votes for the target communicable computing node. After receiving the votes, the number of votes for each node is used to determine the block production rights ratio for each node. A block-producing node is randomly determined from each node according to the block-producing right acquisition ratio; the block-producing node has the right to produce new blocks; Send reward digital resources to the block-producing node so that the block-producing node can allocate the reward digital resources according to the proportion of the votes. The vote composition ratio refers to the ratio between the number of votes received by the block-producing node from the task-initiating node and the number of votes cast by the target computationally limited node.
16. A blockchain-based federated learning processing device, characterized in that, The device is applied to a first computing node in a blockchain network, which further includes a task initiation node and M second computing nodes, where M is a positive integer; the device comprises: The task receiving module is used to receive the federated learning task generated by the task initiating node through the task smart contract, and to obtain the training data associated with the federated learning task as the first training data. The federated learning task includes the task deadline and initial model information. The competition broadcast module is used to determine the first required computing resources corresponding to the federated learning task based on the initial model information; if the first required computing resources are greater than the available computing resources, it is determined that the first computing node does not have the computing capability for federated learning of the first training data; the available computing resources refer to the idle computing resources of the first computing node; if the first required computing resources are less than or equal to the available computing resources, it is determined that the first training duration corresponding to the federated learning of the first training data is determined based on the amount of data in the first training data; if the first training duration is greater than the task deadline, it is determined that the first computing node does not have the computing capability for federated learning of the first training data; if the first computing node does not have the computing capability for federated learning of the first training data, the first training competition request for the federated learning task is broadcast to the M second computing nodes. A node determination module is used to determine, from the M second computing nodes, the second computing node that successfully competes and meets the node credibility condition, as the target computing node; the target computing node has federated learning computing capabilities for the first training data. The data sharing module is configured to send the first training data to the target computing node upon receiving a first data sharing request from the target computing node, so that the target computing node can train the initial model associated with the federated learning task based on the first training data and the stored training data associated with the federated learning task to obtain the first branch training update parameters; the task initiating node is further configured to globally update the initial model using N branch training update parameters; the N branch training update parameters include the first branch training update parameters, where N is a positive integer.
17. A computer device, characterized in that, include: Processor, memory, and network interface; The processor is connected to the memory and the network interface, wherein the network interface is used to provide data communication functions, the memory is used to store program code, and the processor is used to call the program code to execute the method according to any one of claims 1-15.
18. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program adapted to be loaded by a processor and to execute the method according to any one of claims 1-15.
19. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, they can perform the method described in any one of claims 1-15.
Citation Information
Patent Citations
Data processing method, device and system and server
CN112132198A
Federal learning method and system based on block chain and trusted execution environment
CN113837761A
Method for managing computing capacities in a network with mobile participants
US20220050725A1