Blockchain-based distributed ensemble learning method

By designing a three-layer blockchain structure, the implementation of distributed integrated learning tasks is solved, and the problem of low computing power utilization and insufficient decentralization in the existing PoUW protocol is improved, and the robustness and decentralization of the system are improved.

WO2025123956A1PCT designated stage Publication Date: 2025-06-19SOUTHEAST UNIV

Patent Information

Application Number
PCT/CN2024/127661
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-09-09
Filing Date
2024-10-28
Publication Date
2025-06-19

AI Technical Summary

Technical Problem

The problems of low computing power utilization and insufficient decentralization in the existing PoUW protocols are especially easy to introduce central nodes during model aggregation, which affects the robustness and decentralization of the system.

Method used

A three-layer blockchain structure is designed, including mini blocks, integrated blocks and key blocks. Through this structure, the execution of distributed integrated learning tasks is realized, avoiding the introduction of central nodes in the model aggregation process, and maximizing the degree of decentralization of blockchain.

Benefits of technology

It improves the computing power utilization rate of the blockchain network, realizes the automation and distributed execution of the integrated learning process, enhances the robustness and decentralization of the system, and avoids single point failure of the central node.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024127661_19062025_PF_FP_ABST
    Figure CN2024127661_19062025_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed in the present application is a blockchain-based distributed ensemble learning method. On the basis of a designed three-layer blockchain structure, which consists of a mini block, an ensemble block and a key block, by means of a consensus protocol, a miner in a network is enabled to train a base model on a training set after performing sampling with replacement, and aggregates models from other miners; and finally, information of the base model and information of an ensemble model are recorded in a blockchain, such that the whole process of model training, model ensemble and model evaluation is fused into a blockchain consensus, and the whole ensemble learning process is automatically executed in a blockchain network. In this way, the present application can increase the utilization rate of the computing power in a blockchain network by means of a proof-of-useful-work mechanism based on machine learning, and a center node is prevented from being introduced during a model aggregation process, thereby maximizing the degree of decentralization of a blockchain.
Need to check novelty before this filing date? Find Prior Art

Description

A distributed ensemble learning method based on blockchain

[0001] Related applications

[0002] This application claims priority to the Chinese patent application filed with the China Patent Office on September 9, 2024, with application number 2024112560741 and application name “A Distributed Integrated Learning Method Based on Blockchain”, the entire contents of which are incorporated herein by reference. Technical Field

[0003] The present application relates to the field of blockchain technology, and in particular to a distributed ensemble learning method based on blockchain. Background Art

[0004] Blockchain is a distributed data storage technology that uses consensus protocols and incentive mechanisms to ensure data consistency across all nodes in an unreliable peer-to-peer network. Consensus protocols are the cornerstone of blockchain, and their design fundamentally determines the performance characteristics of blockchain systems, including throughput, consistency, scalability, and robustness. Currently, Proof-of-Work (PoW) is one of the most popular consensus protocols. In PoW-based systems, miners generate new blocks by searching for a random number that makes the block's hash value below a target value. However, the mining process consumes a large amount of energy, much of which is wasted in meaningless hash calculations.

[0005] Currently, there are two solutions to the sustainability issue of PoW protocols: one is to reduce the computational effort involved in the consensus process. The other is the Proof-of-Useful-Work (PoUW) protocol. This latter approach uses computational tasks with practical applications as proof of computational effort, thereby satisfying third-party computing power requirements. Existing PoUW protocols primarily use machine learning and optimization tasks as proof of useful work. Cutting-edge research is also integrating federated learning with blockchain consensus, transforming mining pools in PoW systems into clusters of miners competing for machine learning training rewards. The FedAvg algorithm aggregates the models of each mining pool into a global model.

[0006] In the PoUW protocols described above that utilize machine learning and federated learning, most machine learning-based PoUW schemes select a single winning model and discard the remaining models, inevitably leading to low computing power utilization. As for existing consensus mechanisms based on federated learning, although models are trained in a distributed manner by data holders, the robustness of these protocols is limited by the central node, as the models are typically aggregated by one or several central nodes. Therefore, a technology is needed to aggregate the computing power of each node in a decentralized blockchain network.

[0007] Summary of the Invention

[0008] Purpose of the Invention: To address the problems of low computing power utilization and insufficient decentralization in the existing PoUW protocol, the exemplary embodiments of this application aim to design a three-layer blockchain structure consisting of mini blocks, integrated blocks, and key blocks, as well as a distributed integrated learning method based on blockchain, to improve the computing power utilization of the blockchain network, while avoiding the introduction of central nodes in the model aggregation process and maximizing the decentralization of the blockchain.

[0009] Technical solution: To achieve the above-mentioned invention objectives, this application adopts the following technical solution:

[0010] According to one aspect of the present application, a three-layer blockchain structure is proposed to enable the execution of distributed ensemble learning tasks, including:

[0011] MiniBlock, each miniblock corresponds to a unique base model and contains the identifier of the machine learning model parameters, the identifier of the model owner, the hash of the previous key block, the machine learning task hash, and the timestamp.

[0012] Ensemble Block, which contains the performance indicators of the ensemble model on the validation dataset, the identifier of the model integrator, the hash of the mini-block corresponding to the integrated base model, the machine learning task hash, and the timestamp.

[0013] The key block contains the hash of the previous key block, the machine learning task hash, the miner identifier, the performance indicators of the optimal integrated model on the test dataset, the timestamp of the key block generation, the root hash of the hash tree carrying the transaction data, the random number, the task queue, the hash of each integrated block participating in the ranking, and its performance indicators.

[0014] According to one aspect of the present application, a method for generating a three-layer blockchain structure that enables distributed ensemble learning task execution is proposed, comprising the following steps:

[0015] After the miner completes the base model training on the training dataset, it generates and broadcasts a mini-block. Each mini-block corresponds to a unique base model and contains the identifier of the machine learning model parameters, the identifier of the model owner, the hash of the previous key block, the machine learning task hash, and a timestamp.

[0016] After the miners aggregate the base models and evaluate their performance on the validation dataset, they generate and broadcast integration blocks. Each integration block contains the performance metrics of the integrated model on the validation dataset, the identifier of the model integrator, the hash of the mini-block corresponding to the integrated base model, the machine learning task hash, and a timestamp.

[0017] After evaluating the performance indicators of the collected ensemble models on the test dataset, miners select the optimal ensemble model, generate and broadcast a key block; the key block contains the hash of the previous key block, the machine learning task hash, the miner identifier of the block producer, the performance indicators of the optimal ensemble model on the test dataset, the timestamp of the key block generation, the root hash of the hash tree carrying the transaction data, the random number, the task queue, the hash of each ensemble block participating in the ranking, and its performance indicators.

[0018] According to one aspect of the present application, a distributed ensemble learning method based on blockchain is proposed, comprising the following steps:

[0019] Step 1: The task publisher prepares the training, verification, and test datasets in advance and publishes the machine learning task through blockchain transactions. The task includes the hash values ​​of the three datasets and scripts for implementing base model training, model integration, and performance evaluation.

[0020] Step 2: The miner verifies the task information and forwards it to other miners.

[0021] Step 3: After the miner generates the last key block or receives a valid last key block from another miner, it downloads the training dataset from the task publisher. After the download is complete, the miner begins training the base model, generates mini-blocks, and broadcasts them.

[0022] Step ④: The task publisher publishes the verification dataset. After receiving the message, the miners start to aggregate the base model and evaluate its performance on the verification dataset, generate an integrated block and broadcast it.

[0023] Step 5: The task publisher publishes the test dataset. After receiving the message, the miners begin to evaluate the performance indicators of the collected ensemble models on the test dataset, select the optimal ensemble model, calculate the key block hash, generate the key block and broadcast it.

[0024] Step ⑥: The task publisher obtains the performance indicators of the best integrated model and the corresponding base model parameters from the key blocks, and obtains an integrated model with ideal performance by aggregating the base models.

[0025] Preferably, in one embodiment of the present application, the machine learning task published by the user through the blockchain transaction includes a transaction used as training fee, which transfers the training fee to a virtual address.

[0026] Preferably, in one embodiment of the present application, miners who generate key blocks and mini blocks will be rewarded, the training fees of the task publisher will be evenly distributed to all miners who generate the base model used by the winning integrated model, and the miners who generate key blocks will be rewarded with newly generated tokens.

[0027] Optionally, in one embodiment of the present application, miners can use local private data together with public training data provided by the task publisher for base model training, thereby improving the quality of the base model and improving the quality of the integrated model.

[0028] Preferably, in one embodiment of the present application, the integrated model is obtained by aggregating the base models through the Bagging algorithm.

[0029] According to one aspect of the present application, a computer system is proposed, comprising a memory, a processor, and a computer program / instruction stored in the memory and executable on the processor, wherein the computer program / instruction implements the steps of the aforementioned method when executed by the processor.

[0030] According to one aspect of the present application, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the aforementioned method are implemented.

[0031] According to one aspect of the present application, a computer program product is provided, comprising a computer program / instruction, wherein the computer program / instruction implements the steps of the aforementioned method when executed by a processor.

[0032] The distributed ensemble learning method based on blockchain provided in this application has the following beneficial effects:

[0033] (1) This application designs a three-layer blockchain structure and embeds the distributed ensemble learning process into the generation, verification and propagation of three types of blocks, thereby realizing the execution of distributed ensemble learning tasks in the blockchain.

[0034] (2) Compared with the existing machine learning-based PoUW protocol, this application realizes the automation and distributed execution of the integrated learning process in the blockchain network, and forms an integrated model with better performance by aggregating base models trained by multiple miners; by allowing miners to use private data, the quality of the base model and the integrated model is further improved, and even if the private data of different miners are not independent and identically distributed, the integrated model obtained after aggregating the base models still has good performance.

[0035] (3) Compared with the federated learning proof mechanism, this application does not require the mining pool administrator to aggregate the base model, can operate on the public chain, and has a higher degree of decentralization.

[0036] (4) This application encourages miners to participate in maintaining the blockchain, training models, integrating models, and evaluating models by designing an incentive mechanism, so that the blockchain based on this application can operate stably and sustainably.

[0037] (5) The performance of the integrated model generated by this application improves as the number of miners in the blockchain using private data to train the base model increases.

[0038] Additional aspects and advantages of the present application will be given in part in the description below, and in part will become apparent from the description below, or will be learned through practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] The above and / or additional aspects and advantages of the present application will become apparent and easily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which:

[0040] FIG1 is a schematic diagram of a three-layer block structure for enabling distributed ensemble learning task execution according to an embodiment of the present application;

[0041] FIG2 is a general flow chart of a distributed ensemble learning method based on blockchain according to an embodiment of the present application;

[0042] FIG3 is a schematic diagram of various stages of a task execution process in a distributed ensemble learning method based on blockchain according to an embodiment of the present application;

[0043] FIG4 is a schematic diagram showing the performance improvement of the integrated model generated by the embodiment of the present application relative to the base model and the model trained directly on the public data under different proportions of public data;

[0044] FIG5 is a schematic diagram showing the accuracy of the integrated model generated by the embodiment of the present application under different network scales and network connectivity. DETAILED DESCRIPTION

[0045] The technical solution of the present application is described in detail below with reference to the accompanying drawings and specific embodiments. Those skilled in the art should understand that the embodiments described below with reference to the accompanying drawings are exemplary and intended to be used to explain the present application, and should not be understood as limiting the present application.

[0046] The following describes, with reference to the accompanying drawings, a blockchain-based distributed ensemble learning method according to an embodiment of the present application. To address the issues of low computing power utilization and insufficient decentralization in existing PoUW protocols, as mentioned in the background art, this embodiment of the present application provides a blockchain-based distributed ensemble learning method that uses a bagging algorithm to fuse models trained by multiple miners. The disclosed three-layer blockchain structure and distributed ensemble learning method can automatically execute model training, model integration, and model evaluation processes within a fully decentralized public blockchain.

[0047] Specifically, Figure 1 is a three-layer blockchain structure that enables the execution of distributed ensemble learning tasks according to an embodiment of the present application.

[0048] As shown in Figure 1, the three data structures contained in the three-layer blockchain structure and the connection relationship between the three data structures are shown:

[0049] Mini-block (denoted as ), used to record the ownership of the model and the hash of the model parameters. Each mini-block corresponds to a unique base model and contains the identifier M of the model owner. i , the identifier Hash (ω) calculated by connecting the machine learning model parameters with the miner identifier i ||M i ), the hash of the previous key block (KB h-1 ), machine learning task hash Hash (T h ) and timestamp

[0050] Integrated block (expressed as ), used to record the base models required for each ensemble model and the performance indicators of the ensemble model on the validation dataset. The ensemble block points to multiple mini-blocks, and the base models corresponding to the mini-blocks are aggregated to obtain an ensemble model. The ensemble block contains the performance indicators of the ensemble model on the validation dataset D V Performance indicators on Model integrator identifier M k , hash of the aggregated base model Machine Learning Task Hash(T h ) and timestamp

[0051] Key Block (KB h ), used to store blockchain transactions, task queues, and integrated model rankings, etc., including the hash of the previous key block (KB h-1 ), machine learning task hash Hash (T h ), block miner identifier The optimal ensemble model is in the test dataset D E Performance indicators on MTC best , the timestamp Γ(KB) generated by the key block h ), the root hash MKR of the hash tree that carries the transaction data h , random number Nonce, task queue Task Queue, hash of each integrated block participating in the ranking as well as The corresponding integrated model is in the test dataset D E Performance indicators on

[0052] An embodiment of the present application provides a three-layer blockchain structure generation method that enables the execution of distributed ensemble learning tasks. The main steps include: after the miner completes the base model training on the training data set, the miner generates and broadcasts the mini-blocks in the structure shown in Figure 1; after the miner aggregates the base model and performs performance evaluation on the verification data set, the miner generates and broadcasts the integrated block in the structure shown in Figure 1; after the miner evaluates the performance indicators of the collected integrated models on the test data set, the miner selects the optimal integrated model and generates and broadcasts the key blocks in the structure shown in Figure 1.

[0053] As shown in Figures 2 and 3, Figure 2 is a distributed ensemble learning method based on blockchain provided according to an embodiment of the present application, and Figure 3 shows a schematic diagram of each stage of the task execution process in the distributed ensemble learning method based on blockchain.

[0054] In step ① "Task Release" in Figure 2, the task publisher prepares the training, verification, and test data sets in advance and publishes the machine learning task through a blockchain transaction. The task contains the hash values ​​of the three data sets and scripts for implementing base model training, model integration, and performance evaluation. It also includes a transaction to pay the training fee, which transfers the training fee to a virtual address.

[0055] In step ② (Task Verification and On-Chain) in Figure 2, miners verify the task information and forward it to other miners. The task will eventually enter the task queue of the key block.

[0056] Steps ③ through ⑤ in Figure 2 illustrate the task execution process. Phase 1 (Base Model Training) in Figure 3 corresponds to step ③ (Base Model Training) in Figure 2, phase 2 (Ensemble Block Generation) corresponds to step ④ (Ensemble Block Generation) in Figure 2, and phase 3 (Key Model Generation) corresponds to step ⑤ (Key Block Generation) in Figure 2.

[0057] The task execution process in the embodiment of the present application includes the following three stages:

[0058] When the block height is h-1, the key block KB h-1 After being generated by the miner, it will trigger the first stage of the task execution process. i KB received h-1 After that, if KB h-1 After verification, miner M i The base model will start training. If the miner M i Receive one or more valid forks and select the longest fork (i.e. the one with the largest key block height at the end). If the key blocks at the end of these forks have the same height, select the one with the best MTC.best The fork where the key block is located is used as the main chain and the tasks at the current height are performed on the main chain. Once the miner M i Confirm KB h-1 , it will choose in KB h-1 The first task in the task queue is to download the public training dataset D T , D T With private datasets Merge into local dataset Then, miner M i In local dataset Execute the training script and generate the parameters ω of the base model i Once the base model is ready, miner M i Timestamp Miner's identifier M i 、Task Hash(T h ), model parameter identifier Hash(ω i ||M i ) and Hash(KB h-1 ) packaged in and broadcast to the blockchain network All miners receive mini-blocks before the second phase and broadcast them to neighboring miners, but do not transmit the base model to the outside to avoid model theft.

[0059] The task publisher can publicly verify the dataset D at t1 V To trigger the second stage of the task execution process. Once the honest miner M i After receiving the validation dataset, it will refuse to accept the new mini-block and perform the following steps: (1) Download the validation dataset D V And according to The identifier Hash(ω j ||M j ) Get the model parameters ω j ; (2) Miner M i Verify all collected mini-blocks, if the key block pointed to by the mini-block is invalid, or the downloaded model parameters ω j With the identifier Hash(ω j ||M j ) does not match, or the performance index of the base model corresponding to the mini-block on the verification data set is lower than the minimum tolerance value given by the task publisher, then the mini-block is considered invalid, and the miner M i All invalid mini-blocks will be discarded; (3) Miner M i Use the validation dataset D V Evaluate the performance index of the ensemble model formed by aggregating the base models corresponding to all valid mini-blocks (4) Encapsulate timestamp Miner identifier M i 、Task Hash(T h ), the performance indicators obtained by evaluation A hash of the mini-blocks corresponding to all aggregated base models And broadcast to the blockchain network. Before the start of the third phase, if miner M i Receive any integrated blocks from other miners Miner M i verify And broadcast the integrated block to neighboring miners.

[0060] The task publisher can publicly release the test dataset D at t2 E To trigger the third stage of the task execution process. i Receive the test dataset D E After that, we will use the test dataset D E Evaluate all verified integrated blocks and then try to generate a new key block. Assume that miner M i Already verified and point to and Contains Hash(ω j,l ||M l ), where η j yes The number of mini-blocks pointed to, l is no greater than η j A positive integer, ω j,l By miner M l Then miner M i Calculate K performance indicators and from a series of Find the optimal performance indicator MTC best Assumptions is the best performance indicator. Then the miner M i Load all key block components into candidate blocks By changing the random number and calculating the hash value of the candidate block, the miner will eventually find a block that satisfies Valid key blocks And after receiving a valid key block KB from other miners h Former broadcaster If a valid key block is found and broadcast Received valid key block KB from other miners before h , then the miners abandon And start executing the machine learning task T to be executed at block height h h The target is the static threshold that controls the difficulty of generating key blocks, which is equivalent to the target value in the PoW consensus.

[0061] In step ⑥ “Task End” in Figure 2, the task publisher can use the key block KB at block height h to h After downloading the base model parameters from the blockchain network and aggregating the base models, the task publisher will obtain an integrated model with ideal performance.

[0062] 4 and 5 , in order to reveal the performance of the present application in actual work, actual testing and data recording were carried out for a typical embodiment of the present application, and the analysis results are as follows.

[0063] Figure 4 shows the performance improvement of the integrated model (BagChain) generated by this application relative to the base model (Base) and the model (Public) trained solely on public data under different public data ratios. It can be seen that the accuracy (Accuracy) of the integrated model on the test dataset is better than any of their base models. When the public dataset ratio (Public Dataset Ratio) κ is less than 0.3, the model trained on the public dataset performs worse than the base model trained using both private and public data. This can demonstrate the value of this application, because when all miners have private datasets that can be used for machine learning tasks, task publishers can obtain high-quality machine learning models from this application. It can also be observed that as the public dataset ratio κ increases, the accuracy of the base model and the integrated model in this application are improved. This is because the increase in the public dataset ratio κ leads to an increase in the data samples in the miner's local training dataset, thereby generating more accurate base models and integrated models. Furthermore, the difference between the ensemble models obtained after training the base model for 20 and 30 epochs is negligible, indicating that 20 epochs are sufficient to generate a sufficiently good base model, and the aggregated model has ideal performance, thus demonstrating that training the base model in this application is lightweight. Therefore, more resource-constrained nodes can participate in model training without fine-tuning the machine learning model, and their models can be aggregated to improve the performance of the ensemble model in this application.

[0064] In more general cases, the private datasets of different miners may not be independent and identically distributed, which is reflected in the imbalance in the number of labels of each category in the miners' private datasets, and the machine learning models trained on these private datasets will have obvious deviations. However, even if this imbalance exists in the miners' private datasets, the ensemble model generated by this application still has significant performance advantages over the base model and the model trained solely on public data, and when the proportion of public datasets κ is large, the performance of the ensemble model is almost unaffected by the imbalance in the number of labels. That is, in this application, increasing the number of public training data samples can offset the impact of the imbalance in the number of labels in the miners' private datasets to a certain extent.

[0065] Figure 5 shows the accuracy of the integrated model generated by the present application under different network scales and network connectivity. In the figure, a sparse network (Sparse Network) is a network model with worse connectivity than a fully connected synchronous network (Synchronous Network). It can be seen that as the number of miners (Miner Number) increases, the average accuracy on the test data set increases. This is because as the network scale expands, more private data and computing power are invested in the training of the base model. In addition, a higher private data set ratio (Private Dataset Ratio) significantly improves the accuracy of the integrated model generated in this application, indicating that under higher levels of miner participation and private data abundance, the performance of the integrated model in this application is improved. In addition, comparing the accuracy under the two network models shows that in the sparse network model, the performance of the present application is basically unaffected, which highlights the robustness of the present application in a more severe network environment.

[0066] This application discloses a distributed ensemble learning method based on blockchain. It designs a three-layer blockchain structure consisting of mini-blocks, integrated blocks, and key blocks. By designing a consensus protocol, miners in the network train a base model on a training set sampled with replacement and aggregate models from other miners. Ultimately, information about the base model and integrated model is recorded on the blockchain, thereby integrating the entire process of model training, model integration, and model evaluation into the blockchain consensus mechanism. This method can automate and distribute the entire ensemble learning process in a blockchain network, improve the computing power utilization of the machine learning-based proof-of-work mechanism in the blockchain network, and avoid the introduction of central nodes in the model aggregation process, thereby maximizing the decentralization of the blockchain.

[0067] An embodiment of the present application provides a computer system, including a memory, a processor, and a computer program / instruction stored in the memory and executable on the processor, wherein the computer program / instruction implements the steps of the aforementioned method when executed by the processor.

[0068] An embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the aforementioned method are implemented.

[0069] An embodiment of the present application provides a computer program product, including a computer program / instruction, which implements the steps of the aforementioned method when executed by a processor.

Claims

1. A three-layer blockchain structure that enables distributed ensemble learning task execution, including: Mini-blocks, each mini-block corresponds to a unique base model and contains the identifier of the machine learning model parameters, the identifier of the model owner, the hash of the previous key block, the machine learning task hash, and a timestamp; The integration block contains the performance indicators of the integrated model on the validation dataset, the identifier of the model integrator, the hash of the mini-block corresponding to the integrated base model, the machine learning task hash and the timestamp; as well as The key block includes the hash of the previous key block, the machine learning task hash, the block miner identifier, the performance indicators of the optimal integrated model on the test data set, the timestamp of the key block generation, the root hash of the hash tree carrying the transaction data, the random number, the task queue, the hash of each integrated block participating in the ranking, and its performance indicators.

2. A method for generating a three-layer blockchain structure that enables distributed ensemble learning tasks, wherein: The following steps are involved: After the miners complete the base model training on the training data set, they generate and broadcast mini blocks; each mini block corresponds to a unique base model, containing the identifier of the machine learning model parameters, the identifier of the model owner, the hash of the previous key block, the machine learning task hash, and a timestamp; After the miners aggregate the base models and evaluate their performance on the validation dataset, they generate and broadcast integration blocks; each integration block contains the performance indicators of the integrated model on the validation dataset, the identifier of the model integrator, the hash of the mini-block corresponding to the integrated base model, the machine learning task hash, and the timestamp; as well as After the miners evaluate the performance indicators of the collected integrated models on the test data set, they select the optimal integrated model, generate and broadcast the key block; the key block contains the hash of the previous key block, the machine learning task hash, the block miner identifier, the performance indicators of the optimal integrated model on the test data set, the timestamp of the key block generation, the root hash of the hash tree carrying the transaction data, the random number, the task queue, the hash of each integrated block participating in the ranking, and its performance indicators.

3. A distributed ensemble learning method based on blockchain, wherein: The following steps are involved: The task publisher prepares training, validation, and test data sets in advance, and publishes machine learning tasks through blockchain transactions. The tasks contain the hash values ​​of the three data sets and scripts for implementing base model training, model integration, and performance evaluation. The miner verifies the task information and forwards it to other miners. When the miner generates the last key block or receives a valid last key block from other miners, it downloads the training data set from the task publisher. After the download is complete, the miner starts training the base model, generates a mini block and broadcasts it. The mini block contains the identifier of the machine learning model parameters, the identifier of the model owner, the hash of the last key block, the machine learning task hash and a timestamp. The task publisher publishes the verification data set. After receiving the message, the miners start to aggregate the base model and evaluate its performance on the verification data set, generate integrated blocks and broadcast them; The integrated block includes the performance index of the integrated model on the validation data set, the identifier of the model integrator, the hash of the mini-block corresponding to the integrated base model, the machine learning task hash and the timestamp; The task publisher publishes the test data set. After receiving the message, the miners start to evaluate the performance indicators of the collected ensemble models on the test data set, select the optimal ensemble model, calculate the key block hash, generate the key block and broadcast it; Said The key block contains the hash of the previous key block, the machine learning task hash, the miner identifier, the performance index of the optimal integrated model on the test data set, the timestamp of the key block generation, the root hash of the hash tree carrying the transaction data, the random number, the task queue, the hash of each integrated block participating in the ranking, and its performance index; as well as The task publisher obtains the performance indicators of the best integrated model and the corresponding base model parameters from the key blocks, and obtains an integrated model with ideal performance by aggregating the base models.

4. The distributed ensemble learning method based on blockchain according to claim 3, wherein: The machine learning task published by the user through the blockchain transaction includes a transaction used as training fee, which transfers the training fee to a virtual address.

5. The distributed ensemble learning method based on blockchain according to claim 3, wherein: Miners who generate key blocks and mini blocks will be rewarded, and the training costs of the task publisher are evenly distributed to all miners who generate base models used by the winning integrated model. At the same time, miners who generate key blocks are rewarded through newly generated tokens.

6. The distributed ensemble learning method based on blockchain according to claim 3, wherein: Miners use local private data together with the public training data provided by the task publisher for base model training.

7. The distributed ensemble learning method based on blockchain according to claim 3, wherein: The integrated model is obtained by aggregating the base models through the Bagging algorithm.

8. A computer system comprising a memory, a processor, and a computer program / instruction stored in the memory and executable on the processor, wherein: When the computer program / instructions are executed by a processor, the steps of the method according to any one of claims 2 to 7 are implemented.

9. A computer-readable storage medium storing a computer program, wherein: When the computer program is executed by a processor, the steps of the method according to any one of claims 2 to 7 are implemented.

10. A computer program product comprising a computer program / instructions, wherein: When the computer program / instructions are executed by a processor, the steps of the method according to any one of claims 2 to 7 are implemented.

Citation Information

Patent Citations

  • Flexible block chain framework

    CN106775619A

  • Block chain-based trusted computing storage method

    CN114785509A

  • Industrial chain financial risk control model construction method based on block chain

    CN117745409A

  • System and Method for Processing a Database Query

    US20210109917A1

Cited By

  • Algorithm credit scoring system and method based on decentralized continuous effect authentication

    CN120579945A

  • Algorithm-based credit scoring system and method based on decentralized continuous performance verification

    CN120579945B

  • Block chain architecture based on branch and bound proof of workload

    CN121051176A

  • Graph neural network distributed training method, device and system based on decentralized batch processing

    CN121724106A