A Scalable Sharding Consensus Algorithm Based on Clustering

By adopting the scalable shard consensus algorithm KBFT based on clustering in the consortium chain, the problem of communication blocking and low throughput of the PBFT algorithm when the number of nodes increases is solved, and more efficient, secure and scalable system performance is achieved.

CN116094721BActive Publication Date: 2025-06-24XINJIANG UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211518184.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-30
Publication Date
2025-06-24
Estimated Expiration
2042-11-30

AI Technical Summary

Technical Problem

The classic Byzantine fault-tolerant algorithm PBFT in the consortium chain will cause problems of communication blocking and low throughput when the number of nodes increases.

Method used

The scalable shard consensus algorithm KBFT based on clustering is adopted to shard nodes through the K-prototype clustering algorithm, and a consensus algorithm combining BLS aggregation signature and Byzantine fault-tolerant algorithm are used in each shard to dynamically reshape to deal with network changes.

Benefits of technology

It significantly reduces communication complexity, improves system throughput, enhances system security and robustness, and ensures system scalability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116094721B_ABST
    Figure CN116094721B_ABST
Patent Text Reader

Abstract

The present invention is a scalable sharding consensus algorithm based on clustering. A scalable sharding consensus algorithm based on clustering includes the following steps: (1) using the K-prototype clustering algorithm to perform sharding processing on the nodes in the network according to mixed attributes; (2) performing a consensus process, a merging and distribution process on the corresponding transactions of each of the sharding processes; (3) dynamically re-sharding: after the time interval set by the network ends, the K-prototype clustering algorithm performs re-sharding on all the nodes in the network; (4) selecting shard proxy nodes and global proxy nodes. The scalable sharding consensus algorithm based on clustering described in the present invention ensures the security and reliability of the system while significantly reducing the communication complexity and improving the system throughput.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention specifically relates to an extensible sharding consensus algorithm based on clustering. Background Art

[0002] Blockchain technology was born in 2008. As the core technology of the blockchain, the consensus algorithm plays a crucial role in the blockchain. Currently, blockchains can be divided into three types according to the level of openness, namely public blockchains, consortium blockchains, and private blockchains. Therefore, the consensus algorithm is correspondingly divided into public blockchain consensus algorithms, consortium blockchain consensus algorithms, and private blockchain consensus algorithms.

[0003] Public blockchain consensus algorithms are usually proof-based, such as Proof-of-Work (PoW) and Proof-of-Stake (PoS). As the first consensus algorithm in blockchain technology, PoW requires each node to compete for the right to record transactions by solving a mathematical puzzle. However, due to the unfairness caused by the concentration of computing power and the huge amount of electrical resources consumed by the calculation, it has not been widely used. The POS algorithm selects the recording node according to the amount of equity held. Although POS solves the problem of power consumption of the PoW algorithm, it will weaken decentralization. Although the proof-based consensus mechanism has very good node scalability, it will bring problems such as low throughput and long latency.

[0004] With the rapid development of blockchain technology, currently the blockchain has developed from the era of public blockchains to the era of consortium blockchains. Consortium blockchains have become the preferred blockchains for many fields and applications. At the same time, lighter consensus algorithms are more commonly used in consortium blockchains, such as Paxos, Raft, and the classic Byzantine Fault Tolerance algorithm PBFT, etc. Among them, Paxos and Raft are widely used in systems without Byzantine nodes. However, due to the variability and uncertainty of network attacks, there may be Byzantine nodes in the network. At this time, the PBFT algorithm has more advantages compared to the Paxos and Raft consensus algorithms. At the same time, PBFT does not require a large amount of computing, so it has been widely used in consortium blockchains.

[0005] Many scholars have carried out extensive research to improve and perfect the PBFT algorithm. For example, the EPBFT protocol proposed by some scholars uses a verifiable random function (VRF) to implement the selection of consensus nodes, making the protocol applicable to dynamic networks. The DGBFT consensus protocol proposed by some scholars greatly reduces the communication complexity by grouping nodes. The CDBFT consensus protocol proposed by some scholars stimulates the enthusiasm of nodes through a voting reward and punishment scheme and its corresponding credit evaluation scheme. Although the above protocols improve the PBFT algorithm from different focuses, when the network scale increases, there are problems such as high communication complexity or too complex credit models used. Therefore, it is crucial to design a consensus algorithm applicable to large-scale consortium chains. Summary of the Invention

[0006] The purpose of the present invention is to provide a clustering-based scalable sharding consensus algorithm (KBFT). Aiming at the problems of communication blockage and low throughput caused by the increase in the number of nodes in the classical Practical Byzantine Fault Tolerance algorithm (PBFT) mainly used in consortium chains, the operating efficiency of the system is improved, and at the same time, the security and robustness of the system are guaranteed.

[0007] In order to achieve the above purpose, the technical solution adopted is as follows:

[0008] A clustering-based scalable sharding consensus algorithm includes the following steps:

[0009] (1) Use the K-prototype clustering algorithm to perform sharding processing on the nodes in the network according to mixed attributes;

[0010] (2) Perform a consensus process, a merging and distribution process on the corresponding transactions of each shard;

[0011] (3) Dynamically re-shard: After the time interval set by the network ends, the K-prototype clustering algorithm re-shards all the nodes in the network;

[0012] (4) Select shard proxy nodes and global proxy nodes.

[0013] Further, in the step (1), classification is performed according to the numerical attributes and categorical attributes of the nodes.

[0014] Still further, in the step (1), the classification steps are as follows:

[0015] a: Randomly select g nodes as initial prototypes;

[0016] b: Assign the node objects to the nearest cluster according to the dissimilarity, and update the prototype of the cluster after assignment;

[0017] c: Redefine the prototype of the category;

[0018] d: Repeat steps bc until no node samples change category, and return the final sharding result.

[0019] Furthermore, in step (1), the node data set X = {X1, X2, X3, ..., Xn}, n is the number of node objects in the data set X, and each node data in the node data set has m attributes, that is, Xi = {x i1 ,x i2 ,x i3 ,...,x ip , x i(p+1 ),x i(p+2) ,xim}, where the numerical attributes are in the front, totaling p, and the categorical attributes are in the back, totaling mp;

[0020] The dissimilarity formula is as follows:

[0021]

[0022] In the formula, The parameter γ is the classification attribute weight.

[0023] Furthermore, in step (2), the specific operation steps of the consensus process, merging and distribution process are as follows:

[0024] ① Request phase: The client sends the request message to the proxy node of the shard where it is located, and the consensus on the request will be carried out within the shard;

[0025] ② Pre-preparation phase: The proxy node in the shard builds a new block and broadcasts the block to the rest of the consensus nodes in the shard;

[0026] ③ Preparation phase 1: The node verifies the block. If the verification is valid, it uses the BLS algorithm to sign it and feeds the signature back to the proxy node;

[0027] ④ Preparation phase 2: The proxy node waits for and collects valid signatures from other consensus nodes. When it receives at least 2f+1 identical signature messages, it aggregates all individual signatures into a BLS multi-signature, and then broadcasts the BLS multi-signature to all nodes. At this point, the preparation phase ends;

[0028] ⑤ Commitment Phase 1: The node will verify whether the received multi-signature contains at least 2f+1 signers, verify the transactions in the block broadcast by the proxy node in the pre-preparation phase, sign the message received in the preparation phase 2, and send it to the proxy node;

[0029] ⑥Submission phase 2: The proxy node waits for and collects at least 2f+1 valid signatures, aggregates these signatures together to form a BLS aggregate signature, and submits a new block with the BLS aggregate signature, and then broadcasts the new block to all nodes for verification and submission. At this point, the submission phase ends;

[0030] ⑦ Reply phase: After the node submission is completed, a reply message will be sent to the client node. When the client receives at least f+1 identical confirmation messages sent by different nodes, it indicates that the current request has reached the final consensus;

[0031] ⑧ Feedback stage: When the client receives the reply message in the shard, it feeds back all the reply messages to the supervisory node, and the node scores the behavior of the node according to the feedback result of the client;

[0032] ⑨Merge and distribution phase: The proxy nodes of each shard will broadcast the local blocks they have formed to the global proxy node. The global proxy node will merge all the received blocks into a global block and distribute it to each shard proxy node; the proxy node that receives the global block will send it to the node in the shard where it is located.

[0033] Furthermore, in the step (3), while dynamic sharding is being performed, new nodes must be added to the network and nodes must be removed from the network.

[0034] Furthermore, in the step (4), when the network is initialized, g nodes are randomly selected or designated according to actual application as proxy nodes for each shard;

[0035] After the time interval set by the network ends, the shard proxy node and the global proxy node are selected through the node reputation mechanism.

[0036] Compared with the prior art, the present invention is beneficial in that:

[0037] This paper proposes the KBFT algorithm, which is a new consensus protocol that can be used for large-scale alliance chains. For the first time, the consensus algorithm is combined with the K-prototype clustering algorithm. While significantly reducing the communication complexity and improving the system throughput, an efficient and fast reputation mechanism and supervision mechanism are set up to ensure the security and reliability of the system. The details are as follows:

[0038] (1) The present invention first applies the K-prototype clustering algorithm to node sharding in the consortium chain. After the end of the period set by the system, the K-prototype clustering algorithm will be used to re-shard the nodes in the network to prevent behaviors such as nodes colluding to do evil, thus ensuring the security of the system. While re-sharding, new nodes can choose to join the network. The addition of new nodes can not only expand the scale of the network, but also reduce the proportion of Byzantine nodes in the system.

[0039] (2) The present invention uses a consensus algorithm that combines BLS aggregate signature and Byzantine fault tolerance algorithm to perform independent and parallel consensus on transactions within each shard, thereby significantly reducing the communication complexity in the network and linearly increasing the throughput with the number of shards.

[0040] (3) The present invention sets up a simple, efficient credit mechanism and supervision mechanism for the algorithm to score and supervise the behaviors of nodes, thereby further ensuring the security of the system. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1 is a single mesh network topology;

[0042] Figure 2 is multiple mesh network topologies;

[0043] Figure 3 is the overall system model;

[0044] Figure 4 is the interaction process of the KBFT algorithm consistency protocol;

[0045] Figure 5 is the view switching flowchart;

[0046] Figure 6 is the in-shard consensus failure rate under different shard sizes;

[0047] Figure 7 is the analysis result of the relationship between the success rate and P F under different m, the same K and Ν;

[0048] Figure 8 is the two-dimensional diagram of the communication ratio;

[0049] Figure 9 is the three-dimensional diagram of the communication ratio;

[0050] Figure 10 is the simulated consensus result. DETAILED DESCRIPTION OF THE INVENTION

[0051] To further elaborate on a clustering-based scalable sharding consensus algorithm of the present invention and achieve the intended invention purpose, the following, in conjunction with preferred embodiments, details the specific implementation manner, structure, features, and effects of a clustering-based scalable sharding consensus algorithm proposed according to the present invention. In the following description, different "one embodiment" or "embodiments" do not necessarily refer to the same embodiment. In addition, specific features, structures, or characteristics in one or more embodiments can be combined in any suitable form.

[0052] The following will further introduce in detail a clustering-based scalable sharding consensus algorithm of the present invention in conjunction with specific embodiments:

[0053] Practical Byzantine Fault Tolerance (PBFT) is an improved consensus protocol proposed based on the original Byzantine Fault Tolerance algorithm BFT. PBFT solves the problem of low efficiency of the BFT algorithm, and the communication complexity is reduced from exponential level to polynomial level. However, PBFT still has many problems. For example, when the number of nodes is too large, it will cause a huge overhead in network communication, resulting in network congestion. Generally, the number of nodes in the network does not exceed 100. When the primary node fails, a complex view switching protocol will be started to ensure the liveness of the system, so the efficiency of the system will be significantly reduced. There is no certain punishment mechanism for malicious nodes, which will make malicious nodes always exist in the network and continue to perform malicious behaviors.

[0054] In order to improve the throughput of the system, the blockchain in the technical solution of the present invention adopts new technologies such as sharding technology. Sharding technology was first used for database partitioning, dividing a large database into smaller, faster, and easier-to-manage parts. Applying sharding technology to the blockchain, in terms of sharding strategies, sharding technology can be divided into three types: network sharding, transaction sharding, and state sharding. The Elastico protocol proposed by some scholars is the first protocol in blockchain consensus algorithms to use sharding technology. However, since this protocol is designed for public blockchains and requires economic incentives to encourage nodes to verify, it is not suitable for consortium blockchains.

[0055] Unsupervised learning is a popular machine learning technique that has been widely applied in many fields. For example, it can be used to discover abnormal data in a large amount of big data. Many advertising platforms use this technology for user segmentation and various recommendation systems. The most common scenarios of unsupervised learning are clustering and dimensionality reduction. Common clustering algorithms include K-means, DBSCAN, hierarchical clustering, etc. However, these algorithms can only process numerical data, while K-modes only processes categorical attribute data. In engineering, the data processed in this invention mostly includes both numerical data and categorical data. Therefore, a clustering method that can handle two different types of data simultaneously is needed, and K-prototypes is such a method. In this invention, the K-prototypes clustering algorithm will be used to partition nodes according to the mixed attributes of the nodes in the network.

[0056] The differences between the method proposed in this invention and previous work are summarized as follows:

[0057] ① The KBFT algorithm can complete the partitioning work quickly and efficiently through the K-prototype clustering algorithm.

[0058] ② Compared with previous work, the KBFT algorithm designs a more concise and efficient credit mechanism and supervision mechanism on the basis of partitioning, making the network more secure and reliable.

[0059] ③ This invention analyzes in detail the impact of the selection of partition size on security.

[0060] Aiming at the problems of communication blockage and low throughput caused by the increase in the number of nodes in the classical Practical Byzantine Fault Tolerance (PBFT) algorithm mainly used in consortium blockchains, this invention proposes a scalable sharding consensus new algorithm based on clustering (KBFT). This algorithm first uses the K-prototype clustering algorithm to partition the nodes in the network according to their mixed attributes, and then non-overlapping transactions are consensus in parallel in different shards. At the same time, the KBFT algorithm introduces a supervision mechanism and a node reputation mechanism to supervise and score the behavior of nodes and select proxy nodes, improving security. This invention discusses the selection of shard size with the binomial distribution in probability theory, and analyzes the probability that the system can successfully form a global block under different node failure probabilities. Finally, this algorithm is evaluated through theoretical analysis and simulation experiments. The results show that compared with the classical baseline algorithm PBFT in this field, the proposed algorithm in this paper has significantly improved scalability and throughput, and significantly reduced communication complexity, ensuring the operation efficiency of the system while also guaranteeing the security and robustness of the system.

[0061] Embodiment 1.

[0062] The system model of Algorithm A

[0063] (1) System model

[0064] Improve the single mesh network topology into multiple mesh network topologies. The structure is as Figure 1-2 shown, and formalize the model as: The target system consists of a client, N consensus nodes and a supervision node N supervise which are composed. And divide the N nodes into different shards according to the numerical attributes and classification attributes of the nodes by the K-prototype clustering algorithm, denoted by S. Let S = {S1, S2,... S m} represent being divided into m shards, where S m represents the m-th shard. Let T = {K1, K2,..., K j} represent the number of nodes in each shard. Therefore, N = K1 + K2 +... + K j . Select a node as the proxy node in each shard. The role it acts as is similar to the primary node in the PBFT algorithm. However, the algorithm selects a proxy node for each shard, which can greatly reduce the load of a single primary node and prevent the computing power of the primary node from becoming the bottleneck affecting the performance of the blockchain. The proxy nodes in each shard form a proxy committee, denoted by A = {a1, a2,..., a m}.

[0065] The supervision node N set in the KBFT algorithm supervise is the node that performs supervision during the consensus process. For example, whether the consensus nodes actively participate in the consensus and whether there are malicious information modification behaviors, etc. N supervise will store a list of all consensus nodes in the network locally to manage the information of all nodes, including the public keys, IDs, IP addresses, credits, etc. of the nodes. The consensus nodes in the network fully trust the information recorded by N supervise . To maximize the maintenance of the distributed characteristics of the network and the decentralized characteristics of the blockchain, N supervise will not participate in the consensus process in the KBFT consensus algorithm and only serves as an information storage node.

[0066] The client that initiates the transaction signs the transaction and sends it to the proxy node of the shard it belongs to. Inside the shard, the consensus nodes verify this transaction through a Byzantine fault tolerance algorithm based on BLS aggregate signature. After reaching a consensus, the block corresponding to this transaction will be stored locally on the node. However, the design of the KBFT algorithm does not make each shard only retain its own block data. Instead, the global proxy node merges the blocks generated by each shard within the specified time and broadcasts them to all nodes in the network, so that the data saved by the nodes in the entire network is comprehensive. Because losing control of any shard will completely interrupt the blockchain, and the comprehensiveness of the data saved by all nodes in the network also conforms to the original intention of the decentralization of the blockchain. The overall system model is as Figure 3 shown.

[0067] (2) Security Assumptions and Threat Models

[0068] ① Divide the nodes in the network into different shards. The nodes inside each shard are connected to each other through the network and know their basic information, such as IP addresses and public keys. Each shard is disjoint and the nodes in different shards cannot communicate. Assume that the network connection between the nodes within the shard is stable and the network has weak synchronization characteristics. Once a node broadcasts any message, other nodes will arrive within δ t this bounded time delay.

[0069] ② Assume that the computing power of the adversary is limited and cannot break the encryption technology or indefinitely delay the network. Therefore, digital signatures or other encryption technologies can be used to ensure the correctness of the messages. The clients in this model are fault-free, which can be guaranteed through client authentication.

[0070] ③ Stipulate that the nodes joining the network must go through a strict identity admission mechanism to prevent Sybil attacks.

[0071] ④ It is considered that the proposed model is vulnerable to possible attacks that may seriously interfere with the normal operation of the system. The present invention assumes that the probability of each node becoming a Byzantine node is PF and the number of Byzantine nodes in each shard is f, and assumes that the number of nodes in each shard is 3f + 1, that is, the number of Byzantine nodes that can be tolerated in each shard is one-third of the number of nodes in the shard. The behavior of malicious Byzantine nodes may be arbitrary. For example, they may refuse to participate in the consensus, collude with other malicious nodes to attack the system or tamper with information, etc. However, the correct nodes will always follow the algorithm requirements.

[0072] a) If the structure of sharding in the network is fixed, malicious attackers may carry out static cyclic attacks, slow adaptation attacks, etc. on the nodes in the network. Therefore, the KBFT algorithm sets a constant time interval, and after the set time interval ends, the nodes in the network will be "reshuffled" to deal with these attacks.

[0073] B: Design of KBFT Algorithm

[0074] 1) Sharding of Nodes

[0075] The KBFT algorithm first uses the K-prototype clustering algorithm to classify nodes according to their numerical attributes and categorical attributes. The numerical attributes include the ID of the node, the credit value, etc., while the categorical attributes include the company to which the node belongs, the IP address of the node, the geographical location, etc.

[0076] Let the node dataset X = {X1, X2, X3,..., X n}, where n is the number of node objects in the dataset X, and each node data in the node dataset has m attributes, that is, X i = {x i1 , x i2 , x i3 ,..., x ip , x i(p+1) , x i(p+2) , x im}, where the numerical attributes are in the front, a total of p, and the categorical attributes are in the back, a total of m - p. Given a positive integer g, the node dataset X is divided into g shards. The sharding steps are as follows:

[0077] (1) Randomly select g nodes as the initial prototypes (center points), and in practical applications, the initial prototypes and the size of g need to be specified according to the actual application requirements.

[0078] (2) Allocate node objects to the nearest cluster according to the dissimilarity, and update the prototype of the cluster after allocation. The dissimilarity formula is as follows:

[0079]

[0080] where The first term of Equation (1) is the Euclidean squared distance of the numerical attributes, and the second term is the simple matching dissimilarity on the categorical attributes. The parameter γ is the weight of the categorical attributes, and adjusting this parameter can avoid the result being biased towards the numerical attributes or the categorical attributes.

[0081] (3) After the classification of nodes is completed, re-determine the prototypes of the classes. The mean value of the variable sample values of the numerical type is used as the characteristic value of the new prototype, and the mode of the variable sample values of the categorical type is used as the characteristic value of the new prototype.

[0082] (4) Repeat steps (2)-(3) until no node sample changes category, and return the final sharding result.

[0083] 2) KBF algorithm consistency protocol process

[0084] The KBFT algorithm includes the consensus process of each shard processing the corresponding transaction and the process of merging and distributing block header data. The consensus process comes first, and the merging and distribution process comes later. The basic process of the algorithm is as follows: Figure 4 As shown:

[0085] The specific algorithm steps are as follows:

[0086] ① Request phase: The client sends the request message to the proxy node of the shard where it is located, and the consensus on the request will be carried out within the shard.

[0087] ②Preparation phase: The proxy node in the shard builds a new block and broadcasts the block to the rest of the consensus nodes in the shard.

[0088] ③ Preparation Phase 1: The node verifies the block. If the verification is valid, it is signed using the BLS algorithm and the signature is fed back to the proxy node.

[0089] ④ Preparation phase 2: The proxy node waits for and collects valid signatures from other consensus nodes. After receiving at least 2f+1 identical signature messages (including itself), it aggregates all individual signatures into a BLS multi-signature, and then broadcasts this multi-signature to all nodes. At this point, the preparation phase ends.

[0090] ⑤ Commitment Phase 1: The node will verify whether the received multi-signature contains at least 2f+1 signers, verify the transactions in the block broadcast by the proxy node in the pre-preparation phase, sign the message received in preparation phase 2, and send it to the proxy node.

[0091] ⑥Submission phase 2: The proxy node waits for and collects at least 2f+1 valid signatures, aggregates these signatures together to form a BLS aggregate signature, and submits a new block with this aggregate signature. The new block is then broadcast to all nodes for verification and submission. At this point, the submission phase ends.

[0092] ⑦ Reply phase: After the node submission is completed, a reply message will be sent to the client node. When the client receives at least f+1 identical confirmation messages sent by different nodes, it indicates that the current request has reached the final consensus.

[0093] ⑧Feedback stage: When the client receives the reply message in the shard, it will feed back all the reply messages to the supervisory node, and the node will then score the node's behavior based on the client's feedback results.

[0094] ⑨Merge and distribute phase: The proxy nodes of each shard will broadcast the local blocks they have formed to the global proxy node. The global proxy node will merge all the received blocks into a global block and distribute it to each shard proxy node. The proxy node that receives the global block will send it to the node in the shard where it is located. At this point, the consistency protocol process ends.

[0095] 3) Dynamic resharding mechanism

[0096] The KBFT algorithm will use a dynamic resharding mechanism to deal with the attacks on static sharding proposed in the threat model. After the time interval set by the network ends, the K-prototype clustering algorithm will be used to reshard all nodes in the network. After each cycle, the number of nodes in the network, as well as the numerical attributes and classification attributes contained in the same node will be different from those contained before. This ensures that after the K-prototype clustering algorithm is used to shard the nodes, the sharding results are different from the previous results, thereby playing a role in preventing the system from being attacked by static sharding through resharding.

[0097] 4) Selection of shard proxy nodes and global proxy nodes

[0098] When the network is initialized, g nodes will be randomly selected or designated as proxy nodes for each shard according to the actual application. After the set period ends, all nodes in the network will be re-sharded according to the information in the node information list stored by the supervisory node, and the selection of proxy nodes will no longer be random but based on the level of trust in the shard. Nodes with high trust serving as proxy nodes will make the system more secure and the consensus process more robust. The selection of global proxy nodes is similar to the selection of proxy nodes within the shard, that is, after selecting the proxy nodes in each shard, the node with the highest credit will be selected from the selected proxy committee or the global proxy node will be selected according to the actual application.

[0099] 5) New nodes joining and nodes leaving

[0100] While performing dynamic sharding, the number of nodes in the network can be increased or decreased. New nodes can join the network, and nodes can also choose to exit the network at this time. The addition of new nodes will not only improve the scalability of the network, but the increase in the number of nodes in the network will also enhance the robustness of the network. The exit of nodes includes the active exit of nodes and the removal of nodes from the network as a penalty due to malicious behavior in the previous consensus. If malicious nodes continue to exist in the network without being dealt with, then the malicious nodes may gradually erode other good nodes, thereby affecting the security of the network.

[0101] 6) Reputation Model

[0102] The reputation model designed by the KBFT algorithm has three functions. First, the reputation score adds an attribute to each node, so that it can be used when using the K-prototype clustering algorithm for sharding. Secondly, the reputation value of each node will be changed after each consensus and saved by the supervisory node. When re-sharding, you only need to quickly select the proxy node in the shard and the global proxy node based on the reputation value; finally, the reputation can ensure that the node works better and more honestly, thereby reducing the possibility of malicious nodes doing evil. In order to improve the efficiency of the system, there is no complex reputation formula in the KBFT reputation model, but a reputation value is directly increased or decreased according to the behavior of the node. When the honest node's reputation value is increased by 1 after each consensus, the reputation value of the node that does not participate in the consensus is reduced by 1, and the reputation value of the node that has malicious behavior will be directly reduced to 0, and it will no longer participate in the relevant consensus before the global sharding. When a node is found to have committed evil many times, the node will be removed from the network and prohibited from joining the network.

[0103] 7) View switching

[0104] The above describes the normal operation of the KBFT algorithm. However, some failures may occur during the operation of the network. The present invention mainly considers the malicious behavior of proxy nodes in the failure of the shard. Because the consensus in the shard is based on the PBFT algorithm, the algorithm can tolerate no more than one-third of the nodes being malicious at the same time. Therefore, it is more important to consider the situation of proxy nodes being malicious. The failure or malicious behavior may be as follows:

[0105] ① The proxy node within the shard does not broadcast the block information of the new transaction to other consensus nodes within the set time.

[0106] ② After the consensus within the shard is completed, the proxy node does not send the block header to the global proxy node, resulting in the global proxy node not receiving the block header of the legal block number within the set time.

[0107] ③ The global proxy node fails to complete the work of receiving all the blocks in the network and performing merging and distribution within the set time.

[0108] For the above possible situations, the algorithm of the present invention sets corresponding countermeasures to reasonably handle these situations, so as to ensure the security and vitality of the system.

[0109] If a proxy node within a shard fails, the remaining correct nodes within the shard will select a new proxy node by running a local view change and then continue to run the consensus. Similarly, if the global proxy node receives fewer block headers than the minimum threshold set by the network within the specified time, an emergency re-sharding mechanism will be triggered at this time.

[0110] View switching within a shard: The view switching protocol within a shard in the KBFT algorithm is not as complex as the view switching protocol in PBFT. When the consensus within a shard does not complete within the specified time or a replica node does not receive a message sent by the proxy node, the view switching protocol within the shard will be triggered at this time. The specific process is that the replica nodes within this shard only need to send a message requesting view switching to N supervise If N supervise receives the same view switching messages sent by more than half of the nodes within this shard, it will broadcast the node with the highest trust value within this shard to the nodes and clients within this shard, and the shard consensus will start again. The view switching flowchart is as follows Figure 5 shown:

[0111] Global view switching: The global view is also called the emergency re-sharding mechanism, which is different from the re-sharding after the end of the period set by the network. The main difference between the two is the triggering timing. Emergency re-sharding is triggered when the global proxy node receives fewer block headers than the minimum threshold set by the network within the specified time or the global proxy node does not distribute the merged blocks. The sharding methods of the two are the same, that is, they will both perform re-sharding through the K-prototype clustering algorithm to ensure that the sharding results are different from the previous ones. However, in the emergency re-sharding mechanism, if the selected global proxy node after re-sharding is still the same as the previous one, then the proxy node with the second highest trust value in the proxy node set will be appointed as the global proxy node.

[0112] C Security Analysis

[0113] (1) Selection of shard size

[0114] The size of the shard is very important to the security of the system. The Byzantine fault-tolerant consensus algorithm based on BLS aggregate signature is used for transactions generated by the client in each shard, so the number of Byzantine nodes that can be tolerated in each shard does not exceed one-third of the total number of nodes in the shard. The present invention makes the following two assumptions:

[0115] Assumption (1): The number of nodes in each shard is K (K=3f+1, f=1, 2, 3, ...).

[0116] Assumption (2): The probability of each node failing is P F .

[0117] Based on the above assumptions (1) and (2), the probability formula for each shard failing to successfully reach consensus can be obtained as shown in the following formula (2):

[0118]

[0119] The function graph of formula (2) is visualized as Figure 6 shown.

[0120] from Figure 6 It can be seen that when the probability of failure of each node is constant, the consensus failure rate gradually decreases as the shard size increases. When the number of nodes in the shard increases to 88, the probability of containing one-third of malicious nodes is one in a thousand, so the possibility of consensus failure is completely negligible.

[0121] (2) Probability analysis of successful formation of global blocks

[0122] In KBFT, the present invention sets a minimum merge block threshold, which is the minimum sum of the number of blocks sent by the proxy nodes in each shard received by the global proxy node. A complete consensus process is divided into two steps, the first step is the consensus within the shard, and the second step is the merging of blocks. Since the nodes in the network are divided into N / K shards, the present invention needs to analyze the threshold required to reach consensus so that the entire system maintains security and activity.

[0123] The present invention makes the following two assumptions about two events A and B:

[0124] Event A: The number of failed nodes in each shard does not exceed 1 / 3, that is, consensus is successfully reached within the shard, and the probability of failure of each node is also P F .

[0125] Event B: The number of blocks finally merged by the global proxy node is greater than or equal to the set minimum threshold m, that is, the global block is successfully formed.

[0126] Therefore, the probabilities of events A and B can be obtained according to the binomial distribution as follows:

[0127]

[0128]

[0129] Substituting formula (2) into formula (3), the probability of the successful occurrence of the complete event B can be obtained as follows:

[0130]

[0131] To more intuitively observe the relationship between the success rate of successfully forming a global block and PF, the present invention visualizes the function image of formula (5) under different m, the same K and N, as Figure 7 shown.

[0132] From Figure 7 it can be seen that when the number of nodes in the network is the same as the number of nodes in the shard, as the minimum merge block threshold decreases, the inflection point in the probability curve gradually delays. At the same time, it can also be seen that the smaller the probability of each node failing, the higher the success rate. When the node failure probability can be controlled to be less than or equal to 0.2, the KBFT algorithm can ensure a success rate of 100%.

[0133] D Analysis of the KBFT Algorithm

[0134] A comparative analysis is made with the classical PBFT algorithm in terms of communication overhead, throughput, system reliability and robustness, etc., and formula (4) derived in the fifth part is verified by simulating the consensus process.

[0135] (1) Communication Overhead

[0136] Assume that the total number of nodes in the networks of the KBFT and PBFT algorithms is both N, and for the sake of generality, assume that the number of nodes in each shard of the KBFT is K, then the network is divided into N / K shards. The calculated communication overhead of the PBFT and Figure 4 The consistency protocol interaction process of the KBFT algorithm can lead to the following two conclusions:

[0137] Conclusion 1: The communication complexity required for the PBFT algorithm to complete one consensus is Ο(N 2 ), and the specific communication overhead is:

[0138] C PBFT = 2N 2 (6)

[0139] Conclusion 2: The KBFT algorithm divides nodes into shards, and each shard independently conducts Byzantine fault-tolerant consensus based on BLS aggregate signatures. Therefore, the communication complexity of completing one consensus is Ο(Ν), and the specific communication overhead is as follows:

[0140]

[0141] Let the ratio of the communication overheads of the two algorithms be J. From formulas (6) and (7), formula (8) can be obtained as follows:

[0142]

[0143] Figure 8 and 10 are the two-dimensional and three-dimensional function graphs of J respectively, where Figure 8 K takes the values of 4, 7, 10, 13, and 88 respectively. Figure 9 The value range of K in

[0144] is the interval [4, 88]. Figure 8 According to

[0145] it can be seen that when the number of nodes in the network is the same, as the number of nodes within the shard decreases, the communication ratio between the two will become smaller and smaller, indicating that the smaller the number of nodes within the shard, the better the KBFT algorithm can reduce the communication overhead. Secondly, when the shard size is fixed, as the number of network nodes increases, the communication overhead ratio of the two algorithms will gradually decrease, indicating that the larger the scale of the network, the more advantageous the KBFT algorithm is in terms of node communication resources. Figure 9 According to

[0146] it can be seen that when the total number of nodes in the network using the PBFT algorithm and the KBFT algorithm is the same and very large, even if k changes, the communication overhead ratio between the two remains almost horizontal. According to formula (8), it can be explained that when N takes a large value, the value of the denominator will be much larger than the value of the numerator. Therefore, when N is fixed, the change of K has almost no impact on J.

[0147] (2) Throughput

[0148]

[0149] where T x represents Δ time the total number of transactions packed into the block, and Δ timeIt represents the time interval from the completion of transaction creation to the transaction being uploaded to the blockchain. This invention assumes that in a system using the PBFT algorithm and a system using the KBFT algorithm within a unit time, the total transaction volume generated is the same. At the same time, it is assumed that the number of nodes participating in consensus in both systems is the same, and the number of transactions contained in a block is the same. For the former, the processing of transactions will reach an agreement through consensus by all nodes, that is, the transactions consensus by all nodes in the entire network are the same. While for the latter, by sharding the transactions, the transactions processed by each shard are different. Therefore, the throughput of the system using the KBFT algorithm will be N / K times that of the system using the PBFT algorithm, where N / K is the number of shards in the network. And because KBFT uses network sharding and transaction sharding, the total throughput increases linearly with the increase in the number of shards.

[0150] (3) Reliability and robustness

[0151] The proxy nodes in each shard of the KBFT algorithm and the global proxy nodes are selected according to the credit value. Nodes with a high credit value can ensure better leadership in leading other consensus nodes within the shard to complete consensus, effectively promoting the operation efficiency of the system and ensuring the security and reliability of the entire system. At the same time, by scoring nodes according to their consensus behavior, it can encourage honest nodes to work better, while malicious nodes will have their credit values reduced or even be excluded from the network due to malicious behavior. However, in the PBFT algorithm, the main node is selected in a sequential rotation manner, and it is difficult to ensure that the main node is not a Byzantine node. If the main node is a Byzantine node, then a complex view switch will be carried out to re-consensus, resulting in a significant reduction in efficiency. Moreover, the PBFT algorithm does not impose any punishment on the malicious behavior of Byzantine nodes. Therefore, it may lead to continuous malicious behavior of Byzantine nodes without paying any price.

[0152] At the same time, the design of the KBFT algorithm is not specifically for a blockchain network with a fixed number of nodes. When the network is re-sharded, new nodes can choose to join the network. The increase in the number of nodes in the network will reduce the proportion of malicious nodes, thereby ensuring the robustness of the system and also enhancing the security of the system.

[0153] (4) Simulated consensus

[0154] A simulation experiment of the KBFT algorithm was carried out using Python. In the simulated consensus, the failure probability P of a single node F was used as the independent variable, and the success rate P(B) was used as the dependent variable. At the same time, the total number of network nodes N was specified as 880, the number of nodes within a shard K was 88, the minimum thresholds m were 7, 8, 9, 10 respectively, and the conclusion in the fifth part that when the number of nodes within a shard is 88, the failure rate of consensus within the shard will drop to one-thousandth was utilized.

[0155] The results of repeating the simulation experiment 1000 times are as follows Figure 10 as shown

[0156] Through Figure 10 the present invention, it can be seen that the curve of the simulation consensus experiment repeated 1000 times is highly fitted with the curve obtained according to formula (5).

[0157] In view of the fact that the PBFT algorithm and the existing new algorithms improved based on PBFT cannot be applied to the current situation of large-scale consortium chains, the present invention proposes a new KBFT algorithm. This algorithm quickly shards the nodes in the network through the K-prototype clustering algorithm, and a consensus algorithm combining BLS aggregate signature and Byzantine fault tolerance algorithm is used within the shards, which significantly improves the scalability of the network and greatly reduces the communication overhead. At the same time, through transaction sharding, the throughput increases linearly with the increase in the number of shards. The KBFT algorithm also sets up an efficient and simple credit mechanism and supervision mechanism to ensure the security and reliability of the system. For some application scenarios with high requirements for scalability and security, such as financial services, energy trading, Internet of Things, etc., the KBFT algorithm is a good solution.

[0158] The above are only the preferred embodiments of the embodiments of the present invention, and do not impose any formal restrictions on the embodiments of the present invention. Any simple modifications, equivalent changes and modifications made to the above embodiments based on the technical essence of the embodiments of the present invention still fall within the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A scalable sharding consensus algorithm based on clustering, characterized in that The following steps are involved: (1) Use the K-prototype clustering algorithm to shard nodes in the network according to mixed attributes; (2) Processing the corresponding transactions of each shard to carry out consensus process, merging and distribution process; The specific steps of the consensus process, merging and distribution process are as follows: ① Request phase: The client sends the request message to the proxy node of the shard where it is located, and the consensus on the request will be carried out within the shard; ② Pre-preparation phase: The proxy node in the shard builds a new block and broadcasts the block to the rest of the consensus nodes in the shard; ③ Preparation phase 1: The node verifies the block. If the verification is valid, it uses the BLS algorithm to sign it and feeds the signature back to the proxy node; ④ Preparation phase 2: The proxy node waits for and collects valid signatures from other consensus nodes. When it receives at least 2f+1 identical signature messages, it aggregates all individual signatures into a BLS multi-signature, and then broadcasts the BLS multi-signature to all nodes. At this point, the preparation phase ends; the f is the number of Byzantine nodes in each shard; ⑤ Commitment Phase 1: The node will verify whether the received multi-signature contains at least 2f+1 signers, verify the transactions in the block broadcast by the proxy node in the pre-preparation phase, sign the message received in the preparation phase 2, and send it to the proxy node; ⑥Submission phase 2: The proxy node waits for and collects at least 2f+1 valid signatures, aggregates these signatures together to form a BLS aggregate signature, and submits a new block with the BLS aggregate signature, and then broadcasts the new block to all nodes for verification and submission. At this point, the submission phase ends; ⑦ Reply phase: After the node submission is completed, a reply message will be sent to the client node. When the client receives at least f+1 identical confirmation messages sent by different nodes, it indicates that the current request has reached the final consensus; ⑧ Feedback stage: After the client receives the reply message in the shard, it feeds back all the reply messages to the supervisory node, and the supervisory node scores the behavior of the final consensus node according to the feedback result of the client; ⑨Merge and distribute phase: The proxy nodes of each shard will broadcast the local blocks to the global proxy node. The global proxy node will merge all the received blocks into a global block and distribute it to each shard proxy node. The proxy node that receives the global block will send it to the node in the shard where it is located. (3) Dynamic resharding: After the time interval set by the network ends, the K-prototype clustering algorithm reshards all nodes in the network; (4) Selecting shard proxy nodes and global proxy nodes, the steps are as follows: when the network is initialized, randomly select or specify g nodes as proxy nodes for each shard according to actual application; After the time interval set by the network ends, the shard proxy node and the global proxy node are selected through the node reputation mechanism.

2. The sharding consensus algorithm according to claim 1, wherein: In step (1), classification is performed according to the numerical attributes and classification attributes of the nodes.

3. The sharding consensus algorithm according to claim 2, wherein: In step (1), the classification steps are as follows: a: Randomly select g nodes as the initial prototypes; b: Allocate the node objects to the nearest cluster according to the dissimilarity, and update the prototype of the cluster after allocation; c: Re-determine the prototype of the category; d: Repeat steps b-c until no node samples change their categories, and return the final sharding result.

4. The sharding consensus algorithm according to claim 3, wherein: In the said step (1), the node dataset X = {X1, X2, X3,..., Xn}, where n is the number of node objects in the dataset X, and each node data in the node dataset has m attributes, that is, X i = {x i1 , x i2 , x i3 ,..., x ip , x i(p+1) , x i(p+2) , x im}, where there are p numerical attributes in the front and m - p categorical attributes in the back; The dissimilarity formula is as follows: In the formula, the parameter γ is the classification attribute weight.

5. The sharding consensus algorithm according to claim 1, wherein: In step (3), while performing dynamic sharding, new nodes should be added to and nodes should exit from the network.

Citation Information

Patent Citations

  • Block chain adaptive consensus method based on dynamic authorization and network environment perception

    CN108616596A

  • Large-scale node efficient consensus method

    CN114928446A