Reliable broadcasting method with load balancing strategy for DAG consensus
By introducing load balancing strategies and erasable code sharding technology, broadcast tasks are dynamically scheduled, solving the bottleneck of the blockchain system caused by heterogeneous nodes, improving system performance and resource utilization, and reducing upgrade costs.
Patent Information
- Application Number
- CN202511018067.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-23
- Publication Date
- 2025-10-28
AI Technical Summary
In existing blockchain systems, the performance differences between heterogeneous nodes cause bottlenecks, especially when broadcasting large batches of transactions, where some nodes lack sufficient processing capacity, affecting the overall throughput and efficiency of the cluster.
A load balancing strategy is introduced, which monitors the node load status in real time through a reinforcement learning decision engine and a deep learning prediction module. Erasable code sharding technology and VID server are adopted to dynamically schedule broadcast tasks to lightly loaded nodes, thereby achieving dynamic optimal matching of transaction distribution.
It alleviates the performance bottleneck of heterogeneous nodes, improves resource utilization and system elasticity, is compatible with existing consensus protocols, and reduces upgrade costs.
Smart Images

Figure CN120856702A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of blockchain technology, specifically to a reliable broadcasting method with a load balancing strategy for DAG consensus. Background Technology
[0002] In recent years, blockchain technology, with its core advantages such as decentralization, immutability, transparency, and traceability, has attracted widespread attention from academia and industry. It has powerfully promoted the development of cutting-edge fields such as digital currency, smart contracts, metaverse, and digital collectibles, and has been widely applied in areas such as information security, finance and securities, digital rights confirmation, and traceability, achieving distributed storage and efficient utilization of data. As a distributed ledger technology that ensures the trustworthiness of on-chain data, blockchain has injected new vitality into the traditional internet and demonstrated broad application prospects.
[0003] Based on their openness, blockchains are mainly divided into public blockchains, consortium blockchains, and private blockchains. Consortium blockchains, through node authorization mechanisms, build a platform for specific groups that supports multi-party collaborative maintenance and controllable data access permissions, effectively preventing data leakage. In consortium blockchains, the Byzantine Fault-Tolerant (BFT) consensus algorithm is widely used to ensure that the system can still achieve state consistency even with malicious or faulty nodes. To adapt to complex network environments, researchers have proposed various asynchronous Byzantine Fault-Tolerant (ABFT) consensus algorithms, which possess higher fault tolerance and flexibility, ensuring security and liveness in asynchronous networks. In particular, asynchronous consensus algorithms based on Directed Acyclic Graphs (DAGs), such as Narwhal-Tusk, demonstrate outstanding performance in handling high-concurrency transactions by decoupling the transaction pool from the consensus process and optimizing it.
[0004] Research on improving blockchain consensus performance often focuses on optimizing the consensus protocol itself. For example, current state-of-the-art asynchronous consensus protocols like Bullshark and Narwhal generally focus their improvements on the consensus layer, such as designing fast commit channels and implementing dynamic switching between semi-synchronous and asynchronous modes. However, in actual engineering deployments, these protocols often face unexpected challenges. A typical problem is that protocol design is usually based on the assumption of homogeneous node performance, while real-world clusters often contain heterogeneous nodes (such as high-load nodes with extremely high connection counts). In this scenario, a single low-performance node can become a system bottleneck, severely restricting the throughput and efficiency of the entire cluster.
[0005] For example, when Narwal is used for reliable broadcasting, the Narwal protocol broadcasts the entire batch of transactions to all nodes upon receiving it. This approach ignores the real-time load differences among nodes, and especially when broadcasting large batches, the reliable broadcast process may be blocked due to insufficient processing capacity of some nodes. Summary of the Invention
[0006] The purpose of this invention is to address the performance bottleneck problem caused by heterogeneous nodes mentioned in the background art by proposing a reliable broadcast with a load balancing strategy for DAG consensus.
[0007] To achieve the above objectives, the present invention provides the following technical solution:
[0008] A reliable broadcasting method with a load balancing strategy for DAG consensus includes the following steps:
[0009] Step 1: Deploy the environment
[0010] In a blockchain cluster consisting of N heterogeneous nodes, each heterogeneous node is equipped with a load balancing decision engine and a VID server.
[0011] Step Two: Transaction Batch Receiving and Forecasting
[0012] When a node receives a batch of transactions to be broadcast, it calls the LSTM prediction module, inputs the most recent round of global state data, and outputs the expected transaction volume for the future time window.
[0013] Step 3: Reinforcement Learning Decision Distribution
[0014] The load balancing decision engine outputs an on-demand distribution ratio vector;
[0015] Step 4: Reliable Broadcast
[0016] Erasable code fragmentation technology is used to divide the original data into k data blocks, and m fragments are generated through encoding, so that the original data can be completely reconstructed with just any k fragment certificates.
[0017] In a preferred embodiment, the load balancing decision engine in step one includes a reinforcement learning decision engine, a global state collector, and a deep learning prediction module. The reinforcement learning decision engine comprises the following three elements:
[0018] Status: Load balancing metrics include CPU utilization, memory usage, minimum number of connections, and network queue depth;
[0019] Action: Defined as a proportional vector for allocating transactions to N nodes. The action space must satisfy the constraint that the sum of the allocation proportions of each node is 1.
[0020] Rewards: Confirmed transaction throughput is used as the reward signal, which directly reflects the impact of the load balancing decision engine on the overall system performance.
[0021] In a preferred embodiment, the global state collector is responsible for collecting key performance indicators of cluster nodes in real time, including CPU utilization, memory usage, and network bandwidth.
[0022] In a preferred embodiment, the deep learning prediction module uses a lightweight LSTM network model to predict transaction volume, incorporating real historical transaction data from the public blockchain to train a time-series model that can predict future transaction loads.
[0023] In a preferred embodiment, the VID server includes a VID client and a VID server. The VID client is responsible for integrating with the reliable broadcast module and calling its block distribution and reconstruction capabilities to send or retrieve block data to the VID server. The VID server is responsible for the fragmented storage and management of block data and supports reliable access to blocks by the client.
[0024] In a preferred embodiment, step three includes the following steps:
[0025] S1: Status acquisition, the global status collector collects real-time metrics from all nodes;
[0026] S2: Action generation, the reinforcement learning decision engine takes the current state and expected transaction volume as input, and outputs a distribution ratio vector;
[0027] S3: Distribute and execute transactions. Each node reliably broadcasts a number of transactions, and the remaining transactions are split according to the distribution ratio vector mentioned above and delegated to other nodes in a peer-to-peer manner.
[0028] In a preferred embodiment, during the state acquisition in step S1, under the DAG consensus mechanism, each node maintains the same local graph structure. When a node broadcasts its current round certificate in each round, it attaches its own load information to achieve synchronization of the global load view. If a node is missing the load data for the current round, a load estimation mechanism based on time decay is enabled.
[0029] When a transaction is sent to a malicious node that claims to be under low load, but the system throughput does not improve, the reinforcement learning decision engine can gradually identify nodes with abnormal reporting behavior by observing reward feedback over a long period of time, and optimize the transaction distribution strategy accordingly, thereby reducing the reliance on potential fraudulent nodes.
[0030] In a preferred embodiment, step four includes the following steps:
[0031] a: Sharding encoding and storage: Perform Reed-Solomon encoding on transaction batches, input 2*f+1 data blocks, output 3*f+1 shards;
[0032] b: VID client distribution, sending the shards {shard1,...,shardn} to the VID servers on n nodes;
[0033] c: Normal reliable broadcast process is triggered when a node collects storage certificates for 2*f+1 shards;
[0034] d: The VID client retrieves the data after the normal reliable broadcast process is completed. It requests the fragmented data from the corresponding VID server, reconstructs the original data using Lagrange interpolation, and submits the complete transaction batch to the consensus layer.
[0035] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0036] 1. This invention can alleviate the performance bottleneck of heterogeneous nodes: by introducing a dynamic load-aware broadcast routing mechanism, the node load status (such as CPU utilization and network queue depth) is monitored in real time, and broadcast tasks are automatically scheduled from overloaded nodes to lightly loaded nodes, eliminating the systemic bottleneck caused by single-point performance deficiency, and making the throughput of heterogeneous clusters approach the theoretical maximum value.
[0037] 2. This invention can improve resource utilization and system elasticity: Based on node capability profiles (such as historical throughput and hardware configuration), a load balancing model is constructed to achieve dynamic optimal matching between broadcast tasks and node resources;
[0038] 3. This invention offers seamless integration with existing consensus protocols: By encapsulating load balancing logic through a modular broadcast middleware layer, it provides a standard broadcast interface upwards and adapts downwards to mainstream consensus engines (such as Narwal and Bullshark). Performance improvements can be achieved without modifying the consensus protocol itself, reducing the upgrade costs for enterprise-level blockchains. Attached Figure Description
[0039] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings:
[0040] Figure 1 This invention provides a reliable broadcast process with a load balancing strategy.
[0041] Figure 2 This invention provides a load balancing method according to an embodiment of the present invention. Detailed Implementation
[0042] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0043] Please see Figures 1-2 The present invention provides the following technical solution:
[0044] A reliable broadcasting method with a load balancing strategy for DAG consensus includes the following steps:
[0045] Step 1: Deploy the environment
[0046] In a blockchain cluster consisting of N heterogeneous nodes, each heterogeneous node is equipped with a load balancing decision engine and a VID server.
[0047] The load balancing decision engine includes a reinforcement learning decision engine, a global state collector, and a deep learning prediction module. The reinforcement learning decision engine comprises the following three elements:
[0048] Status: Load balancing metrics include CPU utilization, memory usage, minimum number of connections, and network queue depth;
[0049] Action: Defined as a proportional vector for allocating transactions to N nodes. The action space must satisfy the constraint that the sum of the allocation proportions of each node is 1.
[0050] Rewards: Confirmed transaction throughput is used as the reward signal, which directly reflects the impact of the load balancing decision engine on the overall system performance.
[0051] The global status collector is responsible for collecting key performance indicators of cluster nodes in real time, including CPU utilization, memory usage, network bandwidth, etc.
[0052] The deep learning prediction module uses a lightweight LSTM network model to predict transaction volume. It incorporates real historical transaction data from the public chain to train a time-series model that can predict future transaction load. In this framework, when the VID client makes a request action decision, the number of transactions sent is no longer based on the user's current request count, but is dynamically adjusted according to the future transaction volume predicted by the model.
[0053] The VID server includes a VID client and a VID server. The VID client is responsible for integrating with the reliable broadcast module and calling its block distribution and reconstruction capabilities to send or retrieve block data to the VID server. The VID server is responsible for the fragmented storage and management of block data and supports reliable access to blocks by the client.
[0054] Step Two: Transaction Batch Receiving and Forecasting
[0055] When node i receives a batch of transactions to be broadcast, with a size of S transactions, it calls the LSTM prediction module, inputs the global state data of the most recent T rounds, and outputs the expected transaction volume S′ in the future time window of Δt.
[0056] Step 3: Reinforcement Learning Decision Distribution
[0057] The load balancing decision engine outputs an on-demand distribution ratio vector; specifically, it includes the following steps:
[0058] S1: Status acquisition, the global status collector collects real-time metrics from all nodes;
[0059] S2: Action generation, the reinforcement learning decision engine takes the current state and expected transaction volume S′ as input, and outputs the distribution ratio vector A;
[0060] S3: Distribute and execute transactions. Each node reliably broadcasts A[i]*S transactions, and the remaining transactions are split according to the proportion of A and delegated to other nodes through peer-to-peer (P2P).
[0061] In practical use, if a malicious node sends different load information to the nodes, causing the global load obtained by all nodes to be inconsistent, then during the state collection in step S1, under the DAG consensus mechanism, each node maintains the same local graph structure. When a node broadcasts its current round certificate in each round, it attaches its own load information to achieve synchronization of the global load view. If a node is missing the load data of the current round, then the load estimation mechanism based on time decay is enabled. For example, if the bandwidth of the previous round was 1M, then the current round is estimated to be 0.8M.
[0062] If a malicious node fakes its own low load, causing all nodes to broadcast transactions to it but not execute them, the overall TPS will decrease. Therefore, when a transaction is sent to a malicious node claiming low load, but the system throughput does not improve, the reinforcement learning decision engine can gradually identify nodes with abnormal reporting behavior by observing reward feedback over a long period of time, and optimize the transaction distribution strategy accordingly, reducing the reliance on potential fake nodes.
[0063] Step 4: Reliable Broadcast
[0064] Erasable code fragmentation technology is used to divide the original data into k data blocks, and m fragments are generated through encoding, so that the original data can be completely reconstructed with just any k fragment certificates.
[0065] Step four includes the following steps:
[0066] a: Shard encoding and storage, Reed-Solomon encoding is performed on the transaction batch, with 2*f+1 data blocks as input and 3*f+1 shards as output, 3*f+1=2*f+1+f, and the fault tolerance is f shards lost;
[0067] b: VID client distribution, sending the shards {shard1,...,shardn} to the VID servers on n nodes;
[0068] c: Normal reliable broadcast process is triggered when a node collects storage certificates for 2*f+1 shards;
[0069] d: The VID client retrieves the data after the normal reliable broadcast process is completed. It requests the fragmented data from the corresponding VID server, reconstructs the original data using Lagrange interpolation, and submits the complete transaction batch to the consensus layer.
[0070] Finally, it should be noted that the above descriptions are merely preferred embodiments of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A reliable broadcasting method with a load balancing strategy for DAG consensus, characterized in that: Includes the following steps: Step 1: Deploy the environment In a blockchain cluster consisting of N heterogeneous nodes, each heterogeneous node is equipped with a load balancing decision engine and a VID server. Step Two: Transaction Batch Receiving and Forecasting When a node receives a batch of transactions to be broadcast, it calls the LSTM prediction module, inputs the most recent round of global state data, and outputs the expected transaction volume for the future time window. Step 3: Reinforcement Learning Decision Distribution The load balancing decision engine outputs an on-demand distribution ratio vector; Step 4: Reliable Broadcast Erasable code fragmentation technology is used to divide the original data into k data blocks, and m fragments are generated through encoding, so that the original data can be completely reconstructed with just any k fragment certificates.
2. A reliable broadcast method with load balancing strategy for DAG consensus as described in claim 1, characterized in that: The load balancing decision engine in step one includes a reinforcement learning decision engine, a global state collector, and a deep learning prediction module. The reinforcement learning decision engine comprises the following three elements: Status: Load balancing metrics include CPU utilization, memory usage, minimum number of connections, and network queue depth; Action: Defined as a proportional vector for allocating transactions to N nodes. The action space must satisfy the constraint that the sum of the allocation proportions of each node is 1. Rewards: Confirmed transaction throughput is used as the reward signal, which directly reflects the impact of the load balancing decision engine on the overall system performance.
3. A reliable broadcast method with load balancing strategy for DAG consensus as described in claim 2, characterized in that: The global status collector is responsible for collecting key performance indicators of cluster nodes in real time, including CPU utilization, memory usage, and network bandwidth.
4. A reliable broadcast method with load balancing strategy for DAG consensus as described in claim 3, characterized in that: The deep learning prediction module uses a lightweight LSTM network model to predict transaction volume, and incorporates real historical transaction data from the public chain to train a time series model that can predict future transaction load.
5. A reliable broadcast method with load balancing strategy for DAG consensus as described in claim 4, characterized in that: The VID server includes a VID client and a VID server. The VID client is responsible for integrating with the reliable broadcast module and calling its block distribution and reconstruction capabilities to send or retrieve block data to the VID server. The VID server is responsible for the fragmented storage and management of block data and supports reliable access to blocks by the client.
6. A reliable broadcast method with load balancing strategy for DAG consensus as described in claim 5, characterized in that: Step three includes the following steps: S1: Status acquisition, the global status collector collects real-time metrics from all nodes; S2: Action generation, the reinforcement learning decision engine takes the current state and expected transaction volume as input, and outputs a distribution ratio vector; S3: Distribute and execute transactions. Each node reliably broadcasts a number of transactions, and the remaining transactions are split according to the distribution ratio vector mentioned above and delegated to other nodes in a peer-to-peer manner.
7. A reliable broadcast method with load balancing strategy for DAG consensus as described in claim 6, characterized in that: During state acquisition in step S1, under the DAG consensus mechanism, each node maintains the same local graph structure. When a node broadcasts its current round certificate in each round, it attaches its own load information to achieve synchronization of the global load view. If a node is missing load data for the current round, a load estimation mechanism based on time decay is enabled. When a transaction is sent to a malicious node that claims to be under low load, but the system throughput does not improve, the reinforcement learning decision engine can gradually identify nodes with abnormal reporting behavior by observing reward feedback over a long period of time, and optimize the transaction distribution strategy accordingly, thereby reducing the reliance on potential fraudulent nodes.
8. A reliable broadcast method with load balancing strategy for DAG consensus as described in claim 7, characterized in that: Step four Includes the following steps: a: Sharding encoding and storage: Perform Reed-Solomon encoding on transaction batches, input 2*f+1 data blocks, output 3*f+1 shards; b: VID client distribution, sending the shards {shard1,...,shardn} to the VID servers on n nodes; c: Normal reliable broadcast process is triggered when a node collects storage certificates for 2*f+1 shards; d: The VID client retrieves the data after the normal reliable broadcast process is completed. It requests the fragmented data from the corresponding VID server, reconstructs the original data using Lagrange interpolation, and submits the complete transaction batch to the consensus layer.