Alliance chain fragment parameter dynamic optimization method based on deep reinforcement learning
By optimizing the sharding parameters of the government consortium blockchain through deep reinforcement learning, the performance of traditional methods under large-scale government business volume is improved, thereby increasing throughput and storage efficiency while ensuring system security and stability.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- DALIAN NEUSOFT UNIV OF INFORMATION
- Filing Date
- 2026-01-13
- Publication Date
- 2026-04-24
AI Technical Summary
Traditional consortium blockchain consensus methods are insufficient in handling large-scale government business volumes. Existing blockchain optimization schemes based on reinforcement learning have failed to effectively integrate sharding and consensus, and have not considered storage optimization.
A deep reinforcement learning-based approach is adopted to establish the state space and action space for government consortium blockchain systems. By combining the reward function, the number of shards, block size, and full node ratio are optimized. Dynamic parameter adjustment is performed through the DQN model to ensure security and performance.
It improves the throughput and storage efficiency of the government consortium blockchain, ensures that the system does not sacrifice fault tolerance under any circumstances, enhances the integration of sharding and consensus, and reduces the storage burden on nodes.
Smart Images

Figure FT_1 
Figure FT_2 
Figure FT_3
Abstract
Description
Technical Field
[0001] This invention relates to the field of e-government technology, and in particular to a method for dynamic optimization of consortium blockchain sharding parameters based on deep reinforcement learning. Background Technology
[0002] Since my country launched the construction of a nationwide integrated government service platform, scholars have been exploring technologies that can break down information barriers between different levels, regions, and departments of government and achieve interconnection and interoperability of government information systems while ensuring information security. Consortium blockchains, as an emerging information technology, possess irreplaceable advantages due to their immutable and traceable characteristics, which can be used to solve problems related to the reliable flow, sharing, and use of data faced by government service platforms.
[0003] However, government service systems encompass numerous services and handle a massive volume of transactions. According to data released by the Guangdong Provincial Government Service Data Management Bureau in February 2022, the "Yue Sheng Shi" platform alone processed an average of 59.7 million queries and transactions daily, and the number of service items and registered users continues to increase. Processing such a huge volume of transactions on a consortium blockchain is beyond the performance limits of traditional consortium blockchain consensus methods. Therefore, researching optimization methods for government consortium blockchain consensus has significant theoretical and practical value.
[0004] Recently, some scholars have proposed performance optimization schemes based on reinforcement learning in the blockchain field. Reinforcement learning is the learning of intelligent systems from environment to behavior mapping, and it is a well-known method for solving dynamic control problems. DQN is a dynamic programming algorithm that combines traditional reinforcement learning with deep neural networks (DNNs). It is often used in game development, such as Atari. Applying reinforcement learning algorithms to blockchain performance optimization has also yielded significant results. Liu et al. first proposed a blockchain scheme based on DQN for large-scale IoT networks. This scheme optimized the blockchain throughput under various constraints by adjusting blockchain parameters such as time interval, consensus algorithm, and block size. However, it did not use blockchain sharding technology, so there is still much room for improvement in throughput. Yun et al. applied DQN to blockchain sharding, trained a deep reinforcement learning agent to find the optimal parameters in various blockchain states, and adaptively optimized the system throughput and security level. However, it adopted a PoW-based sharding construction method. In real-world applications, cross-shard transactions are frequent, the integration of sharding and consensus is low, and storage optimization is not considered, limiting the optimization to throughput and security. Summary of the Invention
[0005] This invention discloses a method for dynamic optimization of consortium blockchain sharding parameters based on deep reinforcement learning, in order to overcome the above-mentioned technical problems.
[0006] To achieve the above objectives, the technical solution of the present invention is as follows: A method for dynamic optimization of sharding parameters in consortium blockchains based on deep reinforcement learning includes the following steps: S1: Establish the state space and action space of the consortium blockchain based on a deep reinforcement learning model for government consortium blockchain systems; S2: Based on the consortium blockchain constraints for government consortium blockchain systems, establish a reward function for consortium blockchains based on a deep reinforcement learning model for government consortium blockchain systems; S3: Employs an efficient sharding blockchain performance optimization model based on a deep reinforcement learning model to obtain the number of shards, block size, block interval, and full node ratio for sharding the consortium blockchain according to the state space, dynamic space, and reward function, thereby achieving dynamic optimization of the consortium blockchain sharding parameters.
[0007] Furthermore, the state space is represented as follows: (1) In the formula: For the state space of a consortium blockchain; Let be the set of node transmission rates, where , For the first The node to the first The transmission rate of each node, All are nodes. The total number of nodes; For the set of node computing capabilities, where, , For the first The computing power of each node; This is a collection of node consensus histories, where... , For the first The consensus history of each node; Let be the set of node transaction volumes, where , For the first The node and the first Transaction volume between nodes; This is a transpose.
[0008] Furthermore, the action space is represented as follows: (2) In the formula: For the action space of consortium blockchains; For time period Block size at time; For time period Block interval at time; For time period Number of fragments at time; This represents the proportion of all nodes.
[0009] Furthermore, the consortium blockchain constraints for the government consortium blockchain system include latency constraints for intra-shard transaction consensus processes, latency constraints for cross-shard transaction consensus processes, constraints on the number of shards and the number of faulty nodes, and constraints on the node classification ratio.
[0010] Furthermore, the latency constraints of the in-chip transaction consensus process include verification latency and transmission latency, as shown below: (3) In the formula: This is the sum of consensus latency and block interval, i.e., total latency; Block interval; The total consensus delay; The number of consecutive block intervals that conform to the final properties of a consortium blockchain; in: (4) In the formula: Local verification latency; The verification delay for the final consensus; Delayed local transmission; The delay in the dissemination of the final consensus; in, (5) = (6) = (7) In the formula: This represents the total number of fragments. The index for the shard; This is the local verification time. Delay for leaders to verify local consensus; For the verification delay of local consensus by followers; The total number of message authentication codes sent to the final consensus leader for each shard; To verify the computational cost of the block; For the first The number of nodes participating in consensus in each shard; The computational cost of the message authentication code; , The first The average computing power of the follower and leader nodes in each shard; (8) In the formula: To verify the computational cost of the block; This represents the proportion of all nodes; The computational power of the leader in the eventual consensus; (9) In the formula: Index the leader node; This is the index of a node within a shard, and its value is... , For the first The total number of nodes within each shard; Block size; For the first Within the first segment The leader node to the first Within the first segment Transmission rate between follower nodes; For the first Index of the leader node within each shard; For the first Indexes of follower nodes within each shard; For the first Within the first segment The leader node to the first Transmission rate between the leader nodes of the final consensus; An index for the leader node of the final consensus; This serves as the identifier for the final consensus node. This is the longest waiting time; (10) In the formula: This is the index number of the full node, with a value of [value]. ; For the first The leader node of the final consensus to the first Transmission rate between all nodes.
[0011] Furthermore, the latency constraints of the cross-shard transaction consensus process are expressed as follows:
[0012] (11) In the formula: The total consensus latency for cross-shard transactions; Verification latency for cross-segment transactions; For cross-slice transaction transmission latency; This is the sum of consensus latency and block interval for cross-shard transactions;
[0013] (12)
[0014] (13) In the formula: Verification delay for cross-shard transactions; For the first The computing power of the leader of each shard; Delay for verifying the final consensus of the leaders; To accommodate the verification delay of the final consensus; The number of nodes that reach the final consensus; The computational power of the leader in the eventual consensus; The computing power of followers of the final consensus; For cross-slice transaction transmission latency; This is the index of the final consensus follower node.
[0015] Furthermore, the constraints on the number of shards and the number of faulty nodes are expressed as follows: (14) In the formula: The index for the shard; The total number of nodes; This represents the proportion of faulty nodes.
[0016] Furthermore, the node classification ratio constraint is expressed as follows: (15) In the formula: The index for the shard; The total number of nodes; This represents the proportion of faulty nodes.
[0017] Furthermore, the reward function is established as follows: (16) In the formula: The value of the reward function; This represents throughput, which is the number of transactions a blockchain system can process per second. To optimize storage requirements; The number of shards in the blockchain; Block size; The size of the block header; The average size of the transactions; Block interval; For time period lThe average consensus latency of each block in the process. l Index for the time period; To round down; This is the proportionality coefficient; This represents the proportion of all nodes; This represents the total number of fragments. in, (17) (18) In the formula: This represents throughput, which is the number of transactions a blockchain system can process per second. The number of shards in the blockchain; Block size; The size of the block header; The average size of the transactions; Block interval; For time period l The average consensus latency of each block in the process. l Index for the time period; This is for rounding down.
[0018] Beneficial Effects: This invention provides a dynamic optimization method for consortium blockchain sharding parameters based on deep reinforcement learning. Based on the constraints of consortium blockchain systems for government consortium blockchains, it establishes a reward function for consortium blockchains based on a deep reinforcement learning model for these systems. Combining the state space and action space of this model, it employs an efficient sharding blockchain performance optimization model based on deep reinforcement learning to obtain the number of shards, block size, block interval, and full node ratio for sharding the consortium blockchain, thereby achieving dynamic optimization of sharding parameters. This invention establishes the reward function in conjunction with security constraints, rewarding not only high throughput and storage optimization but also ensuring that the system never sacrifices fault tolerance for performance improvements under any circumstances. By introducing full nodes and shard nodes at the node level and optimizing the full node ratio, shard number, block size, and block interval as key dimensions of the DQN action space, this invention significantly improves the integration of sharding and consensus in scenarios with frequent cross-shard transactions. Attached Figure Description
[0019] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0020] Figure 1 This is a flowchart of the dynamic optimization method for consortium blockchain sharding parameters based on deep reinforcement learning, as described in this invention. Figure 2 This is a schematic diagram of the overall logical structure of the consortium blockchain sharding parameter dynamic optimization method based on deep reinforcement learning in this embodiment of the invention. Figure 3 This is a schematic diagram of the segmented structure in an embodiment of the present invention; Figure 4 A schematic diagram of the transaction consensus structure in an embodiment of the present invention; Figure 5 This is a schematic diagram of the neural network structure of the sharding performance optimization model in an embodiment of the present invention; Figure 6 This is a schematic diagram of the dynamic optimization method for consortium blockchain sharding parameters based on deep reinforcement learning in an embodiment of the present invention. Figure 7 This is a schematic diagram of the DQN model structure in an embodiment of the present invention. Detailed Implementation
[0021] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0022] This embodiment introduces a method for dynamic optimization of consortium blockchain sharding parameters based on deep reinforcement learning, including the following steps: Figure 1 As shown: S1: Establish the state space and action space of the consortium blockchain for government consortium blockchain systems based on a deep reinforcement learning model (DQN model); The dynamic optimization method for consortium blockchain sharding parameters based on deep reinforcement learning in this embodiment mainly includes three parts: first, the establishment of a sharding consensus environment; second, performance constraint analysis; and third, an efficient sharding blockchain performance optimization model based on the DQN model.
[0023] This embodiment is geared towards government consortium blockchain systems, and optimizes the performance of dynamic clustering sharding and fault-tolerant consensus mechanisms. The dynamic clustering sharding and fault-tolerant consensus mechanism environment includes three parts: node classification, dynamic sharding, and consensus process. The sharding consensus environment is the carrier for dynamic parameter optimization in this embodiment. Its core issue is how to dynamically select the optimal sharding parameters (such as the number of shards, block size, block interval, and classification ratio) to continuously maintain high throughput, low latency, and high reliability consensus performance under complex government business loads and node states: (1) Node classification: Nodes are divided into full nodes that store complete chain data and shard nodes that only store data within shards to alleviate storage pressure, such as Figure 3 As shown. (2) Dynamic sharding: Based on dynamic state data such as transaction frequency, transmission rate, and computing power between nodes, nodes (including sharded nodes and full nodes) are divided into multiple logical shards through an improved clustering algorithm, aiming to reduce the proportion of cross-shard transactions and improve consensus parallelism. (3) Consensus process: Transactions are divided into intra-shard transactions and cross-shard transactions, which are executed in two phases of CFT consensus within the logical shard or in the final consensus group to ensure data consistency. The consensus process is as follows. Figure 4 As shown.
[0024] Specifically, due to the dynamic nature of consortium blockchains, this embodiment uses DQN, a classic deep reinforcement learning algorithm, and proposes a sharding scheme based on the DQN model. For example... Figure 2 As shown, the parameters to be optimized in the sharding consensus process constitute the action space. The states of each stage of the consortium blockchain are input into the state space to the DQN model. The neural network, through training, obtains the quality function value Q corresponding to each action, selects the action corresponding to the maximum Q value, and parses the optimized parameter results to guide the construction of the next stage of sharding consensus. After receiving the parameters selected by DQN, the sharding consensus environment analyzes the value of the parameters and defines a reasonable reward function. In this embodiment, since the optimization focuses on throughput and storage performance, the reward function is designed based on the system's throughput and storage optimization. Furthermore, to ensure the safe and reliable execution of system consensus, this embodiment provides security constraints on the parameters during the sharding consensus process to limit the reward function. The environment feeds back the reward function to the DQN model, and this process is repeated to allow the DQN model to continuously learn and ultimately select parameters with high reward values to guide the construction of sharding consensus. The reward function, action and state space, network model, and algorithm flow of the optimization model will be described in detail below.
[0025] Specifically, the state space directly affects the convergence speed and final performance of the DQN model; therefore, the choice of state space is crucial. A good state space should start from the ultimate goal, analyzing the essence of the goal and the factors involved, and what information embodies each factor. For the sharded consensus environment of this embodiment, the ultimate goal of optimization is to maximize the system's throughput and reduce the storage burden on nodes while ensuring the stable operation of the consortium blockchain system. This is also the basis for setting the reward function. Therefore, the essence of the optimization goal is related to two factors: throughput and storage optimization. Throughput is directly related to the computing power of nodes and the transmission rate between nodes in the government consortium blockchain system, directly affecting the consensus latency Tcon. Consensus history and the transaction matrix affect throughput and the node classification ratio, which is directly related to storage optimization, by influencing the construction of the sharding structure.
[0026] Therefore, in this embodiment, the transmission rate, node computing power, consensus history, and transaction matrix are selected to constitute the state space of the DQN model, as shown in the following equation: (1) In the formula: For the state space of a consortium blockchain; Let be the set of node transmission rates, where , For the first The node to the first The transmission rate of each node, All are nodes. The total number of nodes; For the set of node computing capabilities, where, , For the first The computing power of each node; This is a collection of node consensus histories, where... , For the first The consensus history of each node; Let be the set of node transaction volumes, where , For the first The node and the first Transaction volume between nodes; For transpose; Specifically, due to Due to differences in order of magnitude and hyperparameter dimensions, this embodiment converts all data into two-dimensional data and normalizes them before constructing the state space. The processed data is then concatenated to form the state space, which is then input into the neural network. The normalization method is a conventional technique for those skilled in the art, and therefore will not be described in detail here.
[0027] Specifically, block size, block interval, number of shards, and the proportion of full nodes are core parameters affecting the performance of government consortium blockchains. In government consortium blockchain systems, manually setting appropriate parameters is difficult, and fixed parameters cannot dynamically adapt to changes in the consortium blockchain environment. Therefore, this embodiment constructs the action space of the DQN model using these core parameters, dynamically finding the most suitable parameters to guide the consortium blockchain sharding consensus process. Here, block size refers to the size of the final block after aggregating the local blocks of each shard, stored in each full node.
[0028] Preferably, the action space is represented as follows: (2) In the formula: For the action space of consortium blockchains; For time period Block size at time; For time period Block interval at time; For time period Number of fragments at time; This represents the proportion of all nodes.
[0029] Specifically, in the implementation process, the neural network is set to output n. 4 4 represents the number of parameters, and n represents the sample size for each parameter. After selecting the corresponding action, it is analyzed to obtain... The combination of values. For example, if n=4, the output of the neural network is 256 actions. Assuming the index corresponding to the action with the highest value is 115, then action=115. The action is then parsed, and the block size is... For example, if the sample set is [1,2,3,4], then the modulo operation of action with respect to n is action%n, that is, 115%4=3. Take the (n+1)th value of action in the sample set, that is =4. The z-th parameter is taken from the sample set ( action / 4 z-1 )%n+1 values.
[0030] In the government consortium blockchain system, the structure guided by the DQN model is based on time periods. The structure remains unchanged within a certain time period. If the parameters guided by the DQN model change in subsequent time periods, the construction of shard consensus is based on the changed parameters.
[0031] S2: Based on the consortium blockchain constraints for government consortium blockchain systems, establish a reward function for consortium blockchains based on the DQN model for government consortium blockchain systems; Specifically, the reward function is the ultimate carrier of the DQN model's optimization logic, and the DQN algorithm optimization is the long-term cumulative return under this reward system. The role of the neural network is to transform the original state information into a form highly correlated with long-term returns after layers of nonlinear refinement, and further guide the generation of action decisions. Therefore, the reward function is the core element of the DQN model algorithm. The setting of the reward function depends on the optimization objective of the DQN model. For the sharded consensus environment of this embodiment, the ultimate goal of optimization is to maximize the system throughput and reduce the storage burden on nodes while ensuring the stable operation of the consortium blockchain system. Therefore, the setting of the reward function is divided into two parts: throughput and storage optimization. The stable operation of the government consortium blockchain is a crucial prerequisite. Therefore, the consortium blockchain constraints for the government consortium blockchain system include: latency constraints for intra-shard transaction consensus processes, latency constraints for cross-shard transaction consensus processes, constraints on the number of shards and the number of faulty nodes, and constraints on the node classification ratio.
[0032] Preferably, the latency constraints of the in-chip transaction consensus process, including verification latency and transmission latency, are expressed as follows: Specifically, this embodiment analyzes consensus latency. Latency is the time required for a transaction to go from being packaged into a block, through the consensus process, and finally committed to an irreversible state. Transactions submitted by nodes are added to either the intra-shard transaction pool or the cross-shard transaction pool based on the sender and receiver's determination. After allocation, verification is performed according to the corresponding consensus process. The transaction process consists of two steps: block interval and total consensus time. The total transaction latency is shown in the following formula.
[0033]
[0034] In the formula: Total transaction latency; Block interval; The total latency for consensus based on K shards, This represents the total number of fragments. in, The calculation method varies depending on the type of transaction. The following is a detailed analysis of the consensus latency for intra-shard transactions and cross-shard transactions.
[0035] Specifically, the total consensus time in the transaction pool of in-chip transactions includes local verification latency and final consensus verification latency, and the local verification latency and final consensus verification latency respectively include transmission latency and verification latency.
[0036] Specifically, transaction consensus latency should be completed within multiple consecutive block intervals to satisfy the finality property of the consortium blockchain. Adding up all time factors forms Constraint 1, the latency constraint of the intra-block transaction consensus process, as shown in the following formula: (3) In the formula: The total latency is the sum of the consensus latency of transactions within the block and the block interval. Block interval; The total consensus latency for in-film transactions; The number of consecutive block intervals that conform to the final properties of a consortium blockchain; (4) In the formula: Local verification latency; The verification delay for the final consensus; Delayed local transmission; The delay in the dissemination of the final consensus; The local verification latency is represented as follows: (5) = (6) = (7) In the formula: This represents the total number of fragments. The index for the shard; This is the local verification time. Delay for leaders to verify local consensus; For the verification delay of local consensus by followers; The total number of message authentication codes sent to the final consensus leader for each shard; To verify the computational cost of the block; For the first The number of nodes participating in consensus in each shard; To verify the computational cost of the block; , The first The average computing power of the follower and leader nodes in each shard; Specifically, the final consensus leader receives and processes the local blocks from each shard. MACs are merged with local blocks to generate... Each MAC (Message Authentication Code) is sent to all full nodes, and a commit message is sent to the shard nodes. Each node replies with a confirmation message; if more than half of the nodes respond, the blockchain is successfully uploaded. The transmission and verification time of the confirmation message is much shorter than that of the block verification and transmission, and therefore can be ignored. The final consensus verification latency is shown in the following formula: (8) In the formula: To verify the computational cost of the block; This represents the proportion of all nodes; The computational power of the leader in the eventual consensus.
[0037] Specifically, propagation latency refers to the time required for a message to reach the target node. To ensure smooth consensus, this embodiment sets a timeout in each consensus step. To prevent unresponsive nodes from excessively delaying the consensus process, local propagation delay is implemented, as shown in the following equation: (9) In the formula: Index the leader node; This is the index of a node within a shard, and its value is... , For the first The total number of nodes within each shard; Block size; For the first Within the first segment The leader node to the first Within the first segment Transmission rate between follower nodes; For the first Index of the leader node within each shard; For the first Indexes of follower nodes within each shard; For the first Within the first segment The leader node to the first Transmission rate between the leader nodes of the final consensus; An index for the leader node of the final consensus; This serves as the identifier for the final consensus node. The longest waiting time; among them, the node that reaches final consensus. (i.e., Final consensus) is used to distinguish the node sequence numbers in the shards.
[0038] Similarly, the propagation delay of the final consensus can be seen as follows: (10) In the formula: This is the index number of the full node, with a value of [value]. ; For the first The leader node of the final consensus to the first Transmission rate between all nodes.
[0039] Preferably, the time constraint of the cross-shard transaction consensus process, i.e. It is expressed as follows:
[0040] (11) In the formula: The total consensus latency for cross-shard transactions; Verification latency for cross-segment transactions; For cross-slice transaction transmission latency; This is the sum of consensus latency and block interval for cross-shard transactions; Specifically, during cross-chain transaction verification, the leaders of each shard collect a batch of data. Each transaction is assigned a message authentication code and its signature is verified. Upon successful verification, a MAC is generated for each transaction and sent to the final consensus leader, who then processes it. After generating MACs, Each MAC is sent to the followers s who will reach the final consensus. Each follower processes the leader's MAC, verifies it, and generates its own MAC, which is then sent to the leader. The leader processes... After processing MACs, based on the signature results, if successful, a [signature] will be generated. Each MAC address is sent to all full nodes. The verification and transmission delays in this process are shown in the following formulas:
[0041] (12)
[0042] (13) In the formula: Verification delay for cross-shard transactions; For the first The computing power of the leader of each shard; Delay for verifying the final consensus of the leaders; To accommodate the verification delay of the final consensus; The number of nodes that reach the final consensus; The computational power of the leader in the eventual consensus; The computing power of followers of the final consensus; For cross-slice transaction transmission latency; This is the index of the final consensus follower node.
[0043] Preferably, the constraint condition between the number of fragments and the number of faulty nodes, i.e. , means as follows: Specifically, from the perspective of horizontal scaling, the larger the total number of shards, the better the concurrency performance. However, in blockchain systems, there are nodes with poor performance. Therefore, when using reinforcement learning techniques to select the optimal shards... When the value is set, without any restrictions, the parameters obtained using DQN Mexico may not be able to carry out a normal consensus process. Therefore, the following constraints are imposed.
[0044] In a consortium blockchain system with strict identity verification, only faulty nodes exist, assuming their probability is... The total number of nodes is Based on the existing technology in this field, it is known that The smallest integer is .
[0045] The environment in this embodiment is a partitioned blockchain system. To prevent all faulty nodes from clustering in the same shard, therefore... , ,and, ,in therefore ,so Therefore, when the blockchain is a partitioned blockchain, that is... When the number of segments is greater than or equal to 2, the constraints on the number of faulty nodes are as follows: (14) In conclusion, when When the constraints of the number of shards and the number of faulty nodes are met, the blockchain consensus can be guaranteed to proceed reliably and securely.
[0046] Preferably, the node classification ratio constraint is expressed as follows: Specifically, since the final consensus needs to be reached by selecting full nodes... Each node is used as a full node (Rep_node), therefore the proportion of full nodes... There are constraints to ensure the consensus process can proceed smoothly. In the worst-case scenario, this... If all nodes are full nodes, then for consensus to be successfully reached, the number of full nodes must meet certain constraints. As shown in the following formula: (15) Specifically, even in Even if all faulty nodes are full nodes, there are still more than [number missing] faulty nodes. A full node enables it to have The final consensus among all nodes is completed normally (the number of effective nodes is greater than half).
[0047] Preferably, when the reward function for each time period satisfies the consortium blockchain constraints for the government consortium blockchain system, the value of the reward function is obtained as follows: (16) In the formula: The value of the reward function; This represents throughput, which is the number of transactions a blockchain system can process per second. To optimize storage requirements; The number of shards in the blockchain; Block size; The size of the block header; The average size of the transactions; Block interval; For time period l The average consensus latency of each block in the process. l Index for the time period; To round down; This is the proportionality coefficient; The total number of nodes; This represents the proportion of all nodes; This represents the total number of fragments. Specifically, storage optimization performance refers to reducing unnecessary redundancy while ensuring system security, enabling nodes to utilize limited storage space to store more valuable data. This is achieved through node classification. The time period and the total number of segments are Then each shard node only stores the total blockchain information. The number of shard nodes is Therefore, the optimizable storage amount can be obtained, as shown in the following formula: (17) The formula shows that optimized storage capacity increases with... The increase The value decreases as the number of nodes increases. However, to ensure the normal execution of the consensus mechanism, this embodiment adjusts the proportion of full nodes. Total number of fragments A reliability analysis was performed.
[0048] Specifically, throughput refers to the number of transactions a blockchain system can process per second. The number of transactions contained in each block of the blockchain is calculated by dividing the size of each block by the average size of the transactions. (The last sentence appears to be incomplete and possibly refers to a time period.) l The number of fragments is Block interval is Block size is There is blockchain The blockchain throughput can be calculated by processing each shard in parallel, as shown in the following formula. The block interval is the interval between the final block being added to the chain. To ensure finality, there is a certain time interval between the addition of a block to the chain and the start of consensus for the next block (the blocks here refer to the final block, which is the block formed by the aggregation of all shard blocks).
[0049] (18) In the formula: This represents throughput, which is the number of transactions a blockchain system can process per second. The number of shards in the blockchain; Block size; The size of the block header; The average size of the transactions; Block interval; For time period l The average consensus latency of each block in the process. l Index for the time period; This is for rounding down.
[0050] As can be seen from the formula, throughput is related to the number of fragments. Block size Positive correlation with block time interval Average consensus latency Negative correlation.
[0051] Specifically, when the reward function does not meet the reward function condition for each time period, the value of the reward function is 0.
[0052] Among them, the parameters of the reward function and the action space , , , The reward function, action space, and state space are closely related, complementing each other, which also ensures the stability of DQN.
[0053] S3: Establish an efficient sharding blockchain performance optimization model based on the DQN model, so as to obtain the number of shards, block size, block interval, and full node ratio for sharding the consortium blockchain according to the state space, dynamic space, and reward function, so as to achieve dynamic optimization of the sharding parameters of the consortium blockchain.
[0054] Specifically, since the data samples to be trained by the network model in this embodiment are three-dimensional, using a fully connected feedforward neural network model would significantly increase the training difficulty due to the large parameter scale. Therefore, this embodiment uses multiple cascaded convolutional neural network structures to compress and extract the fragmented sample data before training it through fully connected layers. The neural network structure diagram is shown below. Figure 5 As shown.
[0055] like Figure 2 As shown, the workflow of the deep reinforcement learning-based consortium blockchain sharding parameter dynamic optimization method consists of two parts: the consortium blockchain environment and the DQN model. This method obtains the state information of the consortium blockchain from the environment and provides it to the DQN model. The DQN model selects the action with the highest Q-value through a neural network and feeds it back to the environment, guiding the environment to construct the next sharding. However, reinforcement learning is very sensitive to dynamic changes in initialization and training processes. With good training examples, it may learn a better policy faster and better. If it does not encounter good training examples at the right time, it may fail to learn a good policy, resulting in some instability. Moreover, the action space proposed in this embodiment is relatively large. Using the traditional uniform sampling method will reduce the probability of obtaining useful information, making the training effect of DQN unstable and reducing the model's convergence speed. Therefore, this embodiment randomly selects actions before training, generates a certain number of samples, and stores them in the experience pool to ensure that the samples fully and evenly cover all action spaces. Then, a ranking-based random priority sampling method is adopted to increase the stability of DQN training.
[0056] This embodiment combines a fault-tolerant consensus mechanism, optimizes the throughput and storage of sharded consensus through DQN, and uses mathematical methods to provide constraints on the latency of the consensus process, the node classification ratio, and the security of the number of shards (constraints on the number of shards and the number of faulty nodes) to constrain the DQN model and ensure the reliability of consensus under the guidance of DQN.
[0057] Specifically, reinforcement learning is the process by which intelligent systems learn a mapping from the environment to actions in order to maximize the reward function. Reinforcement learning judges the quality of decisions by accumulating the reward; if all experienced states are the highest-value states, or all actions are the highest-value actions, then it is considered the optimal policy derived from the current value function. Reinforcement learning includes both model-based and model-free methods. The DQN model is a combination of model-free Q-learning and neural networks.
[0058] The DQN model is trained and updated based on a Markov Decision Process (MDP), which includes a state space, action space, reward function, and state transition probabilities. A DQN structure diagram is shown below. Figure 6 and Figure 7 As shown. During DQN training, the agent interacts with the environment to accumulate samples. The agent inputs state data into the Main-network, selects actions based on the quality function Q-value, and sends them to the environment. The environment changes its state based on the actions, calculates the reward function, and updates the Main-network based on the output of the Q-network, collecting (S) x A x ,R x ,S x+1 The data is stored in the sample pool. To achieve better results, improvements are made to the target network, greedy exploration, and experience recycling of DQN. Among these, S... x At time step x, the environmental state is in which the agent makes its decision; it is the input observed by the agent. x In state S x Below, the actions that the intelligent agent actually selects and executes; R x It is to perform action A. x Then, the immediate feedback obtained from the environment indicates whether the action was good or bad; S x+1 It is to perform action A. x Then, the new state at time step x+1.
[0059] (1) Target Network: Since the DQN model uses both sides of the Bellman equation as labels and predictions respectively, and trains the neural network in a supervised learning manner, if the same network is used for calculation, the estimation of the Q-value and the calculation of the target Q-value will affect each other, leading to algorithm divergence. A target network, namely the Q-network, is introduced to calculate the target Q-value. The parameter weights of the target network are not updated in real time, but are assigned to the parameters of the action network after a period of time. Therefore, the target network is relatively fixed, reducing the variance of the Q-value estimation and improving the stability of the algorithm. The action network is used to calculate the Q-value and is continuously updated and optimized by minimizing the loss function during the training process.
[0060] (2) Experience Recovery: To alleviate the strong correlation between samples and improve sample utilization, at each time step, the DQN agent records the action, reward, and state at time step x, and the state (S) at time x+1. x A x ,R x ,S x+1 The samples are stored in its sample pool and randomly extracted to train the DNN.
[0061] (3) Greedy exploration: In order to prevent getting trapped in local optimization, the concept of exploration is introduced on the basis of development. A greedy algorithm is used in network training, and the exploration probability is set. This probability decreases as the number of training times increases.
[0062] This embodiment has the following beneficial effects: (I) Achieving Dynamic Optimization of Sharding Consensus Parameters: Traditional consortium blockchain sharding parameters (such as the number of shards, block size, and block interval) typically employ static configurations or empirical thresholds, which cannot adapt to the dynamically changing business load, node performance, and network status in government scenarios. This results in the system operating in a suboptimal state for extended periods, with significant performance bottlenecks. This embodiment introduces a deep reinforcement learning (DQN) framework to construct a closed-loop optimization system of "state awareness - intelligent decision-making - environmental feedback." The system can collect node transmission rate, computing power, consensus history, and transaction correlation matrix in real time as state inputs, and dynamically output the optimal number of shards, block size, and other core parameters through a well-trained neural network. Experiments show that this method can improve the throughput of the government consortium blockchain by approximately 37% compared to traditional solutions.
[0063] (II) Embedding security constraints into the optimization mechanism to fundamentally ensure the security and reliability of the consensus process: Existing performance optimization schemes often unilaterally pursue high throughput and low latency, easily ignoring the security boundaries of the sharded system. For example, excessive sharding may lead to insufficient number of nodes in a single shard, resulting in a loss of fault tolerance. In this embodiment, the reinforcement learning optimization framework not only mathematically provides three core security constraints (consensus latency finality constraint, shard number fault tolerance constraint, and full node ratio constraint), but also embeds them as hard conditions into the DQN reward function design. The reward function not only rewards high throughput and storage optimization, but also ensures that the system will never sacrifice fault tolerance for performance improvement under any circumstances.
[0064] (III) Innovative Design of a Joint Optimization Model for Storage Efficiency and Consensus Performance to Significantly Reduce Node Storage Burden: Government consortium blockchain nodes typically need to maintain multiple business chains and basic information chains, resulting in significant storage pressure. Existing sharding solutions mostly focus only on transaction processing performance, neglecting to include storage overhead in the optimization objective. This embodiment proposes a "performance-storage" joint optimization model, introducing an elastic classification mechanism at the node level between full nodes (full storage nodes) and sharded nodes (sharded storage nodes), and optimizing the proportion of full nodes as one of the key dimensions of the DQN action space. The reward function explicitly includes a storage optimization term, incentivizing agents to minimize the proportion of full nodes while ensuring data redundancy and security, allowing more nodes to store only data within their shards. Theoretical analysis and experimental data show that this mechanism can dynamically optimize storage distribution based on the number of shards, saving up to 78% of storage space for the system in the optimal case of 8 shards.
[0065] (iv) Employing priority sampling and a dedicated network structure significantly improves training efficiency and algorithm stability: Applying deep reinforcement learning to complex consortium blockchain parameter optimization faces challenges such as high-dimensional state space, explosive action combinations, and uneven training sample quality, which can easily lead to slow convergence, large fluctuations, and difficulty in practical application. This embodiment addresses the structured characteristics of consortium blockchain state data (such as node relationship matrices) by introducing a hybrid network architecture of "Convolutional Neural Network (CNN) + fully connected layers." The CNN layer effectively extracts local correlations and spatial features between nodes, significantly improving state representation capabilities and reducing the number of network parameters. Simultaneously, a ranking-based random priority experience replay strategy is adopted to replace traditional uniform sampling, prioritizing the training of historical experiences with large prediction errors and high learning value, enabling efficient reuse of valuable exploration experience. These improvements ensure that this method still significantly outperforms the comparative schemes in terms of convergence speed and stability even with a higher action space dimension (4-dimensional).
[0066] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for dynamic optimization of sharding parameters in a consortium blockchain based on deep reinforcement learning, characterized in that, Includes the following steps: S1: Establish the state space and action space of the consortium blockchain based on a deep reinforcement learning model for government consortium blockchain systems; S2: Based on the consortium blockchain constraints for government consortium blockchain systems, establish a reward function for consortium blockchains based on a deep reinforcement learning model for government consortium blockchain systems; S3: Employs an efficient sharding blockchain performance optimization model based on a deep reinforcement learning model to obtain the number of shards, block size, block interval, and full node ratio for sharding the consortium blockchain according to the state space, dynamic space, and reward function, thereby achieving dynamic optimization of the consortium blockchain sharding parameters.
2. The method for dynamic optimization of consortium blockchain sharding parameters based on deep reinforcement learning according to claim 1, characterized in that, The state space is represented as follows: (1) In the formula: For the state space of a consortium blockchain; Let be the set of node transmission rates, where , For the first The node to the first The transmission rate of each node, All are nodes. The total number of nodes; For the set of node computing capabilities, where, , For the first The computing power of each node; This is a collection of node consensus histories, where... , For the first The consensus history of each node; Let be the set of node transaction volumes, where , For the first The node and the first Transaction volume between nodes; This is a transpose.
3. The method for dynamic optimization of consortium blockchain sharding parameters based on deep reinforcement learning according to claim 1, characterized in that, The action space is represented as follows: (2) In the formula: For the action space of consortium blockchains; For time period Block size at time; For time period Block interval at time; For time period Number of fragments at time; This represents the proportion of all nodes.
4. The method for dynamic optimization of consortium blockchain sharding parameters based on deep reinforcement learning according to claim 1, characterized in that, The consortium blockchain constraints for the government consortium blockchain system include latency constraints for intra-shard transaction consensus processes, latency constraints for cross-shard transaction consensus processes, constraints on the number of shards and the number of faulty nodes, and constraints on the node classification ratio.
5. The method for dynamic optimization of consortium chain sharding parameters based on deep reinforcement learning according to claim 4, characterized in that, The latency constraints of the in-chip transaction consensus process include verification latency and transmission latency, as shown below: (3) In the formula: This is the sum of consensus latency and block interval, i.e., total latency; Block interval; The total consensus delay; The number of consecutive block intervals that conform to the final properties of a consortium blockchain; in: (4) In the formula: Local verification latency; The verification delay for the final consensus; Delayed local transmission; The delay in the dissemination of the final consensus; in, (5) = (6) = (7) In the formula: This represents the total number of fragments. The index for the shard; This is the local verification time. Delay for leaders to verify local consensus; For the verification delay of local consensus by followers; The total number of message authentication codes sent to the final consensus leader for each shard; To verify the computational cost of the block; For the first The number of nodes participating in consensus in each shard; The computational cost of the message authentication code; , The first The average computing power of the follower and leader nodes in each shard; (8) In the formula: To verify the computational cost of the block; This represents the proportion of all nodes; The computational power of the leader in the eventual consensus; (9) In the formula: Index the leader node; This is the index of a node within a shard, and its value is... , For the first The total number of nodes within each shard; Block size; For the first Within the first segment The leader node to the first Within the first segment Transmission rate between follower nodes; For the first Index of the leader node within each shard; For the first Indexes of follower nodes within each shard; For the first Within the first segment The leader node to the first Transmission rate between the leader nodes of the final consensus; An index for the leader node of the final consensus; This serves as the identifier for the final consensus node. This is the longest waiting time; (10) In the formula: This is the index number of the full node, with a value of [value]. ; For the first The leader node of the final consensus to the first Transmission rate between all nodes.
6. The method for dynamic optimization of consortium chain sharding parameters based on deep reinforcement learning according to claim 4, characterized in that, The latency constraints of the cross-shard transaction consensus process are expressed as follows: (11) In the formula: The total consensus latency for cross-shard transactions; Verification latency for cross-segment transactions; For cross-slice transaction transmission latency; This is the sum of consensus latency and block interval for cross-shard transactions; (12) (13) In the formula: Verification delay for cross-shard transactions; For the first The computing power of the leader of each shard; Delay for verifying the final consensus of the leaders; To accommodate the verification delay of the final consensus; The number of nodes that reach the final consensus; The computational power of the leader in the eventual consensus; The computing power of followers of the final consensus; For cross-slice transaction transmission latency; This is the index of the final consensus follower node.
7. The method for dynamic optimization of consortium chain sharding parameters based on deep reinforcement learning according to claim 4, characterized in that, The constraints on the number of shards and the number of faulty nodes are expressed as follows: (14) In the formula: The index for the shard; The total number of nodes; This represents the proportion of faulty nodes.
8. The method for dynamic optimization of consortium blockchain sharding parameters based on deep reinforcement learning according to claim 4, characterized in that, The node classification ratio constraint is expressed as follows: (15) In the formula: The index for the shard; The total number of nodes; This represents the proportion of faulty nodes.
9. The method for dynamic optimization of consortium chain sharding parameters based on deep reinforcement learning according to claim 1, characterized in that, The reward function is established as follows: (16) In the formula: The value of the reward function; This represents throughput, which is the number of transactions a blockchain system can process per second. To optimize storage requirements; The number of shards in the blockchain; Block size; The size of the block header; The average size of the transactions; Block interval; For time period l The average consensus latency of each block in the process. l Index for the time period; To round down; This is the proportionality coefficient; This represents the proportion of all nodes; This represents the total number of fragments. in, (17) (18) In the formula: This represents throughput, which is the number of transactions a blockchain system can process per second. The number of shards in the blockchain; Block size; The size of the block header; The average size of the transactions; Block interval; For time period l The average consensus latency of each block in the process. l Index for the time period; This is for rounding down.
Citation Information
Cited By
Construction of alliance chain sharding based on dynamic clustering and consensus method
CN122262724A