A Method for Optimizing Blockchain Sharding Strategy Based on Model Reinforcement Learning
Through the blockchain sharding strategy optimization method based on model reinforcement learning, the optimal sharding strategy is generated, which solves the problem of low sampling efficiency of blockchain sharding method in the existing technology and improves the throughput of blockchain.
Patent Information
- Application Number
- CN202411437115.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-15
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2044-10-15
AI Technical Summary
In the prior art, the blockchain sharding method frequently interacts with the blockchain system through model-free reinforcement learning (MFRL), resulting in low sampling efficiency and limiting the improvement of blockchain throughput.
The blockchain sharding strategy optimization method based on model reinforcement learning is adopted. By obtaining blockchain state data, the initial policy network and cross-entropy algorithm are used to generate the optimal behavioral trajectory, the optimal policy network is trained, and the optimal sharding strategy is generated, which reduces the number of interactions with the real environment.
It improves the sampling efficiency of blockchain when learning the optimal sharding strategy, enhances the throughput of blockchain, reduces the number of interactions between servers and real environments, and improves sampling efficiency.
Smart Images

Figure CN119383192B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of blockchain sharding, and particularly to a method for optimizing blockchain sharding strategy based on model reinforcement learning. Background Art
[0002] As a distributed ledger jointly maintained by multiple participants, the blockchain has the characteristics of being tamper-proof and forgery-proof. Globally, blockchain technology has attracted extensive attention in the academic and industrial fields. However, when applied to large-scale systems, the main challenge faced by the blockchain lies in its scalability.
[0003] In the prior art, the blockchain network is usually divided into multiple parallel processing shards through blockchain sharding technology (such as Zilliqa) to improve the throughput of the blockchain. For example, model-free reinforcement learning (MFRL) is used to solve the optimal blockchain sharding (OBCS) problem to achieve dynamic sharding.
[0004] However, in the blockchain sharding method of the prior art, MFRL usually needs to interact with the blockchain system a large number of times, which to a certain extent limits the improvement of the sampling efficiency and results in a low throughput of the blockchain. Summary of the Invention
[0005] Based on this, it is necessary to provide a method for optimizing blockchain sharding strategy based on model reinforcement learning to solve the above technical problems, which improves the sampling efficiency when learning the optimal sharding strategy of the blockchain, thereby improving the throughput of the blockchain.
[0006] The present invention adopts the following technical solutions:
[0007] The present invention provides a method for optimizing blockchain sharding strategy based on model reinforcement learning, including:
[0008] Obtain the current state data of the blockchain, and input the current state data into the initial policy network to obtain the behavior data of the blockchain; the behavior data of the blockchain is used to guide blockchain sharding;
[0009] Perform state transition on the blockchain according to the behavior data of the blockchain to obtain the state data of the blockchain at the next moment;
[0010] Obtain the initial state data of the prediction model, and input the initial state data into the cross-entropy algorithm. Through the cross-entropy algorithm, call the prediction model to predict the state data at future moments, and call the model policy network to predict the behavior data corresponding to the state data, so as to obtain the optimal behavior trajectory of the blockchain; the optimal behavior trajectory includes the state data and behavior data of the blockchain at multiple moments;
[0011] Obtain the initial state behavior pairs from the optimal behavior trajectory, and train the initial policy network through the initial state behavior pairs to obtain the optimal policy network; the optimal policy network is used to generate the optimal sharding policy of the to-be-sharded blockchain.
[0012] Preferably, input the current state data into the initial policy network to obtain the behavior data of the blockchain, including:
[0013] Input the current state data into the initial policy network to obtain the probability values of each discrete behavior of the blockchain; the discrete behaviors include the block size of the blockchain, the number of shards, and the time required to generate a block;
[0014] Determine the behavior data of the blockchain according to the probability values of each discrete behavior.
[0015] Preferably, through the cross-entropy algorithm, call the prediction model to predict the state data at future moments, and call the model policy network to predict the behavior data corresponding to the state data, so as to obtain the optimal behavior trajectory of the blockchain, including:
[0016] Through the cross-entropy algorithm, call the prediction model to predict the state data at future moments, and call the model policy network to predict the behavior data corresponding to the state data, so as to obtain multiple different behavior trajectories; each behavior trajectory includes the state data, behavior data, and corresponding reward values of the blockchain at multiple moments;
[0017] According to the multiple reward values in each behavior trajectory, determine the total return value of each behavior trajectory;
[0018] Determine the preset number of behavior trajectories with the highest total return value as the optimal behavior trajectory of the blockchain.
[0019] Preferably, through the cross-entropy algorithm, call the prediction model to predict the state data at future moments, and call the model policy network to predict the behavior data corresponding to the state data, so as to obtain multiple different behavior trajectories, including:
[0020] For any one behavior trajectory, input the initial state data into the model policy network to obtain the behavior data corresponding to the initial state data, and calculate the reward value of the behavior data;
[0021] Input the initial state data into the prediction model to obtain the state data at the next moment;
[0022] Determine the behavior data and reward value at the next moment according to the state data at the next moment;
[0023] Determine the behavior trajectory according to the state data, behavior data and reward value at each moment.
[0024] Preferably, calculating the reward value of the behavior data includes:
[0025] Judge whether the blockchain corresponding to the behavior data meets the preset constraint conditions;
[0026] When the blockchain corresponding to the behavior data meets the constraint conditions, substitute the behavior data into the preset reward function to obtain the reward value;
[0027] When the blockchain corresponding to the behavior data does not meet the constraint conditions, determine that the reward value is 0;
[0028] The constraint conditions are:
[0029]
[0030] TI + TI K ≤ uTI, K = 1, 2,..., K max
[0031] Where N is the number of nodes in the blockchain, P t is the number of malicious nodes in the blockchain, TI is the time required to generate a block, and TI K is the total consensus time when there are K shards in the blockchain, and u is the block interval.
[0032] Preferably, the method further includes:
[0033] Train the model policy network through the initial state behavior pair to update the model policy network.
[0034] Preferably, obtaining the initial state behavior pair from the optimal behavior trajectory includes:
[0035] Obtain the behavior data corresponding to the initial state data in the optimal behavior trajectory;
[0036] Determine the initial state behavior pair according to the initial state data and the corresponding behavior data.
[0037] Preferably, the construction process of the prediction model includes:
[0038] Interact the blockchain with the initial policy network for multiple time steps to obtain multiple data quadruples, and store the multiple data quadruples in the first experience replay buffer; each data quadruple includes the current state data, action data, reward value, and state data at the next moment of the blockchain.
[0039] Train the initial prediction model based on the multiple data quadruples in the first experience replay buffer to obtain a prediction model.
[0040] Preferably, the method further includes:
[0041] After obtaining the state data of the blockchain at the next moment, substitute the action data of the blockchain into the reward function to obtain a reward value, and store the current state data, action data, reward value, and state data at the next moment of the blockchain as a data quadruple in the second experience replay buffer.
[0042] Train the prediction model with the data quadruples in the first experience replay buffer and the second experience replay buffer to obtain an updated prediction model; the updated prediction model is used for predicting the state data of the blockchain.
[0043] Preferably, obtaining the initial state data of the prediction model includes:
[0044] Randomly select a state data of the blockchain from the first experience replay buffer and the second experience replay buffer as the initial state data of the prediction model.
[0045] The present invention provides a blockchain sharding strategy optimization device based on model reinforcement learning, including:
[0046] An acquisition module, configured to acquire the current state data of the blockchain and input the current state data into the initial policy network to obtain the action data of the blockchain; the action data of the blockchain is used to guide blockchain sharding.
[0047] A state transition module, configured to perform state transition on the blockchain according to the action data of the blockchain to obtain the state data of the blockchain at the next moment.
[0048] A solution module, configured to obtain the initial state data of the prediction model, input the initial state data into the cross-entropy algorithm, and call the prediction model through the cross-entropy algorithm to predict the state data at a future moment, and call the model policy network to predict the action data corresponding to the state data, so as to obtain the optimal action trajectory of the blockchain; the optimal action trajectory includes the state data and action data of the blockchain at multiple moments.
[0049] A training module, configured to obtain an initial state-action pair from an optimal behavior trajectory and train an initial policy network with the initial state-action pair to obtain an optimal policy network; the optimal policy network is configured to generate an optimal sharding policy for a blockchain to be sharded.
[0050] The present invention provides a computer-readable storage medium storing a computer program, which when executed by a processor implements the above-mentioned blockchain sharding policy optimization method based on model reinforcement learning.
[0051] The present invention provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, where when the processor executes the program, it implements the above-mentioned blockchain sharding policy optimization method based on model reinforcement learning.
[0052] At least one of the above technical solutions adopted by the present invention can achieve the following beneficial effects:
[0053] In the present invention, first, the server interacts with the blockchain to obtain the state data of the blockchain in the real environment, and then further learns the optimal behavior trajectory of the blockchain through the cross-entropy algorithm. The cross-entropy algorithm can optimize the discrete behavior space, thereby exploring and determining an optimal behavior trajectory that can maximize the throughput of the blockchain sharding system. And during the decision interval, most of the time is used to learn the optimal policy through the cross-entropy algorithm. The cross-entropy algorithm generates a model data stream based on the current simulated environment, rather than directly obtaining data from the real environment. Since the behavior data generation step is only a forward pass of the policy network, its time cost is negligible compared to the length of the decision interval. Therefore, such an implementation process effectively combines the simulated environment with the blockchain, reduces the number of interactions between the server and the real environment, thereby greatly improving the sampling efficiency when learning the optimal sharding policy of the blockchain, and thus improving the throughput of the blockchain. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of the present invention. The illustrative embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:
[0055] Figure 1 It is a schematic flow chart of a blockchain sharding policy optimization method based on model reinforcement learning provided by the present invention;
[0056] Figure 2 It is a schematic flow chart of another blockchain sharding policy optimization method based on model reinforcement learning provided by the present invention;
[0057] Figure 3Schematic diagram of the structure of a blockchain system provided by the present invention;
[0058] Figure 4 Schematic diagram of the structure of a blockchain sharding strategy optimization system based on model reinforcement learning provided by the present invention;
[0059] Figure 5 Schematic diagram of the process of another blockchain sharding strategy optimization method based on model reinforcement learning provided by the present invention;
[0060] Figure 6 Schematic diagram of the process of another blockchain sharding strategy optimization method based on model reinforcement learning provided by the present invention;
[0061] Figure 7 Block diagram of the structure of a blockchain sharding strategy optimization device based on model reinforcement learning provided by the present invention;
[0062] Figure 8 Schematic diagram of a computer device for implementing the blockchain sharding strategy optimization method based on model reinforcement learning provided by the present invention. Detailed implementation manners
[0063] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below in conjunction with the specific embodiments and corresponding drawings of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0064] The blockchain sharding strategy optimization method based on model reinforcement learning provided by the embodiments of the present application can be applied to a server, which can be a server set up in a business platform or a device such as a desktop computer or a laptop computer that can execute the solution of the present invention.
[0065] The following details the technical solutions provided by each embodiment of the present invention in conjunction with the drawings.
[0066] Figure 1 Schematic diagram of the process of a blockchain sharding strategy optimization method based on model reinforcement learning in the present invention, specifically including the following steps:
[0067] S101. Obtain the current state data of the blockchain and input the current state data into the initial policy network to obtain the behavior data of the blockchain; the behavior data of the blockchain is used to guide blockchain sharding.
[0068] Among them, the state data of the blockchain may include the set of computing capabilities of the nodes in the blockchain, the set of transmission rates between the nodes, and the set of consensus histories. Optionally, the state data of the blockchain at time t can be defined as follows:
[0069] S t = [C t , V t , H t (1)
[0070] Among them, C t = {C t,i}, i ∈ N is the set of computing capabilities of the nodes, and N represents the number of nodes in the blockchain. V t = {V t,ij}, i, j ∈ N is the set of transmission rates between the nodes. H t = {H t,i}, i ∈ N is the set of consensus histories of the previous epoch (an epoch usually refers to a time period or a batch of block generations); when the consensus opinion of node i is valid, H t,i = 0, otherwise, H t,i = 1.
[0071] The initial policy network can be a neural network model that predicts the behavior space of the blockchain through the state data of the blockchain. Therefore, the way to input the current state data into the initial policy network to obtain the behavior data of the blockchain can be to input the current state data into the initial policy network to obtain the probability values of each discrete behavior of the blockchain; the discrete behaviors include the block size, the number of shards, and the time required to generate a block of the blockchain; and determine the behavior data of the blockchain according to the probability values of each discrete behavior.
[0072] The behavior space of the blockchain at time t can be defined as follows:
[0073] a t = [B t , K t , TI t (2)
[0074] Among them, B t = {1, 2, 3,..., B max} represents the block size, K t = {1, 2, 3,..., K max} represents the number of shards of the blockchain, and TI t = {1, 2, 3,..., TI max} represents the time required to generate a block. It should be noted that B t , K t , TI t are all discrete variables. Bmax can be the maximum block size, K max can be the maximum number of shards into which the blockchain can be sharded, TI max is the longest time required to generate a block.
[0075] Therefore, input the current state data into the initial policy network, and the initial policy network can output the probability values of each discrete behavior, that is, output B t , K t and TI t 's probability values; taking B t as an example, for example, B t = {1, 2, 3}, output the probability values of B t = 1, B t = 2, and B t = 3.
[0076] In this way, according to the probability values of each discrete behavior, determine the behavior data of the blockchain, including: for example, for any behavior, determine the value with the largest probability value as the value of the corresponding behavior, and take the value of each behavior as the value of the behavior data; or, randomly sample the values of each discrete behavior to determine the behavior data of the blockchain.
[0077] S102, perform a state transition on the blockchain according to the behavior data of the blockchain to obtain the state data of the blockchain at the next moment.
[0078] Interact with the blockchain system through the behavior data a t of the blockchain to transfer the blockchain to the next state to obtain the state data of the blockchain at the next moment.
[0079] At the same time, determine the reward value of the blockchain according to the behavior data of the blockchain, that is, substitute the behavior data of the blockchain into the reward function to obtain the reward value, and use the current state data, behavior data, reward value, and state data at the next moment of the blockchain as the data quadruple (s t , a t , r t , s t+1 ) and store it in the second experience replay buffer; the second experience replay buffer can be represented by D RL .
[0080] S103, obtain the initial state data of the prediction model, and input the initial state data into the cross-entropy algorithm. Through the cross-entropy algorithm, call the prediction model to predict the state data at future moments, and call the model policy network to predict the behavior data corresponding to the state data, to obtain the optimal behavior trajectory of the blockchain.
[0081] Among them, the optimal behavior trajectory includes the state data and behavior data of the blockchain at multiple moments.
[0082] Before sharding, a prediction model can be constructed first, and the state data of the blockchain at the next moment can be predicted through the prediction model. Preferably, the construction process of the prediction model includes: interacting the blockchain with the initial policy network for multiple time steps to obtain multiple data quadruples, and storing the multiple data quadruples in the first experience replay buffer; each data quadruple includes the current state data, action data, reward value, and state data at the next moment of the blockchain; training the initial prediction model according to the multiple data quadruples in the first experience replay buffer to obtain the prediction model.
[0083] Specifically, the initial policy network can be interacted with the blockchain system for G time steps to generate a series of data quadruples composed of state data, action data, reward value, and the next state (s g ,a g ,r g ,s g+1 ), g ∈ {0, 1,..., G - 2}. These data quadruples can be stored in the first experience replay buffer D RAND in. D RAND is used to train the prediction model f (prediction model input: current state data; output: state data at the next moment), and sample experiences in the initial stage.
[0084] The way to obtain the initial state data of the prediction model can be to randomly select a state data of the blockchain from the first experience replay buffer and the second experience replay buffer as the initial state data of the prediction model. For example, randomly select an experience (s i ,a i ,r i ,s i+1 ) from the first experience replay buffer and the second experience replay buffer. i is a certain moment, and s i is used as the initial state data s'0 for the prediction model to roll. Starting from s'0, the cross-entropy algorithm (Cross-Entropy Method, CEM) can be used to obtain the behavior trajectory ζ = (s i ,a'0,r'0,s' i+1 ...,s i+T-1 ,a' T-1 ,r' T-1 ,s' i+T ).
[0085] Preferably, as Figure 2As shown, input the initial state data into the cross-entropy algorithm. Through the cross-entropy algorithm, call the prediction model to predict the state data at future moments, and call the model policy network to predict the behavior data corresponding to the state data, to obtain the optimal behavior trajectory of the blockchain, including the following steps:
[0086] S201, input the initial state data into the cross-entropy algorithm. Through the cross-entropy algorithm, call the prediction model to predict the state data at future moments, and call the model policy network to predict the behavior data corresponding to the state data, to obtain multiple different behavior trajectories.
[0087] Among them, each behavior trajectory includes the state data, behavior data, and corresponding reward values of the blockchain at multiple moments.
[0088] Preferably, for any behavior trajectory, the initial state data can be input into the model policy network to obtain the behavior data corresponding to the initial state data, and calculate the reward value of the behavior data. Then, input the initial state data into the prediction model to obtain the state data at the next moment; and determine the behavior data and reward value at the next moment according to the state data at the next moment; finally, determine the behavior trajectory according to the state data, behavior data, and reward value at each moment.
[0089] The model policy network can be a policy network with the same structure as the initial policy network but different weight parameters; starting from the initial state data s' j begin, input the initial state data s' j into the model policy network π * to obtain the behavior data a' * output by the model policy network π j , and calculate the reward value r' j of the behavior data a' j , then use the prediction model f to predict the state data s' j+1 at the next moment, and determine the behavior data and reward value at the next moment according to the state data at the next moment. This process can continue for T steps to generate a complete behavior trajectory ζ=(s i ,a'0,r'0,s' i+1 ...,s i+T-1 ,a' T-1 ,r' T-1 ,s' i+T ).
[0090] By repeating this complete process S times, S different behavior trajectories can be obtained. It should be noted that the initial state data corresponding to each behavior trajectory can be different. By obtaining different initial state data, multiple different behavior trajectories are constructed.
[0091] Optionally, calculate the reward value of the behavior data, including: determining whether the blockchain corresponding to the behavior data meets a preset constraint condition; substituting the behavior data into a preset reward function to obtain a reward value when the blockchain corresponding to the behavior data meets the constraint condition; and determining that the reward value is 0 when the blockchain corresponding to the behavior data does not meet the constraint condition.
[0092] The reward function is shown in formula (3).
[0093]
[0094] Among them, r t represents the reward value of the state data at time t, B h represents the size of the block header, b represents the average size of a transaction, and (B t -B h ) represents the size of the transactions processed in each block.
[0095] The constraint condition is:
[0096]
[0097] TI + TI K ≤ uTI, K = 1, 2,..., K max (6)
[0098] Among them, N is the number of nodes in the blockchain, P t is the number of malicious nodes in the blockchain, TI is the time required to generate a block, TI K is the total consensus time when there are K shards in the blockchain, and u is the block interval.
[0099] It should be noted that formulas (4) and (5) are security constraints, aiming to ensure that after sharding, malicious nodes within a shard do not threaten the security of the entire shard and prevent blocks generated by malicious nodes from being uploaded to the blockchain. Formula (6) is a latency constraint, aiming to meet the ultimate property of the blockchain, that is, the latency should be completed within several consecutive block intervals u, TI K represents the consensus time within a blockchain shard, and the calculation of the malicious node probability P t depends on the consensus history.
[0100] S202. Determine the total return value of each behavior track according to the multiple reward values in each behavior track.
[0101] For any behavior track, the multiple reward values of each behavior track can be substituted into the objective function to obtain an objective function value, and the objective function value obtained for each behavior track is used as the total return value of each behavior track. The objective function is shown in formula (7).
[0102]
[0103] Among them, T represents the number of reward values, and γ t represents the discount factor corresponding to the behavior data at time t.
[0104] S203. Determine the optimal behavior trajectories of the blockchain by identifying the preset number of behavior trajectories with the highest total return values.
[0105] Select the preset number of behavior sequences with the highest total return values from the generated behavior trajectories as the optimal behavior trajectories; for example, select the top e behavior sequences with the highest total return from the generated S trajectories to form the optimal behavior trajectories.
[0106] S104. Obtain the initial state-action pairs from the optimal behavior trajectories and train the initial policy network using the initial state-action pairs to obtain the optimal policy network.
[0107] Among them, the optimal policy network is used to generate the optimal sharding policy for the blockchain to be sharded.
[0108] Obtaining the initial state-action pairs from the optimal behavior trajectories includes: obtaining the behavior data corresponding to the initial state data in the optimal behavior trajectories; determining the initial state-action pairs based on the initial state data and the corresponding behavior data.
[0109] The initial state-action pairs can include one state data and behavior data. The initial state data can be used as the state data in the initial state-action pairs, and the behavior data corresponding to the behavior trajectory with the highest total return value among the behavior data corresponding to the initial state data in the optimal behavior trajectories can be used as the behavior data in the initial state-action pairs. That is, the initial state-action pair is: (s i , a'0).
[0110] The initial state-action pairs can be stored in the third experience replay buffer D model and the initial policy network is trained based on the initial state-action pairs in D elite The trained initial policy network is determined as the optimal policy network. Optionally, during the training process, the following defined loss function can be used to optimize the initial policy network: Among them, θ is the parameter of the initial policy network
[0111]
[0112] where θ is the parameter of the initial policy network of.
[0113] Optionally, taking the preset number as e as an example, the initial state-action pairs can also include the initial state data and all the behavior data corresponding to the initial state data in the optimal behavior trajectories. That is, the initial state-action pair is: The model policy network can be trained through this initial state-action pair to update the model policy network. Specifically, the initial state-action pair can be stored in the elite experience replay buffer D elite and the model policy network π elite is trained according to the formula (9) and the initial state-action pairs in D * .
[0114]
[0115] where θ - is the parameter of the policy network π * and α is the learning rate. This process increases the probability that π * takes actions within D elite (representing the high-return action space). By iterating the entire process Z times, the cross-entropy algorithm can better explore the solution space and approximate the optimal solution to the OBCS problem. During the iteration process, the value of the parameter S gradually decreases, enabling the algorithm to focus more on the high-return region in the later stage. After the last iteration, the model selects the optimal action trajectory ζ with the best action sequence.
[0116] After obtaining the optimal policy network, the prediction model can also be trained using the data quadruples in the first experience replay buffer and the second experience replay buffer to obtain an updated prediction model; the updated prediction model is used for predicting the state data of the blockchain.
[0117] Optionally, the state data of the blockchain to be sharded can be input into the optimal policy network to obtain the sharding policy of the blockchain to be sharded.
[0118] where the blockchain to be sharded is the blockchain to be traded; the state data of the blockchain to be sharded is input into the optimal policy network to obtain the probability value of each discrete action of the blockchain to be sharded; the probability value of each discrete action is determined as the sharding policy of the blockchain to be sharded.
[0119] Optionally, the value of each discrete action of the blockchain to be sharded can also be determined according to the probability value of the discrete action, that is, the value with the largest probability value in each discrete action is used as the value of each discrete action, and then the value of each discrete action is used as the action data and determined as the sharding policy of the blockchain to be sharded.
[0120] After obtaining the sharding policy of the blockchain to be sharded, the time taken for each node in the blockchain to be sharded to solve the proof-of-work puzzle is obtained, and the nodes are sorted in ascending order of the time; the nodes with the first specific number of durations are assigned to the directory committee; the remaining nodes are assigned to different shards of the blockchain to be sharded according to the last preset number of bits of their respective identifiers.
[0121] Specifically, at time t, each node in the blockchain to be sharded participates in solving the Proof of Work (PoW) puzzle according to its computing power, determines the duration for each node to solve the PoW puzzle, and assigns a specific number of nodes that can solve the PoW puzzle faster to the Directory Committee (DC), while the remaining nodes are assigned to different shards according to the last preset number of bits of their respective identifiers (Identifier, ID) (for example, if the preset number is 2, the last two bits represent the shard number).
[0122] As Figure 3 shown, Figure 3 Fig. shows a blockchain system composed of N nodes and K shards. Within each shard, a block producer is designated as the primary node (faster at solving puzzles / randomly), responsible for packing transactions into blocks and performing the Practical Byzantine Fault Tolerance (PBFT) consensus algorithm for verification to defend against potential Byzantine node attacks.
[0123] After block verification is completed within the shard, the verified blocks are transmitted to the DC. As the primary node, the DC is responsible for centrally executing the PBFT consensus and undertaking the final block packing process to ensure the finality and consistency of the blocks. Once a block passes the PBFT consensus verification, the DC packs it into the final block and broadcasts it to the entire blockchain network. At the same time, the DC synchronizes the final block information with each shard node to ensure real-time update and synchronization of system data.
[0124] The blockchain sharding strategy optimization method based on model-based reinforcement learning in the present invention is essentially a sharding control algorithm (Model-Based Policy Optimization with Batch Sampling, MBPOBS) based on Model-Based Policy Optimization (MBPO), as Figure 4 shown, Figure 4 Fig. is a schematic structural diagram of a blockchain sharding strategy optimization system based on model-based reinforcement learning; this system controls the optimal sharding of the blockchain through two experience streams: the real experience stream and the simulated experience stream. The real experience stream refers to the experience obtained from real interactions within the blockchain, and these experiences are directly used to train the prediction model; while the simulated experience stream is generated by the prediction model and is used to learn the optimal strategy, which can reduce the number of real interactions and improve the sampling efficiency.
[0125] Specifically, in the blockchain environment, the state data st Input to the optimal policy network to determine the action data a t , and then the action data a can be executed under the state data s t . By means of sharding information and network monitoring, calculate the consensus inconsistency and malicious node rate, and then calculate the reward value r t and the state data s at the next moment t , so as to obtain the data quadruple (s t+1 , a t , r t , s t ), and store the data quadruple (s t+1 , a t , r t , s t ) in the experience replay buffer D t+1 . RL .
[0126] Randomly sample M state data M<s RL > in D and input them into the model policy network π t to obtain the corresponding action data * and calculate the consensus inconsistency and malicious node rate by means of sharding information and network monitoring, and then calculate the reward value r ′, and predict the state data s′ at the next moment through the prediction model f t . This process is iteratively calculated for T steps to generate a complete action trajectory. The construction process of this action trajectory can be iterated S times to generate S action trajectories, calculate the total return of each action trajectory, and select the top e action sequences with the highest total return as the optimal action trajectories and store them in the elite experience replay buffer D t+1 . D elite , D elite is used to train the model policy network π * .
[0127] Moreover, store the first action data and state data of the action sequence with the highest total return in the experience replay buffer D model . D model is used to train the optimal policy network
[0128] The process of the entire workflow is as Figure 5 shown. The process is as follows:
[0129] S501, Initialize three experience replay buffers D RAND , D RL , D mod el , and initialize the initial policy network and the prediction model.
[0130] S502 interacts with the blockchain system for G time steps to generate a series of data quadruples consisting of state, action, reward, and next state, and these data are subsequently stored in the experience replay buffer D RAND , and through D RAND train the prediction model.
[0131] S503, obtain the current state data of the blockchain in the system environment, input the current state data into the initial policy network to obtain action data, and then execute the action data in the blockchain and calculate the reward value.
[0132] S504, obtain the state data of the blockchain at the next moment, generate data quadruples, and store the data quadruples into D RL .
[0133] S505, randomly select a data quadruple from D RL and D RAND and obtain the optimal action trajectory by combining the state data in the data quadruple through the cross-entropy algorithm.
[0134] S506, store the initial state-action pair in the optimal action trajectory into D mod el .
[0135] S507, train the initial policy network through D mod el and train the prediction model through D RL and D RAND train the prediction model.
[0136] S508, clear D mod el .
[0137] S509, determine the initial policy network as the optimal policy network.
[0138] It should be noted that the above steps S503 - S508 need to be iterated N times, and the above steps S505 - S506 need to be iterated M times.
[0139] In the blockchain sharding system, the present invention adopts different prediction methods for three state values with different characteristics: (1) node computing power C t , (2) node - to - node transmission rate V t , (3) consensus history H t . Since C t and V t have randomness and continuity, the prediction model selects the Gaussian Processes Regression (GPR) model. And for H tThe update, since it is accurately calculated based on known state transition rules and the current state, we directly use the state transition rules in the system for prediction. This method not only ensures the accuracy of the consensus historical value but also avoids the uncertainty brought by randomness.
[0140] Gaussian process regression assumes that the observations of state C t and V t follow a Gaussian process, that is, the joint distribution of any finite number of observations follows a multivariate Gaussian distribution. Given the current state values of C t and V t the GPR model f will output a Gaussian distribution, whose mean and variance are determined by the mean function and covariance function respectively. The predicted next state is the mean of this predicted distribution. The covariance function uses a Radial Basis Function (RBF) kernel.
[0141] After making a decision, the remaining time within the decision interval will be used for planning. During this period, the cross-entropy algorithm generates a simulated experience stream based on the model. Since the behavior generation step (step S503) only involves one forward pass through the initial policy network its time cost is negligible compared to the length of the decision interval. Therefore, assuming there is enough time for planning. The planning process (step S505) can be implemented by calling Figure 6 the CEM algorithm outlined in
[0142] As Figure 6 shown, the implementation method of obtaining the optimal behavior trajectory by combining the state data in this data quadruple through the cross-entropy algorithm in step S505 above includes the following steps:
[0143] S601, Initialize the model policy network and the elite experience replay buffer.
[0144] S602, Initialize an array X of length S, as well as the trajectory w and the trajectory set W.
[0145] S603, Clear the elite experience replay buffer, X, w, and W.
[0146] S604, Use the state data as the initial state data.
[0147] S605, Input the current state data into the model policy network, obtain the behavior data, execute the behavior data and calculate the reward value; predict the state data at the next moment according to the prediction model; finally store the state data, behavior data, reward value, and the state data at the next moment into w.
[0148] S606, Calculate the total return of w, store the total return into X, and store w into W.
[0149] S607, the e trajectories with the highest total rewards in X, add the data corresponding to the e trajectories to the elite experience replay buffer, and use the action data and state data corresponding to the trajectory with the highest total reward as the initial state-action pair.
[0150] S608, update the model policy network through the elite experience replay buffer.
[0151] S609, select the optimal action trajectory from W and output it.
[0152] It should be noted that the above steps S603 - S608 need to be iteratively executed Z times, steps S604 - S606 need to be iteratively executed S times, and step S605 needs to be iteratively executed T times.
[0153] The core idea of the present invention is to use model-based reinforcement learning to find higher-quality sharding strategies, so as to more effectively solve the OBCS problem. In addition, MBPOBS innovatively integrates Gaussian process regression for accurate prediction of future states, and on this basis, performs imitation learning through the demonstrator of the cross-entropy algorithm to learn the best sharding strategy. Compared with MFRL, MBPOBS has faster efficiency, stronger generalization ability, and higher stability in learning the optimal strategy.
[0154] When applying the blockchain sharding strategy optimization method based on model-based reinforcement learning provided by the present invention, it is not necessary to execute according to Figure 1 the order of the steps shown. The specific execution order of each step can be determined as needed, and the present invention does not limit this.
[0155] The above is the blockchain sharding strategy optimization method based on model-based reinforcement learning provided by one or more embodiments of the present invention. Based on the same idea, the present invention also provides a corresponding blockchain sharding strategy optimization device based on model-based reinforcement learning, as Figure 7 shown.
[0156] Figure 7 is a structural block diagram of a blockchain sharding strategy optimization device based on model-based reinforcement learning provided by the present invention. The device 700 includes:
[0157] An acquisition module 701, configured to acquire the current state data of the blockchain, and input the current state data into the initial policy network to obtain the action data of the blockchain; the action data of the blockchain is used to guide blockchain sharding.
[0158] A state transition module 702, configured to perform state transition on the blockchain according to the action data of the blockchain to obtain the state data of the blockchain at the next moment.
[0159] A solution module 703, configured to obtain initial state data of a prediction model, input the initial state data into a cross-entropy algorithm, call the prediction model through the cross-entropy algorithm to predict state data at a future moment, and call a model policy network to predict behavior data corresponding to the state data, so as to obtain an optimal behavior trajectory of the blockchain; the optimal behavior trajectory includes state data and behavior data of the blockchain at multiple moments.
[0160] A training module 704, configured to obtain initial state-action pairs from the optimal behavior trajectory, and train an initial policy network through the initial state-action pairs to obtain an optimal policy network; the optimal policy network is used to generate an optimal sharding policy for the to-be-sharded blockchain.
[0161] For specific limitations on the blockchain sharding device, reference may be made to the limitations on the blockchain sharding policy optimization method based on model reinforcement learning in the above text, which will not be elaborated here. Each module in the above blockchain sharding policy optimization device based on model reinforcement learning can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in or independent of a processor in a computer device in the form of hardware, or stored in a memory in the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.
[0162] The present invention further provides a computer-readable storage medium, which stores a computer program, and the computer program can be used to execute the above Figure 1 provided blockchain sharding policy optimization method based on model reinforcement learning.
[0163] The present invention further provides Figure 8 a schematic structural diagram of the computer device shown in, as Figure 8 shown, at the hardware level, the computer device includes a processor, an internal bus, a network interface, a memory, and a non-volatile memory. Of course, there may also be other hardware required for other services. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to implement the above Figure 1 provided blockchain sharding policy optimization method based on model reinforcement learning.
[0164] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above various methods. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided by the present invention can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical memory, etc. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.
[0165] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered to be within the scope recorded by the present invention.
Claims
1. A method for optimizing blockchain sharding strategies based on model reinforcement learning, characterized in that The method includes: Obtain the current state data of the blockchain, and input the current state data into the initial policy network to obtain the behavior data of the blockchain; the behavior data of the blockchain is used to guide the sharding of the blockchain; Perform a state transition on the blockchain according to the behavior data of the blockchain to obtain the state data of the blockchain at the next moment; train based on the current state data, behavior data, and state data at the next moment of the blockchain to obtain a prediction model; Obtain the initial state data of the prediction model, and input the initial state data into the cross-entropy algorithm. Through the cross-entropy algorithm, call the prediction model to predict the state data at a future moment, and call the model policy network to predict the behavior data corresponding to the initial state data and the state data at the future moment respectively, to obtain the optimal behavior trajectory of the blockchain; the optimal behavior trajectory includes the state data and behavior data of the blockchain at multiple moments; Obtain the initial state behavior pairs from the optimal behavior trajectory, and train the initial policy network through the initial state behavior pairs to obtain an optimal policy network; the optimal policy network is used to generate the optimal sharding policy of the blockchain to be sharded.
2. The method according to claim 1, wherein The step of inputting the current state data into the initial policy network to obtain the behavior data of the blockchain includes: Input the current state data into the initial policy network to obtain the probability values of each discrete behavior of the blockchain; the discrete behaviors include the block size of the blockchain, the number of shards, and the time required to generate a block; Determine the behavior data of the blockchain according to the probability values of each discrete behavior.
3. The method according to claim 1 or 2, characterized in that, The step of calling the prediction model by the cross-entropy algorithm to predict the state data at a future moment, and calling the model policy network to predict the behavior data corresponding to the initial state data and the state data at the future moment respectively, to obtain the optimal behavior trajectory of the blockchain includes: Call the prediction model by the cross-entropy algorithm to predict the state data at a future moment and call the model policy network to predict the behavior data corresponding to the initial state data and the state data at the future moment respectively, to obtain multiple different behavior trajectories; each behavior trajectory includes the state data, behavior data, and corresponding reward values of the blockchain at multiple moments; Determine the total return value of each behavior trajectory according to the multiple reward values in each behavior trajectory; Determine the preset number of behavior trajectories with the highest total return value as the optimal behavior trajectory of the blockchain.
4. The method according to claim 3, wherein The step of calling the prediction model by the cross-entropy algorithm to predict the state data at a future moment and calling the model policy network to predict the behavior data corresponding to the initial state data and the state data at the future moment respectively, to obtain multiple different behavior trajectories includes: For any one behavior trajectory, input the initial state data into the model policy network to obtain the behavior data corresponding to the initial state data, and calculate the reward value of the behavior data; Input the initial state data into the prediction model to predict the state data at the next moment; Determine the behavior data and reward value at the next moment according to the predicted state data at the next moment; Determine the behavior trajectory according to the state data, behavior data and reward value at each moment.
5. The method according to claim 4, wherein Calculating the reward value of the behavior data includes: Determine whether the blockchain corresponding to the behavior data meets the preset constraint conditions; When the blockchain corresponding to the behavior data meets the constraint conditions, substitute the behavior data into a preset reward function to obtain the reward value; When the blockchain corresponding to the behavior data does not meet the constraint conditions, determine that the reward value is 0; The constraint condition is: ; Among them, N is the number of nodes in the blockchain, is the number of malicious nodes in the blockchain, is the time required to generate a block, is in the blockchain K is the total consensus time when there are shards, is the block interval.
6. The method according to claim 4, wherein The method further includes: Train the model policy network through the initial state behavior pair to update the model policy network.
7. The method according to claim 1 or 2, characterized in that Obtaining the initial state behavior pair from the optimal behavior trajectory includes: Obtain the behavior data corresponding to the initial state data in the optimal behavior trajectory; Determine the initial state behavior pair according to the initial state data and the corresponding behavior data.
8. The method according to claim 1 or 2, characterized in that, The construction process of the prediction model includes: Interact the blockchain with the initial policy network for multiple time steps to obtain multiple data quadruples, and store the multiple data quadruples in the first experience replay buffer; each data quadruple includes the current state data, behavior data, reward value and the state data at the next moment of the blockchain; Train the initial prediction model according to the multiple data quadruples in the first experience replay buffer to obtain the prediction model.
9. The method according to claim 8, wherein The method further includes: After obtaining the state data of the blockchain at the next moment, substitute the behavior data of the blockchain into the reward function to obtain the reward value, and store the current state data, behavior data, reward value and the state data at the next moment of the blockchain as a data quadruple in the second experience replay buffer; Train the prediction model through the data quadruples in the first experience replay buffer and the second experience replay buffer to obtain an updated prediction model; the updated prediction model is used for predicting the state data of the blockchain.
10. The method according to claim 9, wherein Obtaining the initial state data of the prediction model includes: Randomly select a state data of the blockchain from the first experience replay buffer and the second experience replay buffer as the initial state data of the prediction model.
Citation Information
Patent Citations
Intelligent block chain multi-fragment prediction and control method
CN114841070A
Industrial data storage-oriented overlapping self-organizing block chain fragmentation method
CN116886323A