Block chain address allocation method and system based on reinforcement learning
By adopting the Markov decision-making process of deep reinforcement learning in the sharded blockchain, optimizing address allocation, the problems of frequent cross-shash transactions and unbalanced loads are solved, and the performance and efficiency of sharded chains are improved.
Patent Information
- Application Number
- CN202410087618.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-22
- Publication Date
- 2025-07-22
AI Technical Summary
The existing technology cannot effectively predict future address distribution in sharded blockchain, resulting in frequent cross-shash transactions, affecting performance and load imbalance, and the existing heuristic solutions have poor results.
The Markov decision-making process based on deep reinforcement learning is adopted to receive transactions through transaction entrance services, and address allocation is used to use reinforcement learning models. Combining the state information of the transaction queue and shard chain, address allocation strategy is optimized, cross-shash transactions are reduced and load balancing.
Significantly reduce the cross-shash transaction ratio, improve the performance of shard chains, reduce computing and storage overhead, and realize load balancing, which is suitable for real shard chain systems.
Smart Images

Figure CN120358240A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the fields of information technology, blockchain, reinforcement learning, and deep learning, and particularly relates to a blockchain address allocation method and system based on reinforcement learning. Background Art
[0002] A narrow sense blockchain is a chain data structure formed by combining data blocks in a sequential connection manner according to the time sequence, and is a distributed ledger guaranteed by cryptography to be tamper-proof and non-forgeable. The general blockchain technology is a new distributed infrastructure and computing paradigm that uses a block chain data structure to verify and store data, uses a distributed node consensus algorithm to generate and update data, uses cryptography to ensure the security of data transmission and access, and uses smart contracts composed of automated script codes to program and operate data. Blockchains are divided into three types: public blockchains, consortium blockchains, and private blockchains. Public blockchains have no access mechanism and are completely decentralized, but their performance is average. Consortium blockchains have a certain access mechanism, and their performance is better than that of public blockchains, but they lose a certain degree of decentralization.
[0003] Reinforcement learning (RL) is a field in machine learning that emphasizes how to act based on the environment to obtain the maximum expected benefit. A common model is the standard Markov Decision Process (MDP). Reinforcement learning problems have application cases in fields such as information theory, game theory, and automatic control, and are used to explain the equilibrium state under conditions of bounded rationality, design recommendation systems, and robot interaction systems. The Markov Decision Process (MDP) is a mathematical model for sequential decision-making, used to simulate the stochastic strategies and rewards that an agent can achieve in an environment where the system state has the Markov property. It is widely used in various fields such as reinforcement learning, predictive analysis, and optimization, providing an effective method for decision-makers to make decisions in complex dynamic environments. In the field of machine learning, Transformer is a sequence model based on the attention mechanism, initially proposed by a research team at Google and applied to the machine translation task. It is mainly used in the fields of natural language processing (NLP) and computer vision (CV), and can process sequence data, such as translation and text summarization in natural language processing. Different from traditional recurrent neural networks (RNNs) and convolutional neural networks (CNNs), Transformer only uses the self-attention mechanism to process the input sequence and output sequence, so it can perform parallel computing, greatly improving the computing efficiency.
[0004] A sharded blockchain, simply referred to as a shard chain, is a blockchain network that uses sharding technology to process data. It divides the data in a large database into many small, manageable parts, and then stores the data shards on different servers respectively to reduce the data access pressure on each server, thereby improving the performance of the entire database system. Introducing sharding technology in a blockchain can solve scalability and latency problems. Each node in the blockchain only has a part of the data on the blockchain rather than all the information, so it can process more transactions simultaneously and improve the processing capacity of the network. There are various implementation methods for blockchain sharding technology. Technologically, it can be divided into network sharding, transaction sharding, state sharding, etc. Network sharding divides the entire blockchain network into multiple sub-networks to process different transactions in the network in parallel. Transaction sharding performs sharding on transactions to improve the processing capacity of the network. State sharding divides the blockchain state into multiple shards, and each node only stores part of the state information of the blockchain. In the present invention, the shard chain refers to state sharding.
[0005] "SkyChain: A Deep Reinforcement Learning-Empowered Dynamic Blockchain Sharding System". This article introduces a new type of dynamic blockchain sharding system - SkyChain. The system features the use of deep reinforcement learning technology to optimize the sharding and consensus mechanisms to improve the performance and scalability of the blockchain. By adopting deep reinforcement learning algorithms, SkyChain can adaptively adjust the sharding and consensus mechanisms under different network conditions to achieve optimal performance and security. "Sharded Blockchain for Collaborative Computing in the Internet of Things: Combined of Dynamic Clustering and Deep Reinforcement Learning Approach". This article mainly explores how to utilize the combination of sharded blockchain and deep reinforcement learning to achieve related issues of collaborative computing in the Internet of Things field. The proposed solution in the article combines the K-Means clustering algorithm with reinforcement learning technology to present a sharded blockchain solution based on dynamic clustering and deep reinforcement learning, which can effectively solve the collaborative computing problem in the Internet of Things environment and improve the overall performance and efficiency. The deficiencies in the above two articles are common. Firstly, the state space selection in their reinforcement learning parts is relatively simple and does not reflect the data states of each sharded chain. In addition, a bigger problem lies in the selection of actions, including adjusting the sharding reconfiguration interval, the number of shards, and the block size. In blockchain, these values cannot be arbitrarily modified. For example, increasing the number of transactions included in a block will increase the block size, thereby increasing the transaction transmission time, which may cause security attacks such as double-spending.
[0006] The patent "Method for Optimizing the Performance of a Blockchain Sharding System Combining Deep Reinforcement Learning" (CN115102867B) proposes a method for optimizing the performance of a blockchain sharding system by combining deep reinforcement learning (DRL). This method optimizes the blockchain sharding system through the DRL algorithm to achieve higher system performance and efficiency. It is worth mentioning that the technical route and design scheme of this patent are highly similar to SkyChain, so they have similar problems. The patent "Method for Adaptive Optimization of Blockchain Performance Based on Hierarchical Consensus and Reinforcement Learning" (CN115378788A) proposes a method for adaptive optimization of blockchain performance based on hierarchical consensus and reinforcement learning. This method combines hierarchical consensus and reinforcement learning technologies, and divides the nodes in the consensus process into a main consensus group and a sub-consensus group cluster through a network node hierarchical module. The sub-consensus group cluster includes several sub-consensus groups. Reinforcement learning is mainly used in this patent to adjust the parameters in the blockchain network, that is, the parameters composed of the block size, block generation time, and the number of nodes in the consensus group are used as the action space. However, its application scenario is only a single blockchain, not a sharded chain, and there are also problems with the setting of the action space of the agent in this patent, which cannot be used in a real system.
[0007] Blockchain is a highly dynamic distributed system. As new transactions and users emerge, the state of the blockchain changes over time. Users use addresses to represent their identities in the blockchain, which is a string containing numbers and letters. Due to the pseudo-anonymous nature of the blockchain and the fact that creating new addresses does not require cost, users tend to change their addresses regularly. This leads to an increasing number of new addresses being continuously added to the blockchain. Sharded blockchains exhibit more dynamics than general blockchains. Sharded blockchains are often affected by cross-shard transactions. When the participants in a transaction are distributed across different shards of a sharded blockchain, a special type of transaction needs to be processed, called a cross-shard transaction. A large number of cross-shard transactions will significantly affect the performance of the sharded blockchain because processing these transactions takes more time. Cross-shard transactions are common and frequent in sharded blockchains. Research shows that when there are more than 16 shards, more than 95% of the transactions are cross-shard transactions.
[0008] Since each shard is only responsible for processing transactions and storing a part of the entire blockchain, an inappropriate sharding strategy will lead to uneven distribution of transactions and addresses among shards, and thus result in a large number of cross-shard transactions. Therefore, reducing the number of cross-shard transactions is crucial for further improving the throughput of sharded blockchains. However, there is a trade-off between maintaining workload balance among shards and reducing cross-shard transactions in sharded blockchains. For example, placing all addresses on the same shard can eliminate cross-shard transactions, but it will regress to the situation of having only a single blockchain. How to reduce cross-shard transactions while maintaining load balance among shards has become a problem. This problem is called the Address Placement problem. Currently, existing work mainly solves it through heuristic schemes.
[0009] Heuristic solutions, such as Monoxide and Shard Scheduler, are the basic solutions adopted in early sharded blockchains. In these solutions, addresses are assigned to shards according to some rules, such as assigning addresses to the shard with the fewest addresses, which is simple but has a poor effect on solving the address placement problem. The reason why these heuristic schemes have a poor effect is that they cannot effectively predict the shard where the addresses that will have transaction relationships with the assigned addresses in the future are located. If the future address distribution can be effectively predicted, then the address placement problem can be solved more effectively. Currently, there are many works on analyzing transaction data on the blockchain. Through the analysis of a large amount of transaction data, these works have revealed this future address distribution feature to a certain extent. This feature is related to time because it takes into account future blocks. At the same time, this feature is also related to space because it reveals the distribution of addresses among different shards. Therefore, this patent refers to this feature as the temporal feature of blockchain transactions. Summary of the Invention
[0010] Address assignment in a dynamic sharded blockchain can be regarded as a sequential decision-making problem suitable for being solved by reinforcement learning methods because a block can be regarded as a transaction sequence containing many addresses. To solve the address assignment problem in a dynamic sharded blockchain, while considering the temporal feature of blockchain transactions and minimizing the computational overhead of blockchain nodes, the present invention proposes a blockchain address assignment method and system based on deep reinforcement learning.
[0011] The technical solution adopted by the present invention is as follows:
[0012] A blockchain address assignment method based on reinforcement learning, comprising the following steps:
[0013] Receiving a transaction submitted by a user through a transaction entry service and saving the transaction in a transaction queue sorted by timestamp;
[0014] The transaction entry service extracts a fixed number of transactions from the transaction queue at fixed time intervals and applies a reinforcement learning model to allocate the transactions to each shard chain;
[0015] Each shard chain processes the transactions allocated by the transaction entry service and counts the number of in-shard transactions and cross-shard transactions;
[0016] The transaction entry service updates the reinforcement learning model by using the number of in-shard transactions and cross-shard transactions counted by each shard chain.
[0017] Furthermore, the transaction entry service is a transaction entry service responsible by a third-party trusted entity or a smart contract.
[0018] Furthermore, the reinforcement learning model is a Markov decision process model, and its goal is to reduce the proportion of cross-shard transactions and ensure the load balance of transactions through address allocation; the reinforcement learning model includes an overall objective R(θ), a state space S, an action space A, a reward function R, and a state transition function P; the overall objective R(θ) is defined as follows:
[0019]
[0020] where T is the total number of steps of the task, and each step corresponds to a block; γ is a discount factor; r t is the reward for each step; represents the expected value.
[0021] Furthermore, the state space S includes a state observation, which represents the observation of the agent and is a 3×k-dimensional vector, represented as follows:
[0022] obs = [num_tx1,...,num_tx k ,
[0023] cross_tx1,...,cross_tx k ,
[0024] sender_pos1,...,sender_pos k ,
[0025] where k is the number of shards in the shard chain, num_txi represents the total number of transactions on the i-th shard, cross_tx i represents the number of cross-shard transactions on the i-th shard, and sender_posi represents the positions of all senders associated with new addresses in the current block; at the beginning of each reinforcement learning training step, based on the positions of the shards where the existing addresses are located, num_tx iand cross_tx i is initialized.
[0026] Furthermore, the action in the action space A is a k-dimensional vector, denoted as:
[0027] action = [a1, a2, …, a k
[0028] where action is a one-hot encoding; a i has a value of 0 or 1, indicating whether a new address is assigned to the i-th shard.
[0029] Furthermore, the reward function R is defined as follows:
[0030]
[0031]
[0032]
[0033]
[0034] where k i is the total number of transactions in the i-th shard; k acerage is the average value of all k i [m i is a k-dimensional vector used to represent whether the transaction load in the i-th shard exceeds the average value; NET is the number of shards in the sharded chain whose workload does not exceed the average number of transactions; λ is a weight parameter used to balance the proportion of cross-shard transactions and the weight of workload balance in the reward.
[0035] Furthermore, the state transition function P is defined as the probability of reaching a new state s t when taking an action a t under a given state s t+1 The state transition function P is updated as follows: After the transaction entry service selects a batch of transactions each time, the total number of transactions num_tx and the cross-shard transaction number cross_tx in the state vector are only changed according to the states in each shard at the beginning of each training step, the sender position sender_pos changes according to the actions of each agent in the block, and after the transaction entry service puts the address into the sharded chain, the agents therein update the sender_pos related to that address.
[0036] Furthermore, the training process of the reinforcement learning model includes:
[0037] Denote the mapping from state to action as π(a|s), which represents the probability of taking action a in state s. The goal of reinforcement learning is to learn the optimal policy π that maximizes the expected cumulative reward. * :
[0038]
[0039] where T represents the length of a training set, and γ ∈ [0, 1] is the discount factor, which is used to adjust the influence of future rewards on the current decision. represents the expected value;
[0040] Solve the optimal policy π using the value function V(s). * , and solve the optimal value function V through the Bellman equation using dynamic programming or iterative methods. * .
[0041] Furthermore, the reinforcement learning algorithm adopted by the reinforcement learning model is the PPO algorithm, which includes an actor network and a critic network.
[0042] Furthermore, use surrogate policy gradient to optimize the actor network.
[0043] Furthermore, use temporal difference TD to optimize the critic network.
[0044] Furthermore, the actor network and the critic network are implemented using a fully connected network or a Transformer network based on the attention mechanism.
[0045] A blockchain address allocation system based on reinforcement learning includes a transaction entry service module and shard chains; the transaction entry service module receives transactions submitted by the user side and saves the transactions in a transaction queue sorted by timestamp; the transaction entry service module extracts a fixed number of transactions from the transaction queue at fixed time intervals and allocates the transactions to each shard chain using the reinforcement learning model; each shard chain processes the transactions allocated by the transaction entry service module and counts the number of intra-shard transactions and cross-shard transactions; the transaction entry service module updates the reinforcement learning model using the number of intra-shard transactions and cross-shard transactions counted by each shard chain.
[0046] The beneficial effects of the present invention are as follows:
[0047] The present invention can significantly reduce the proportion of cross-shard transactions, ensure the balance of the workload, and ultimately improve the overall performance of the shard chain. In addition, the design scheme of this patent requires less storage space and computational overhead, reducing the burden on nodes. Finally, the Markov decision process designed by the present invention is more practical and can be applied to a real shard chain system. Description of the Drawings
[0048] Figure 1 It is a schematic diagram showing that an agent determines which block a new address is assigned to by reading transactions in a block.
[0049] Figure 2 It is an overall flowchart of sharded chain address allocation and transaction processing.
[0050] Figure 3 It is a flowchart of PPO algorithm training.
[0051] Figure 4 It is a diagram of a fully connected neural network architecture.
[0052] Figure 5 It is a diagram of a Transformer network architecture based on the attention mechanism. Detailed Implementation Manner
[0053] To make the above objects, features, and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below through specific embodiments and drawings.
[0054] Figure 1 It is a schematic diagram showing that an agent determines which block a new address is assigned to by reading transactions in a block. The present invention models the state allocation problem in a sharded chain through a Markov decision process, and uses a state allocation algorithm based on deep reinforcement learning in the sharded blockchain to reduce the number of cross-shard transactions, while ensuring the balance of transactions in each shard, and ultimately improving the overall performance of the sharded blockchain.
[0055] The overall process of the present invention is as Figure 2 shown. In the sharded chain involved in the present invention, the processing of transactions is completed by each sharded chain, and the allocation of addresses is completed by a dedicated third-party trusted entity or a transaction entry service of a smart contract. Specifically, the specific steps of address allocation and transaction completion are as follows:
[0056] Step 1: After a user's transaction is submitted, it is first sent to the transaction entry service.
[0057] Step 2: The transaction entry service saves the transaction in a transaction queue sorted by timestamp. Subsequently, at fixed time intervals, the transaction entry service extracts a fixed number of transactions from the transaction queue and applies a reinforcement learning model to allocate them. The allocation result means which shards these transactions are sent to. The allocation result is sent to different shards in the form of multiple transaction arrays.
[0058] In blockchains with an account model like Ethereum, accounts are represented by addresses, which are strings containing numbers and letters. Addresses are used to represent users' identities in the blockchain. Each transaction has a sender and a recipient. The sender is an existing address, while the recipient may be a new address. If the consensus node discovers that the recipient address is a new address, it creates a new state to store the relevant information of the account. Therefore, the generation of new addresses / accounts is closely related to transactions. In a sharded blockchain, the transaction carrying the new address is assigned to a particular shard, and a new address will be generated on that shard.
[0059] Step 3: Each sharded blockchain receives the transactions sent by the transaction entry service and processes them. During the processing, each shard counts the types of transactions it processes. Transactions are divided into intra-shard transactions and cross-shard transactions. The block-producing nodes in the sharded blockchain will count the quantities of these two types of transactions and record the results in the block.
[0060] Step 4: The transaction entry service observes the blocks from each sharded blockchain and counts information such as cross-shard transactions and total transaction quantities. This information will be used in the training of the reinforcement learning model.
[0061] Step 5: The transaction entry service uses the data observed in Step 4 to update the reinforcement learning model (referred to as the agent model in Figure 2 ).
[0062] The above Steps 2, 4, and 5 involve making address allocation decisions using the reinforcement learning model. The design, training process, and network architecture of the reinforcement learning model involved in the present invention are as follows:
[0063] I. Reinforcement learning model design:
[0064] The Markov Decision Process (MDP) provides a powerful mathematical framework that captures the essence of sequential decision-making under uncertainty. The MDP model is particularly suitable for the address allocation task in sharded blockchains because it can represent the state space, action space, state transition, and reward structure of the problem as a tuple (S, A, P, R), where:
[0065] 1. S represents the state space, which is the set of all possible states.
[0066] 2. A represents the action space, which is the set of all possible actions.
[0067] 3. P represents the state transition probability (or state transition function), where P(s′|s,a) represents the probability of transitioning to state s′ after taking action a in state s.
[0068] 4. R represents the Reward function, where R(s, a, s′) represents the reward obtained after taking action a in state s and transitioning to state s′.
[0069] In the context of the sharded blockchain address allocation task, the objective of the present invention is to reduce the proportion of cross-shard transactions and ensure load balancing of transactions. To solve the address allocation problem, the present invention defines it within the MDP framework, which requires specifying the overall objective R(θ), state space S, action space A, reward function R, and state transition function P.
[0070] The overall objective R(θ) is defined as follows:
[0071]
[0072] where T is the total number of steps of the task, with each step corresponding to a block. γ is the discount factor, and r t is the reward for each step, represents the expected value.
[0073] 1) State, S
[0074] Different from previous work, the state space in the present invention considers the current situation in each shard. In addition, the state space should not use as much information as the solution based on the transaction graph, because the transaction graph usually contains tens of thousands of nodes (addresses) and more edges (transactions). Such a transaction graph tends to occupy a large amount of space and also requires a large amount of time to process. When the transaction entry service selects a batch of transactions, the agent in the transaction entry service can use the information in the block, such as the sender location of each receiver, and the situation in all sharded chains. The situation in the sharded chains includes the distribution of the number of transactions and the number of cross-shard transactions.
[0075] Observation is usually used to emphasize a subset of the environmental information where the agent is located, that is, the state that the agent can perceive. In the address allocation task, the state of the environment is equivalent to the state observed by the agent. In the present invention, in order to more accurately describe the state space, the observation of the agent is used as a substitute for the environmental state. Considering the above, in order to reduce the cross-shard transaction rate and avoid excessive imbalance in all shards, the state observation is a 3×k-dimensional vector, which is expressed as follows:
[0076] obs = [num_tx1,...,num_tx k ,
[0077] cross_tx1,...,cross_tx k ,
[0078] sender_pos1,...,sender_pos k ,
[0079] where k is the number of shards in the shard chain, num_tx i represents the total number of transactions on the i-th shard, cross_tx i represents the number of cross-shard transactions on the i-th shard, sender_pos i represents the positions of all senders associated with new addresses in the current block. Since the senders and receivers of some transactions have been processed in previous transactions, these transactions have been placed into their respective shards. At the start of each reinforcement learning training step, based on the positions of the shards where the existing addresses are located, num_tx i and cross_tx i are initialized. The above information can well represent the situation in each shard and each new address without imposing a heavy storage burden.
[0080] 2) Action, A
[0081] The principle of designing the action is that this design should be feasible in a real sharded blockchain. Adjusting the block size or shard reconfiguration interval is not appropriate because these are hyperparameters in a real blockchain system and cannot be easily changed during runtime. Changing these parameters requires consensus. The action set in the present invention is a k-dimensional vector, expressed as:
[0082] action = [a1, a2,..., a k
[0083] action is a one-hot encoding, where there is only one 1. In the action vector, a i has a value of 0 or 1, indicating whether to allocate a new address to the i-th shard. The action space in the present invention is concise and straightforward, directly solving the address placement problem.
[0084] 3) Reward, R
[0085] The reward, or return, is used to guide how the agent behaves to achieve the design goal of the present invention. Therefore, the present invention defines two metrics to represent the cross-shard ratio and the workload balance across all shards. The reward function R is defined as follows:
[0086]
[0087]
[0088]
[0089]
[0090] where k i is the total number of transactions in the i-th shard, and k average is the average value of all k i . [m i is a k-dimensional vector used to represent whether the transaction load in the i-th shard exceeds the average value. NET is the number of shards in the sharded chain whose workload does not exceed the average number of transactions, and λ is a weight parameter used to balance the proportion of cross-shard transactions and the weight of workload balance in the reward.
[0091] 4) State transition, P
[0092] The state transition function P defines the probability of reaching a new state s t after taking an action a t in a given state s t+1 . The state vector is updated by performing an action, that is, which shard to select to allocate a new address. The function P depends on the current timing characteristics of the blockchain and is an attribute of the environment.
[0093] Since the agent continuously processes the arriving transactions, the state transition function can be updated in the following way: After the transaction entry service selects a batch of transactions each time, the total number of transactions num_tx and the cross-shard transaction number cross_tx in the state vector are changed only according to the states in each shard at the beginning of each training step. For the sender position sender_pos, it changes according to the actions of each agent within the block. After the transaction entry service puts the address into the sharded chain, the agents in it will update the sender_pos related to that address.
[0094] II. Design of the training process of the reinforcement learning model:
[0095] For deep reinforcement learning, a policy is a mapping from states to actions, denoted as π(a|s), representing the probability of taking an action a in state s. The goal of reinforcement learning is to learn the optimal policy π that maximizes the expected cumulative reward * :
[0096]
[0097] where T represents the length of a training set (i.e., the number of steps T, that is, the length of the training set is measured by the number of steps), γ ∈ [0,1] is the discount factor used to adjust the influence of future rewards on the current decision, represents the expected value.
[0098] To solve the optimal policy, the present invention uses a value function, which satisfies the Bellman equation, and its derivation is as follows:
[0099]
[0100] Among them, V(s) represents the value function value corresponding to state s, and V(s’) represents the value function value corresponding to state s’. represents the expected value when the action is a. represents the expected value when the state is s’.
[0101] Through the Bellman equation, the optimal value function V can be solved by using dynamic programming or iterative methods * .
[0102] Deep reinforcement learning usually uses neural networks to represent policies and value functions, allowing the agent to handle complex state spaces and action spaces. The main reinforcement learning algorithm used in the present invention is PPO (Proximal Policy Optimization). PPO is a deep reinforcement learning algorithm based on policy gradients and the actor-critic framework. It consists of an actor network and a critic network. The agent uses the existing policy π θ (actor network) to interact with the environment and collect a batch of data. Once a complete batch of data is obtained, the actor network and the critic network will learn from the sampled data according to the Figure 3 process. Specifically, the surrogate policy gradient is used to optimize the actor network as follows:
[0103]
[0104]
[0105] A t =δ t +(γ·λ)·δ t+1 +…+(γ·λ) T-t ·δ T ,
[0106] δ t =r t +γ*V(s t+1 )-V(s t ),
[0107] Among them, A t is the generalized advantage estimate, which measures taking action a t under state s tRelative advantage compared to the average case; t represents a time step (i.e., a training moment); L(θ) represents the objective function of the PPO algorithm, and θ represents the parameters of the actor network; represents the expected value; ratiot represents the ratio (taking the log of the new and old test values, so the final calculation method is a ratio, or quotient); clip represents the clip function, which is used to limit the value range of the first parameter between the second and third parameters, i.e., [1 - ∈, 1 + ∈]; ∈ represents the clipping ratio, which is used to determine the upper and lower bounds of the clip function; π θ represents the policy function of the actor network; represents the policy function of the old actor network; δ t represents the TD error at time step t; T represents the total number of training steps; γ represents the discount factor of future rewards; λ is a parameter used to control the trade-off between the bias and variance of the advantage estimate; V(s t+1 ) represents the value function value of state s t+1 ; V(s t ) represents the value function value of state s t .
[0108] The critic network can be optimized using Temporal Difference (TD):
[0109]
[0110]
[0111] where L critic represents the objective function of the critic network, represents the expected value at time step t, V target represents the target network value function, represents the expected value corresponding to the action a taken at time step t, represents the expected value corresponding to the state at time step t + 1.
[0112] By using the above objective function, the parameters of the actor network and the critic network can be continuously updated in multiple episodes, thereby improving the performance of the agent in the environment.
[0113] III. Deep learning network architecture
[0114] In the present invention, in order to achieve a balance between the method effect and the computational cost, two different types of neural network architectures are designed to implement the above-mentioned actor network and critic network. One of them is a fully connected network, and the other is a Transformer network based on the attention mechanism.
[0115] (1) Fully connected network. In the present invention, a feedforward neural network with three linear layers and two ReLU activation functions is selected in the fully connected network. The input layer contains n D neurons. Both of these hidden linear layers have an n neuron -dimensional feature space, and the output layer maps n neuron neurons to out D neurons. The ReLU activation function is used between each linear layer. The structure of these layers is shown as Figure 4 follows:
[0116] 1. Input layer: The dimension of this layer is equal to the state space dimension (n D ), which is the state observation of the agent.
[0117] 2. Hidden layer 1: A fully connected layer followed by a ReLU activation function.
[0118] 3. Hidden layer 2: Another fully connected layer also followed by a ReLU activation function.
[0119] 4. Output layer: A fully connected layer whose dimension is equal to the action space dimension of the actor network (out D ), or a single output neuron of the critic network.
[0120] (2) Transformer network based on the attention mechanism. Since the fully connected network treats all addresses in the selected transactions of the trading entry service equally, it is difficult to extract and use important information from the operations of the previously assigned addresses that contain rich recent temporal features in the current block. Since Transformer has shown strong temporal modeling ability in the field of time series prediction, the present invention adopts it to further improve the effectiveness of the agent in utilizing the temporal features of transactions. The structure of Transformer is shown as Figure 5 follows.
[0121] The input of Transformer is X, where X ∈ R N×C×D . N represents the batch size. C represents the length of the sequence, which is the state in the example of the present invention. D is the input dimension. The architecture of this model starts from an input layer that accepts data with the dimension specified by n D . First, the following positional encoding is performed on the input data:
[0122] X′ = X + PositionalEncoding(X),
[0123] where the calculation formula for each element (i, j) is:
[0124]
[0125] Next, X' is passed to Y ∈ R by the Transformer encoder N×C×D . The Transformer encoder layer consists of 2 layers. Each of them consists of two main sub-layers: a multi-head self-attention mechanism with n head heads and a position-wise feed-forward network ( Figure 5 the feed-forward layer in neuron ). The hidden size of the feed-forward network is set to n
[0126] Y = TransformerEncoder(X').
[0127] Specifically, the Transformer encoder integrates a multi-head self-attention mechanism. For each head h, there is a learned weight set W Q h , W K h , W V h used to calculate the query, key, and value. For each input X', its query, key, and value are first calculated. In the Transformer model, the attention representation is calculated by assigning a weight to each position in the input sequence to calculate the representation of the input sequence. This weight is calculated through the self-attention mechanism, which treats each position in the input sequence as a query, key, and value to calculate their similarity. Then, by applying the weight of each position to the elements in the input sequence, a global representation can be calculated, which can capture the important features and dependencies in the input sequence.
[0128]
[0129] where, Q h represents the query vector corresponding to the head labeled h in the multi-head attention mechanism; K h represents the key vector corresponding to the head labeled h in the multi-head attention mechanism; V h represents the value vector corresponding to the head labeled h in the multi-head attention mechanism.
[0130] When the query, key, and value of each head are calculated, the attention mechanism representation is calculated as:
[0131]
[0132] After that, the present invention connects all the attention representations of the heads and calculates them through a linear layer:
[0133] MultiHead(Q, K, V) = concat(Attention1,...,\(Attention_{h}\)) \(W^{O}\) H ) \(W^{O}\) O .
[0134] where \(W^{O}\) o represents the weight matrix learned by the neural network and is used for the calculation of the linear layer.
[0135] After obtaining the output of the multi-head attention mechanism, the present invention passes this result matrix \(Y\) through a feed-forward neural network to generate a new output, denoted as \(Z\in\mathbb{R}^{N\times d_{model}}\) N×C×D
[0136] \(Z = FFN(Y)\),
[0137] where \(FFN(Y)\) is a feed-forward neural network composed of two linear layers and a ReLU activation function. Subsequently, \(Z\) takes the average value in the time dimension to obtain \(Z'\in\mathbb{R}^{N\times 1}\) N×D
[0138]
[0139] where \(C\) represents the length of the sequence.
[0140] Finally, \(Z'\) is transformed into the output dimension through a linear layer to obtain the final output \(action\):
[0141] \(action = Linear(z')\),
[0142] In summary, the present invention proposes an address allocation algorithm based on deep reinforcement learning. It first models the Markov decision process of the address placement problem and applies the attention mechanism to dynamic sharding. The present invention can reduce the proportion of cross-shard transactions without overly affecting the workload balance between different shards. In addition, the method of the present invention utilizes the temporal characteristics of transaction data, thereby improving the throughput of the entire system.
[0143] The main innovation points of the present invention include:
[0144] 1. Using deep reinforcement learning to solve the address allocation problem in the context of sharded chains. The present invention first models the Markov decision process of the address allocation problem in sharded chains. Existing work has not systematically modeled this problem in sharded chains, so it is impossible to solve the address allocation problem in sharded blockchains.
[0145] 2. The present invention first utilizes the temporal characteristics of transaction data in the Markov modeling process of reinforcement learning. The temporal characteristics can be used to reveal the future distribution of addresses, so as to allocate more relevant addresses to the same shard, thereby reducing cross-shard transactions. In addition, by using the temporal characteristics, the balance of transaction allocation can also be taken into account, thus improving the performance of the sharded blockchain as a whole.
[0146] 3. When designing the state space of the agent in the sharded blockchain for the first time, the present invention takes into account the state of each specific shard. For the agent, understanding the state of each shard in the sharded chain system helps the agent make more reasonable decisions. The existing solutions consider all the sharded chains as a whole, which cannot achieve good results.
[0147] 4. The present invention first considers the feasibility of combining the design of the agent's action space with the sharded chain. In the existing solutions, the actions of the agent are mostly to adjust the block size and the interval occurring during the shard reconfiguration phase. These parameters are all related to the consensus of the blockchain, and any modification may cause potential security problems and is difficult to apply in a real system.
[0148] 5. The attention mechanism is introduced into the scenario of the sharded blockchain for the first time. Currently, the reinforcement learning technology applied in the sharded chain is mainly based on a fully connected neural network. This neural network architecture is relatively simple and cannot effectively utilize the temporal characteristics of the input data. The present invention first introduces the attention mechanism into the sharded blockchain, improving the utilization degree of the temporal characteristics of the blockchain transaction data, and thus enhancing the solution effect for the address allocation problem.
[0149] 6. A reasonable framework for combining reinforcement learning with the sharded blockchain is proposed. By unifying the transaction entry and the deployment of the agent into the transaction entry service, users can complete the allocation of transactions by the agent and the update of the agent without perceiving the existence of the entry service.
[0150] Another embodiment of the present invention provides a blockchain address allocation system based on reinforcement learning, which includes a transaction entry service module and a sharded chain; the transaction entry service module receives transactions submitted by the user terminal and saves the transactions in a transaction queue sorted based on timestamps; the transaction entry service module extracts a fixed number of transactions from the transaction queue at fixed time intervals and applies a reinforcement learning model to allocate the transactions to each sharded chain; each sharded chain processes the transactions allocated by the transaction entry service module and counts the number of in-shard transactions and cross-shard transactions; the transaction entry service module updates the reinforcement learning model by using the number of in-shard transactions and cross-shard transactions counted by each sharded chain. Among them, the transaction entry service module can be a separate computer or server.
[0151] The specific embodiments of the present invention disclosed above are intended to help understand the content of the present invention and implement it accordingly. Those of ordinary skill in the art can understand that various substitutions, changes, and modifications are possible without departing from the spirit and scope of the present invention. The present invention should not be limited to the content disclosed in the embodiments of this specification, and the protection scope of the present invention shall be subject to the scope defined by the claims.
Claims
1. A blockchain address allocation method based on reinforcement learning, characterized in that, It includes the following steps: Receive the transactions submitted by users through the transaction entry service, and save the transactions in a transaction queue sorted by timestamp; The transaction entry service extracts a fixed number of transactions from the transaction queue at fixed time intervals, and applies a reinforcement learning model to allocate the transactions to each shard chain; Each shard chain processes the transactions allocated by the transaction entry service, and counts the number of in-shard transactions and cross-shard transactions; The transaction entry service updates the reinforcement learning model using the number of in-shard transactions and cross-shard transactions counted by each shard chain.
2. The method according to claim 1, wherein The transaction entry service is a transaction entry service responsible by a third-party trusted entity or a smart contract.
3. The method according to claim 1, characterized in that, The reinforcement learning model is a Markov decision process model, and its goal is to reduce the proportion of cross-shard transactions through address allocation and ensure the load balance of transactions; the reinforcement learning model includes an overall objective R(θ), a state space S, an action space A, a reward function R, and a state transition function P; the overall objective R(θ) is defined as follows: where T is the total number of steps of the task, with each step corresponding to a block; γ is the discount factor; r t is the reward for each step; represents the expected value.
4. The method according to claim 3, characterized in that The state space S includes a state observation, representing the observation of the agent, which is a 3×k-dimensional vector and is represented as follows: obs = [num_tx1,..., num_tx k , cross_tx1,..., cross_tx k , sender_os1,..., sender_pos k , where k is the number of shards in the shard chain, num_txi represents the total number of transactions on the i-th shard, cross_tx i represents the number of cross-shard transactions on the i-th shard, and sender_posi represents the positions of all senders associated with the new address in the current block; at the start of each reinforcement learning training step, based on the position of the shard where the existing address is located, num_tx i and cross_tx i are initialized; The action in the action space A is a k-dimensional vector, represented as: action = [a1, a2,..., a k Among them, action is a one-hot encoding; the value of a i is 0 or 1, indicating whether a new address is assigned to the i-th shard; The reward function R is defined as follows: where k i is the total number of transactions in the i-th shard; k average is the average value of all k i ; [m i is a k-dimensional vector used to represent whether the transaction load in the i-th shard exceeds the average value; NET is the number of shards in the sharded chain where the workload does not exceed the average number of transactions; λ is a weight parameter used to balance the proportion of cross-shard transactions and the weight of workload balance in the reward; The state transition function P is defined as the probability of taking action a t and reaching a new state s t given the state s. t+1 The state transition function P is updated as follows: After a batch of transactions is selected at each transaction entry service, the total number of transactions num_tx and the number of cross-shard transactions cross_tx in the state vector are changed only at the beginning of each training step according to the states in each shard. The sender position sender_pos changes according to the actions of each agent within the block. After the transaction entry service places the address into the shard chain, the agents therein update the sender_pos associated with that address.
5. The method according to claim 3, wherein The training process of the reinforcement learning model includes: Denote the mapping from state to action as π(a|s), which represents the probability of taking action a in state s. The goal of reinforcement learning is to learn the optimal policy π that maximizes the expected cumulative reward. * : where \(T\) represents the length of a training set, and \(\gamma\in[0,1]\) is the discount factor, which is used to adjust the influence of future rewards on the current decision, represents the expected value; Solve the optimal policy π using the value function V(s) * , and solve the optimal value function V by using the Bellman equation and adopting dynamic programming or iterative methods * .
6. The method according to claim 5, wherein The reinforcement learning algorithm adopted by the reinforcement learning model is the PPO algorithm, including an actor network and a critic network.
7. The method according to claim 6, wherein Optimize the actor network using surrogate policy gradients as follows: A t = δ t + (γ·λ)·δ t+1 + … + (γ·λ) T - t·δ T , δ t = r t + γ * V(s t+1 ) - V(s t ), where At is the generalized advantage estimate, measuring the relative advantage of taking action a t in state s t relative to the average case; t represents a certain time step; L(θ) represents the objective function of the PPO algorithm; denotes the expected value; clip represents the clip function, which is used to limit the value range of the first parameter between the second and third parameters, i.e., [1 - ∈, 1 + ∈], where ∈ represents the clipping ratio; π θ represents the policy function of the actor network; represents the policy function of the old actor network; δ t denotes the TD error at time step t; T represents the total number of training steps; γ represents the discount factor of future rewards; λ is a parameter used to control the trade-off between the bias and variance of the advantage estimate; V(s t+1 ) represents the value function value of state s t+1 , and V(s t ) represents the value function value of state s t .
8. The method according to claim 7, wherein Optimize the critic network using temporal difference TD as follows: Among them, L critic represents the objective function of the critic network, represents the expected value at time step t, V target represents the target network value function, represents the expected value corresponding to the action a taken at time step t, represents the expected value corresponding to the state at time step t + 1.
9. The method according to claim 6, wherein The actor network and the critic network are implemented using a fully connected network, or implemented using a Transformer network based on an attention mechanism.
10. A blockchain address allocation system based on reinforcement learning, characterized in that, It includes a transaction entry service module and shard chains; the transaction entry service module receives the transactions submitted by the user side and saves the transactions in a transaction queue sorted by timestamp; the transaction entry service module extracts a fixed number of transactions from the transaction queue at fixed time intervals and applies a reinforcement learning model to allocate the transactions to each shard chain; each shard chain processes the transactions allocated by the transaction entry service module and counts the number of in-shard transactions and cross-shard transactions; the transaction entry service module updates the reinforcement learning model using the number of in-shard transactions and cross-shard transactions counted by each shard chain.
Citation Information
Patent Citations
A Performance Optimization Method for Blockchain Sharding Systems Combining Deep Reinforcement Learning
CN115102867B
Block chain performance adaptive optimization method based on hierarchical consensus and reinforcement learning
CN115378788A
Cited By
Fragmented block chain state distribution method
CN120743549A