Transaction processing method and system for reinforcement learning and fragmentation block chain integration
By introducing the architecture of agent sharding and transaction sharding in the shard chain, combining the PBFT protocol and relay model, the decentralized deployment problem of reinforcement learning in the shard chain is solved, the system performance and security are improved, and it is suitable for existing and future shard chain technologies.
Patent Information
- Application Number
- CN202410358962.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-01-29
- Filing Date
- 2024-03-27
- Publication Date
- 2025-07-29
AI Technical Summary
The existing reinforcement learning and sharded chain combination schemes fail to effectively solve the problems of decentralized deployment and training, making it difficult to apply in actual scenarios.
The architecture of agent sharding and transaction sharding is adopted, combined with Byzantine fault-tolerant PBFT protocol and relay model, decentralized reinforcement learning training and deployment methods are realized, and transaction processing is optimized through the consensus stage and the reconstruction stage.
It improves the performance and efficiency of shard chains, while ensuring the security and privacy protection of the system, and is suitable for existing and future shard chain technologies.
Smart Images

Figure CN120387890A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of blockchain, and particularly relates to a transaction processing method and system for integrating reinforcement learning and sharded blockchain. Background Art
[0002] Narrow blockchain is a chain data structure formed by combining data blocks in sequence according to time order, and is a distributed ledger that is tamper-proof and forgery-proof guaranteed by cryptography. Generalized blockchain technology is a new distributed infrastructure and computing paradigm that uses a block chain data structure to verify and store data, uses a distributed node consensus algorithm to generate and update data, uses cryptography to ensure the security of data transmission and access, and uses smart contracts composed of automated script codes to program and operate data. Blockchain is divided into three types: public chain, consortium chain, and private chain. The public chain has no admission mechanism and is completely decentralized, but its performance is average. The consortium chain has a certain admission mechanism, and its performance is better than that of the public chain, but it loses a certain degree of decentralization.
[0003] The PBFT protocol, namely the Practical Byzantine Fault Tolerance system, is a fault tolerance mechanism in a distributed system implemented based on the BFT (Byzantine Fault Tolerance) algorithm. The principle of the PBFT protocol is based on the concept of a state replication machine, which divides the communication and state synchronization between nodes into multiple steps such as preprocessing, primary witness selection, request processing, and result return. Through multiple rounds of voting and confirmation mechanisms, the consistency of each block in the blockchain is ensured. Specifically, the core idea of the PBFT protocol is that all nodes jointly maintain a state, and all nodes take consistent actions. In the PBFT system, nodes are divided into two types: primary nodes (Lead) and secondary nodes (Follow). There is only one primary node, and all other remaining nodes are secondary nodes. Before the PBFT system function is implemented, all nodes are secondary nodes, and after an election, a primary node is generated. Each node has the right to be elected and the right to vote, and the probability of each node in the election process is equal. The operation of the PBFT protocol is divided into multiple stages, including sequence allocation (Pre-prepare), mutual interaction (Prepare), and confirmation (Commit). In the sequence allocation stage, the primary node broadcasts the request and its own signature to all nodes. In the mutual interaction stage, nodes will communicate multiple times, exchange messages, and reach an agreement. In the confirmation stage, nodes will confirm or reject the received messages.
[0004] Reinforcement Learning (RL) is a field in machine learning that emphasizes how to act based on the environment to maximize the expected benefit. A common model is the standard Markov Decision Process (MDP). Reinforcement learning problems have been discussed in fields such as information theory, game theory, and automatic control, and are used to explain equilibrium states under bounded rationality, design recommendation systems, and robot interaction systems. The Markov Decision Process (MDP) is a mathematical model for sequential decision-making, used to simulate the stochastic policies and rewards achievable by an agent in an environment where the system state has the Markov property. It is widely applied in various fields such as reinforcement learning, predictive analytics, and optimization, providing an effective method for decision-makers to make decisions in complex dynamic environments.
[0005] A sharded blockchain, simply referred to as a shard chain, is a blockchain network that uses sharding technology to process data. It divides the data in a large database into many small, manageable parts, and then stores the data shards on different servers respectively to reduce the data access pressure on each server, thereby improving the performance of the entire database system. Introducing sharding technology in the blockchain can solve scalability and latency problems. Each node in the blockchain only has a part of the data on the blockchain rather than all the information, so it can process more transactions simultaneously, improving the processing capacity of the network. There are various implementation methods for blockchain sharding technology, which are classified into network sharding, transaction sharding, and state sharding according to technology. Network sharding divides the entire blockchain network into multiple sub-networks to process different transactions in the network in parallel. Transaction sharding shards transactions to improve the processing capacity of the network. State sharding divides the blockchain state into multiple shards, and each node only stores part of the state information of the blockchain. In this article, the shard chain refers to state sharding.
[0006] A Merkle Tree is a hash-based data structure that can be a binary tree or a multi-way tree. The Merkle Tree is a tree-like structure where each leaf node is the hash value of a data block, and non-leaf nodes are the hashes of the hashes of their child nodes. The value on each node of the Merkle Tree is calculated according to the values of the leaf nodes below it using a certain algorithm. In Bitcoin and other cryptocurrencies, the Merkle Tree is used to more efficiently and securely encode blockchain data. By using the Merkle Tree, the present invention can effectively calculate and compare the hash values of data without the need to access all the data. At the same time, it generates a unique identifier (Merkle root) to represent the entire data set, which enables the present invention to prove whether a certain data block belongs to a certain data set without providing the entire data set. The Merkle Tree has many applications in computer science and cryptography, especially in the Bitcoin network, where it is used to summarize all transactions in a block and generate a digital fingerprint of the entire transaction set, providing an effective way to verify whether a certain transaction exists in a block.
[0007] Currently, reinforcement learning techniques have been widely applied in the field of sharded blockchains to improve their performance. However, when applying reinforcement learning to sharded blockchains, the key problem to be solved is how to design the sharded blockchain system architecture to adapt to the decentralized deployment and training of agents. The solution to this key problem determines the feasibility of applying reinforcement learning techniques on sharded blockchains and belongs to a fundamental key problem. Existing solutions that combine reinforcement learning and blockchains mostly use simulation-based methods. These methods can demonstrate the role of reinforcement learning in sharded blockchains, but do not consider the difficulties that may be faced in actual deployment. Specifically, simulation-based methods usually assume that agents in reinforcement learning are centrally deployed and trained. In traditional reinforcement learning, the deployment and training of agents are both completed by a central node or cluster. However, sharded blockchains require a decentralized environment, so the traditional centralized deployment method is no longer applicable to sharded blockchains. Although these simulation-based reinforcement learning deployment solutions have certain effects, they are difficult to apply to real scenarios. In the related work of combining reinforcement learning and sharded blockchains, some studies have proposed some decentralized reinforcement learning training and agent deployment solutions. However, these solutions do not combine with the original consensus mechanism in sharded blockchains, resulting in unnecessary complexity in the process. Summary of the Invention
[0008] To solve these problems, the present invention proposes a decentralized reinforcement learning training and deployment method that can be combined with a sharded blockchain system. This method enables reinforcement learning to be more effectively combined with sharded blockchains and to be deployed and trained under the premise of decentralization.
[0009] The technical solution adopted by the present invention is as follows:
[0010] A transaction processing method for integrating reinforcement learning and sharded blockchain, comprising the following steps:
[0011] Establish a sharded chain, which includes an agent shard and a transaction shard;
[0012] In the consensus stage, the consensus nodes in the agent shard send the corresponding transactions to the transaction shard by running a reinforcement learning algorithm;
[0013] In the reconstruction stage, the agent shard completes the reconstruction of the transaction shard according to the situation in each transaction shard.
[0014] Further, the number of agent shards is 1, and the number of transaction shards is k, where k is an integer greater than 1; the transaction shard is responsible for processing the transactions distributed to its own shard, and each transaction shard is only responsible for storing 1 / k of the users in the entire blockchain ledger; the agent shard serves as a bridge connecting the transaction shard and the user side, and users do not have direct contact with the transaction shard.
[0015] Further, both the agent shard and the transaction shard adopt the PBFT protocol based on Byzantine fault tolerance as the in-shard consensus protocol, and adopt a relay-based cross-shard transaction processing model to implement cross-shard transactions.
[0016] Further, the in-shard consensus protocol adopted by the agent shard and the transaction shard includes:
[0017] In each round of consensus, the leader of the agent shard selects transactions from its transaction pool, generates the allocation result of the transactions in the form of k transaction batches, and after reaching a consensus on the allocation result of the transactions, the transaction batches will be sent to the corresponding transaction shards;
[0018] The transaction shard verifies and executes the transactions, incorporates the valid transactions into the block, and records the new addresses corresponding to the new states and the number of cross-shard transactions in the new block;
[0019] The consensus nodes in the agent shard record the distribution of addresses, the number of cross-shard transactions, and the total number of transactions in each transaction shard for training the agent, and the consensus nodes in the agent shard record the relevant information in the entire consensus stage for the agent decision-making in the reconstruction stage.
[0020] Furthermore, the agent shard adopts a decentralized agent deployment and training method. Each consensus node in the agent shard maintains a copy of the agent with the same initial parameters. The hash value of the genesis block in the agent shard is used as a random seed for all agents. The update of the agent and the state placement operations of all agents are determined through a consensus process. In the consensus stage, a consensus node in the agent shard is selected as the leader based on epoch randomness and proposes a new block.
[0021] Furthermore, the data structure of the block in the agent shard includes a block header and a block body; the block header includes metadata, votes from other consensus nodes, and the root node and leaf nodes of the Merkle Patricia tree, and the leaf nodes of the Merkle Patricia tree store the mapping from the address to the transaction shard ID.
[0022] Furthermore, the agent sharding implements decentralized training based on PBFT using the following steps:
[0023] At the beginning of a consensus round, the leader broadcasts a Pre-prepare message containing a proposed block to other consensus nodes; during the consensus phase, the block header of the block proposed by the leader contains the state S of the transaction shard observed by the leader. * t , the block contains the leader based on S * t Placement result A * t During the reconstruction phase, the block header contains the basis for the agent's decision-making, namely the status of each transaction shard in the past few block production processes, and the block body contains the number of the transaction shard to which each consensus node has been reassigned. After receiving the Pre-prepare message, the consensus node broadcasts a Prepare message.
[0024] After the consensus node receives Prepare information for the same block from more than 2 / 3 of the consensus nodes, the consensus node will calculate the status S' of each transaction shard based on the local observation. t Calculate an action A' of your own agent t , and then set a predefined threshold φ in the agent shard a , through a predefined threshold φ a Avoid inconsistent agent actions due to inconsistent observation states; A' t and A * t The differences between them were evaluated by the Jaccard index;
[0025] Leaders and other nodes accumulate Commit messages for the new block. If more than two-thirds of the consensus nodes send Commit messages for the new block, all consensus nodes commit the block and update the local model, and the leader sends the transactions to the corresponding transaction shards.
[0026] A transaction processing system for integrating reinforcement learning and sharded blockchains, characterized by comprising agent shards and transaction shards; in the consensus stage, the consensus nodes in the agent shards send the corresponding transactions to the transaction shards by running a reinforcement learning algorithm; in the reconstruction stage, the agent shards complete the reconstruction of the transaction shards according to the situations in each transaction shard.
[0027] The beneficial effects of the present invention are as follows:
[0028] In the scenario of sharded blockchains, the present invention proposes a decentralized reinforcement learning deployment and training method, which can make the combination of sharded blockchains and reinforcement learning technologies more feasible.
[0029] The decentralized reinforcement learning training and deployment method designed by the present invention can effectively improve the performance and efficiency of sharded blockchains while ensuring the security and privacy protection of the system. This method can not only be applied to existing sharded blockchain systems, but also provide new ideas and methods for future sharded blockchain technologies. Brief Description of the Drawings
[0030] Figure 1 is a schematic diagram of the overall architecture of the system of the present invention.
[0031] Figure 2 is the block content of the agent shard of the present invention. Detailed Embodiment
[0032] The present invention will be further described in detail below through specific embodiments and the accompanying drawings.
[0033] 1. Basic Settings of the System
[0034] In the sharded blockchain, the present invention defines a time unit - an epoch. Each epoch contains two stages: a consensus stage and a reconstruction stage. The consensus stage is the stage in the sharded blockchain for confirming transactions and adding new transactions to the blockchain. In the consensus stage, the consensus nodes in each shard reach a consensus by executing a series of consensus algorithms to ensure the legality of the transactions and add them to the blockchain. The consensus nodes in each shard are randomly assigned to this shard.
[0035] The reconstruction phase is the phase to verify and receive newly joined nodes before the end of each epoch. In this phase, according to certain reconstruction rules, some nodes will randomly join another new shard. This helps maintain the diversity and security of the shard committee, ensuring that the shard chain system can continue to expand and prevent a certain shard from being controlled by malicious attackers. In the reconfiguration phase of each epoch, each shard chain application generates an unpredictable and bias-resistant random source using a verifiable random function, which is called epoch randomness. Epoch randomness is a mechanism used in shard blockchains to enhance security and privacy protection. Its main role is to introduce randomness at the beginning of each epoch, making it difficult for attackers to predict or control the members of a specific shard committee.
[0036] Blockchains are decentralized, and there may be malicious nodes launching attacks to disrupt the blockchain. Assume that the adversary cannot control more than f = 1 / 3 of the consensus nodes and cannot forge signatures. In practice, this can be achieved through various mechanisms to prevent Sybil attacks, such as those adopted by mature blockchains like Bitcoin and Ethereum, including proof-of-work and proof-of-stake. In the shard chain architecture proposed in the present invention, proof-of-work is used to prevent Sybil attacks. Proof-of-work requires nodes seeking to join the blockchain to solve a puzzle, and the last few bits of the solution string indicate which shard the node belongs to. All nodes in the shard chain architecture proposed in the present invention are connected by a partially synchronous peer-to-peer network. In a partially synchronous network, the message passing time is uncertain. In a partially synchronous network, it is defined that messages arrive within an unknown time upper bound Δ. In addition, this partially synchronous network may experience network partitions, but will return to normal after an unknown time.
[0037] 2. Shard Chain System Model
[0038] As Figure 1 shown, the shard chain architecture proposed in the present invention consists of two types of shards: agent shards and transaction shards. The number of agent shards is 1, and the number of transaction shards is k, where k is an integer greater than 1, for example, preferably greater than or equal to 8 or 16.
[0039] Transaction Shards: Transaction shards are responsible for verifying and processing transactions. Similar to traditional blockchains, transaction shards are responsible for processing transactions distributed to their own shards. The difference from traditional blockchains is that each transaction shard is only responsible for storing 1 / k of the users in the entire blockchain ledger. Users do not have a direct connection with transaction shards and thus do not perceive the existence of transaction shards.
[0040] Agent Sharding: Agent sharding is responsible for receiving users' transaction requests and storing these transactions in the local transaction pool. In the consensus phase, the consensus nodes in the agent sharding run a reinforcement learning algorithm to make the action of which transaction shard to send the transaction to (i.e., determine which transaction shard to send the transaction to), and send the corresponding transaction to the transaction shard. In the reconstruction phase, the agent sharding completes the relevant instructions for transaction shard reconstruction according to the situations in each transaction shard.
[0041] In summary, from the user's perspective, the user only needs to send the transaction to the agent sharding and then wait for the transaction to be processed and chained, without having to handle other messages. As a bridge connecting the transaction sharding and the client, the agent sharding plays an important role in the application of reinforcement learning in sharded blockchains. Running a sufficient number of consensus nodes in the agent sharding can ensure the stable operation of the agent sharding, and it can still operate normally even if a small number of nodes crash.
[0042] Both the agent sharding and the transaction sharding adopt the PBFT protocol based on Byzantine fault tolerance as the in-shard consensus protocol. In addition, a relay-based cross-shard transaction processing model is also adopted. In this model, the transaction is first processed on the source shard. After being executed on the source shard, the result is relayed to the target shard, and finally the cross-shard transaction is completed.
[0043] Figure 1 The working process of the in-shard consensus protocol shown above is described as follows:
[0044] 1) In each round of consensus, the leader of the agent sharding selects N transactions from its transaction pool and generates the distribution result of the transactions in the form of k transaction batches (i.e., the transactions sent to k transaction shards). After reaching a consensus on the distribution result of the transactions, the transaction batches are sent to the corresponding transaction shards.
[0045] 2) The transaction shard validates and executes the transaction, and only valid transactions will be included in the block. In addition, the new addresses corresponding to the new state and the number of cross-shard transactions are also recorded in the new block.
[0046] 3) Finally, by observing the new transaction shard blocks, the consensus nodes in the agent sharding record the address distribution, the number of cross-shard transactions, and the total number of transactions in each transaction shard for further training of the agent. In addition, the relevant information in the entire consensus phase is also recorded by the consensus nodes in the agent sharding for the agent decision-making in the reconstruction phase.
[0047] Specifically, during the reconstruction phase, nodes are reallocated to different shards according to the new sharding strategy or network requirements. This helps ensure that each shard has sufficient computing power, storage space, and security. Additionally, during the process of reallocating nodes, it may be necessary to migrate data from one shard to another. This includes transaction records, smart contract states, and other important information. Data migration needs to ensure data consistency and integrity and perform verification and repair if necessary. At this time, based on the results of previous multiple transaction allocations and observations of the load status in each shard's blockchain, the agent makes decisions on the selection of nodes that need to be reallocated and decides which shard these nodes will be assigned to.
[0048] Among them, an agent is a core concept in reinforcement learning. It is an entity with decision-making and action capabilities that learns how to achieve a certain goal through interaction with the environment. The agent performs actions in the environment and observes the feedback from the environment to learn how to optimize its behavior. Specifically, the agent makes decisions based on the observed environmental state and reward signal and executes corresponding actions. Through continuous attempts and learning, the agent gradually optimizes its decision-making strategy to maximize cumulative rewards or achieve the expected goal.
[0049] 3. Decentralized Training and Deployment Process Based on PBFT
[0050] To avoid the problem of centralization, the present invention proposes a decentralized agent deployment and training method. Each consensus node in the agent shard maintains a copy of the agent with the same initial parameters. Additionally, the hash value of the genesis block (the first block) in the agent shard is used as the random seed for all agents. The update of the agent and the state placement operation of all agents are determined through the consensus process. During the consensus phase, a node (the consensus node in the agent shard) is selected as the leader according to epoch randomness and proposes a new block.
[0051] The data structure of the block in the agent shard is as Figure 2 shown, including two parts: the block header and the block body. The block header includes metadata (the hash value of the previous block, the address of the leader node, etc.), votes from other consensus nodes, and the root and leaf nodes of the Merkle Patricia Tree (MPT). The leaf nodes of the MPT store the mapping from the address to the transaction shard ID. Using the MPT, the existence and location of an address can be efficiently queried with a complexity of O(logn), where n is the number of addresses.
[0052] Since the validity of all content in the blocks of the agent sharding can be verified through transactions, the basic security and effectiveness of the underlying consensus protocol are not complex. Therefore, the present invention proposes a training process based on PBFT. This process also reveals how the leader in the agent sharding proposes a block and helps all consensus nodes to update the agent consistently.
[0053] 1) Pre-prepare. At the beginning of a round of consensus, the leader broadcasts a Pre-prepare message containing the proposed block to other consensus nodes. For the consensus phase, as Figure 2 shown, the block header of the block proposed by the leader contains the state S of the transaction sharding observed by the leader * t , while the block body contains the placement result A based on S * t . S * t and A * t are components related to reinforcement learning, where S * t represents the state observed by the agent, and A * t represents the action taken by the agent for the current state. The reconstruction phase is similar. At this time, the content in the block header is still the basis for the agent to make decisions, that is, the states of each transaction sharding in the past several block production processes, and the content in the block body is related to sharding reconstruction, that is, the numbers of the transaction shards to which each consensus node is reappointed. After receiving the Pre-prepare message, the consensus node will broadcast a Prepare message, the content of which is the hash value of the block proposed by the master node received. In this way, it can be determined that the data received by each node is consistent. * t
[0054] 2) Prepare. After the consensus node receives the Prepare messages for the same block from more than 2 / 3 of the consensus nodes, the consensus node will calculate an action A' of its own agent according to the state S' of each transaction sharding observed locally t . Subsequently, a predefined threshold φ t is set in the agent sharding. Due to the characteristics of the distributed system, the states observed by different consensus nodes may be different. In this case, through the predefined threshold φ a , the inconsistency of the agent actions caused by inconsistent observed states can be avoided. Since the decision result of the agent is represented as k batches in the form of an array, A' a and A t * t The difference between them can be evaluated by the Jaccard index.
[0055]
[0056] Given two sets A and B, the Jaccard coefficient is defined as the ratio of the size of the intersection of A and B to the size of the union of A and B. If the difference is within φ a the consensus nodes will broadcast a commit message to confirm the validity of the proposed block.
[0057] 3) Commit. The leader and other nodes accumulate commit messages for the new block. If more than 2 / 3 of the consensus nodes send commit messages for the new block, all consensus nodes will commit the block and update the local model. In addition, the leader will send the transaction to the corresponding transaction shard along with a proof, indicating that the placement result is reached through a round of consensus.
[0058] During the consensus process of PBFT, once the nodes reach the prepared state, that is, they have verified the transactions and are ready to commit, they will enter the commit phase. In this phase, the nodes will broadcast commit messages to all other nodes. This commit message contains the necessary information related to the transaction. When a node receives commit messages from other nodes, it will verify the correctness of these messages. Only when certain conditions are met will the node accept these commit messages and record them in the local log. When a node receives valid commit messages from a sufficient number of other nodes (usually more than 2 / 3 of the total number), it can consider that this transaction has been confirmed by the entire system and can enter the committed state. This means that the transaction has been verified and committed by a sufficient number of honest nodes and can be safely executed and the result returned to the client.
[0059] The above sharded chain architecture and the PBFT-based agent training process combine sharded blockchain and reinforcement learning, enabling sharded blockchain to also complete the training of reinforcement learning in a decentralized manner without interfering with its original consensus process.
[0060] The key points of the present invention include:
[0061] 1. The present invention proposes a novel sharded chain architecture, including two types of shards, namely transaction shards and agent shards, for adapting the application of reinforcement learning in the sharded chain.
[0062] 2. The present invention proposes a reinforcement learning training process based on PBFT. By combining the process of reinforcement learning training and model update with the consensus process of the blockchain, each node in the blockchain can reach a consensus on the model update under the premise of decentralization.
[0063] 3. The present invention designs a block data structure suitable for deploying and training reinforcement learning in the blockchain ( Figure 2 ). By including the observation state of the agent and the actions taken by the agent in the block, each slave node receiving the block can independently verify the model update, thereby further ensuring the decentralization and security of the reinforcement learning deployment and training process.
[0064] 4. Compared with the existing combination of reinforcement learning and sharded blockchains, this solution starts from the system architecture and consensus protocol, and proposes a more practical and feasible system deployment method for integrating reinforcement learning and sharded blockchains.
[0065] Another embodiment of the present invention provides a transaction processing system for integrating reinforcement learning and sharded blockchains, which is characterized by including an agent shard and a transaction shard; in the consensus stage, the consensus nodes in the agent shard send the corresponding transactions to the transaction shard by running a reinforcement learning algorithm; in the reconstruction stage, the agent shard completes the reconstruction of the transaction shard according to the situations in each transaction shard.
[0066] For the specific processing processes of the agent shard and the transaction shard and the training process of the agent, refer to the description of the method of the present invention above.
[0067] Another embodiment of the present invention provides a computer device (such as a computer, a server, a smart phone, etc.), which includes a memory and a processor. The memory stores a computer program, and the computer program is configured to be executed by the processor. The computer program includes instructions for executing the steps in the method of the present invention.
[0068] Another embodiment of the present invention provides a computer-readable storage medium (such as ROM / RAM, a disk, an optical disc). The computer-readable storage medium stores a computer program, and when the computer program is executed by the computer, the various steps of the method of the present invention are implemented.
[0069] The specific embodiments of the present invention disclosed above are intended to help understand the content of the present invention and implement it accordingly. Those of ordinary skill in the art can understand that various substitutions, changes, and modifications are possible without departing from the spirit and scope of the present invention. The present invention should not be limited to the content disclosed in the embodiments of this specification, and the protection scope of the present invention is subject to the scope defined by the claims.
Claims
1. A transaction processing method for integrating reinforcement learning with a sharded blockchain, characterized in that, Including the following steps: Establish a shard chain, where the shard chain includes agent shards and transaction shards; In the consensus phase, the consensus nodes in the agent shards send the corresponding transactions to the transaction shards by running a reinforcement learning algorithm; In the reconstruction phase, the agent shards complete the reconstruction of the transaction shards according to the situations in each transaction shard.
2. The method according to claim 1, wherein The number of the agent shards is 1, and the number of the transaction shards is k, where k is an integer greater than 1; the transaction shards are responsible for processing the transactions distributed to their own shards, and each transaction shard is only responsible for storing 1 / k of the users in the entire blockchain ledger; the agent shards serve as a bridge connecting the transaction shards and the user side, and users do not have direct contact with the transaction shards.
3. The method according to claim 1, wherein Both the agent shards and the transaction shards adopt the PBFT protocol based on Byzantine fault tolerance as the in-shard consensus protocol, and adopt a cross-shard transaction processing model based on relay to implement cross-shard transactions.
4. The method according to claim 3, wherein The in-shard consensus protocol adopted by the agent shards and the transaction shards includes: In each round of consensus, the leader of the agent shard selects transactions from its transaction pool, generates the allocation result of the transactions in the form of k transaction batches, and after reaching a consensus on the allocation result of the transactions, the transaction batches will be sent to the corresponding transaction shards; The transaction shards verify and execute the transactions, incorporate the valid transactions into the block, and record the new addresses corresponding to the new states and the number of cross-shard transactions in the new block; The consensus nodes in the agent shards record the address distribution, the number of cross-shard transactions, and the total number of transactions in each transaction shard for training the agent, and the consensus nodes in the agent shards record the relevant information in the entire consensus phase for the agent decision-making in the reconstruction phase.
5. The method according to claim 1, wherein The agent shards adopt a decentralized agent deployment and training method. Each consensus node in the agent shards maintains a copy of the agent with the same initial parameters. The hash value of the genesis block in the agent shards is used as the random seed for all agents. The update of the agent and the state placement operation of all agents are determined through the consensus process. In the consensus phase, a consensus node in the agent shards is selected as the leader according to the epoch randomness and proposes a new block.
6. The method according to claim 5, characterized in that, The data structure of the block in the agent shards includes a block header and a block body; the block header includes metadata, votes from other consensus nodes, and the root node and leaf nodes of the Merkle Patricia tree. The leaf nodes of the Merkle Patricia tree store the mapping from the address to the transaction shard ID.
7. The method according to claim 6, wherein The agent shards adopt the following steps to achieve decentralized training based on PBFT: At the beginning of a consensus round, the leader broadcasts a Pre-prepare message containing a proposed block to other consensus nodes; during the consensus phase, the block header of the block proposed by the leader contains the state S of the transaction shard observed by the leader. * t , the block contains the leader based on S * t Placement result A * t During the reconstruction phase, the block header contains the basis for the agent’s decision-making, namely the status of each transaction shard in the past few block production processes, and the block body contains the number of the transaction shard to which each consensus node was reassigned; After receiving the Pre-prepare message, the consensus node broadcasts a Prepare message; After the consensus node receives the Prepare information for the same block from more than 2 / 3 of the consensus nodes, the consensus node calculates an action A' of its own agent according to the status S' of each transaction shard observed locally t and then sets a predefined threshold φ in the agent shard t . By means of the predefined threshold φ a , the inconsistency of the agent actions caused by inconsistent observed states is avoided; the difference between A' a and A t is evaluated by the Jaccard index * t The leader and other nodes accumulate Commit messages for the new block. If more than 2 / 3 of the consensus nodes send Commit messages for the new block, all consensus nodes submit the block and update the local model, and the leader sends the transactions to the corresponding transaction shards.
8. A transaction processing system for integrating reinforcement learning and sharded blockchain, characterized in that, It includes agent sharding and transaction sharding; in the consensus stage, the consensus nodes in the agent sharding send the corresponding transactions to the transaction sharding by running a reinforcement learning algorithm; in the reconstruction stage, the agent sharding completes the reconstruction of the transaction sharding according to the situations in each transaction sharding.
9. A computer device, characterized in that, It includes a memory and a processor, the memory stores a computer program, the computer program is configured to be executed by the processor, and the computer program includes instructions for executing the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and when the computer program is executed by a computer, the method according to any one of claims 1 to 7 is implemented.