An iot blockchain sharding method and device

By employing a dynamic sharding method driven by deep reinforcement learning, combined with a strategy committee and trustee scoring table, the resource allocation and cross-shard transactions of the IoT blockchain network are optimized, solving the scalability and security issues of traditional blockchains in the IoT and achieving efficient and secure network sharding.

CN119854361BActive Publication Date: 2025-11-25BEIJING UNIV OF POSTS & TELECOMM +2
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411454511.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-17
Publication Date
2025-11-25
Estimated Expiration
2044-10-17

AI Technical Summary

Technical Problem

Traditional blockchains suffer from insufficient scalability, severe storage and computing bottlenecks when dealing with large-scale IoT devices and data. Dynamic sharding technology cannot adapt to changes in network conditions, and cross-shard transaction processing is inefficient and difficult to guarantee security.

Method used

A dynamic sharding method based on deep reinforcement learning is adopted. Nodes are selected through a strategy committee, a dual deep reinforcement learning model is constructed, relay blocks are introduced to optimize cross-shard transactions, and a creditor scoring table is combined to prevent malicious nodes, thereby achieving efficient resource utilization and secure sharding.

Benefits of technology

It improves the transaction processing speed of IoT blockchain networks, reduces network latency, optimizes storage resource utilization, reduces energy consumption and computing costs, while ensuring data security and privacy protection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119854361B_ABST
    Figure CN119854361B_ABST
Patent Text Reader

Abstract

The application provides an Internet of Things blockchain sharding method and device, after nodes are selected into an Internet of Things blockchain network through a preset mechanism, a committee is screened based on the computing power and credit reputation of the nodes to supervise and execute transaction control and blockchain sharding. A deep reinforcement learning (DRL) actor network is used to select sharding actions based on the environment state, two state spaces and their corresponding action spaces are established for transaction and blockchain sharding respectively, a double deep reinforcement learning model is constructed, a first target network is introduced to stabilize the first target value predicted by a first evaluation network, a second target network is introduced to stabilize the second target value predicted by a second evaluation network, and transaction and blockchain sharding are simultaneously regulated to adapt to the dynamic changes of the network environment and realize efficient network sharding.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of blockchains, and particularly relates to an Internet of Things blockchain sharding method and device. BACKGROUND

[0002] With the rapid development of the Internet of Things, the number of devices worldwide is increasing, and it is estimated that by 2030, the number of Internet of Things devices connected to the Internet will reach hundreds of billions. These devices generate and exchange massive amounts of data, driving the development of applications such as smart homes, smart cities, autonomous vehicles, and industrial automation. However, the expansion of the Internet of Things also faces enormous challenges, particularly in terms of data security, privacy protection, storage capacity, bandwidth, and computing power. Traditional centralized architectures are increasingly unable to cope with the scale of devices and data volume, and are difficult to meet the needs of decentralization, security, and scalability.

[0003] In this context, blockchain technology, with its decentralized, secure, and tamper-proof characteristics, is considered a potential solution to the challenges of the Internet of Things. Through blockchain, Internet of Things devices can securely transmit and transact data in a decentralized network, avoiding dependence on centralized servers, thereby improving system reliability and attack resistance. However, blockchain itself also faces scalability issues, particularly when dealing with large-scale Internet of Things data, traditional blockchains face storage and computing bottlenecks.

[0004] To address these issues, innovative solutions are emerging in the blockchain field. Blockchain sharding technology is one of the key breakthroughs. Sharding technology divides the blockchain network into multiple independent sub-networks, called shards, each responsible for handling different transactions, thereby improving the transaction concurrency capability of the blockchain. Dynamic sharding further improves this technology, which can automatically adjust the number and size of shards based on network traffic, load, and other factors, ensuring optimal utilization of network resources.

[0005] Traditional sharding technology cannot adjust the expansion of the blockchain according to the network situation when facing large-scale networks, and has limitations in improving the number of transactions per second, resulting in decreased performance. Another difficulty in traditional sharding methods is the handling of cross-shard transactions, which can increase latency and complexity, affecting the efficiency of the blockchain system. The use of certain algorithms such as community division algorithms in sharding can reduce the occurrence of cross-shard transactions, reduce the interaction cost between shards, and improve the overall performance of the system. However, such division of shards has a high probability of causing the number of corrupt nodes in some community shards to exceed the normal range, thereby affecting the security of the entire blockchain. Therefore, optimizing cross-shard transactions while ensuring security is a major challenge.

[0006] Therefore, there is an urgent need for a dynamic sharding technology for Internet of Things blockchain to overcome the challenges brought by the explosive growth of the number of Internet of Things devices and the explosive growth of data exchange to the traditional blockchain architecture. SUMMARY

[0007] In view of this, the embodiments of the present application provide an Internet of Things blockchain sharding method and device to eliminate or improve one or more defects in the prior art, solve the problem that the existing Internet of Things blockchain sharding technology cannot adapt to network condition changes for sharding to achieve reasonable resource utilization, resulting in slow processing speed and high network delay.

[0008] One aspect of the present application provides an Internet of Things blockchain sharding method, which is executed by an agent of an Internet of Things blockchain network; when performing cross-shard transactions, the Internet of Things blockchain network introduces a relay block to relay the cross-shard transactions from the source shard to the target shard, and completes the subsequent operation of the transaction in the target shard; in each epoch, the method comprises the following steps:

[0009] A preset mechanism is used to select nodes to join the Internet of Things blockchain network, and based on the computing resources of each node and the score in the trust scorer table, a plurality of nodes are selected to form a policy committee;

[0010] The policy committee acts as an agent to build the number of transaction pool transactions, the number of completed transactions, and the total number of transaction transactions into a first state space, and build the epoch length, the block size, the number of shards, and the block interval into a first action space; the shard state of each node in the Internet of Things blockchain network is built into a second state space, and the shard adjustment of each node is built into a second action space;

[0011] The policy committee acts as an agent, and a first actor network based on deep reinforcement learning takes the parameters of the first state space as input and outputs the prediction of the first action in the first action space; a first evaluation network is used to predict a first target value of transaction throughput reward according to the first action, and the first target value introduces a first target network to predict the transaction throughput reward as a stable reference; the structure of the first target network is consistent with the first actor network and the first evaluation network;

[0012] The policy committee acts as an agent, and a second actor network based on deep reinforcement learning takes the parameters of the second state space as input and outputs the prediction of the second action in the second action space; a second evaluation network is used to predict a second target value of cross-shard transaction optimization reward according to the second action, and the second target value introduces a second target network to predict the cross-shard transaction optimization reward as a stable reference; the structure of the second target network is consistent with the second actor network and the second evaluation network;

[0013] The first actor network and the first evaluation network are updated according to the first target value, and the second actor network and the second evaluation network are updated according to the second target value; and the first target network and the second target network are soft-updated;

[0014] The policy committee performs the first action, and performs the second action when the number of cross-shard transactions in the Internet of Things blockchain network is higher than a set number of total transaction numbers.

[0015] In some embodiments, based on the computing resources of each node and the score in the trust scorer table, a plurality of nodes are selected to constitute a policy committee, including:

[0016] The historical behavior reputation is calculated as follows:

[0017]

[0018] wherein, represents the historical behavior reputation of node i, represents the number of transactions handled by node i in the time period t, represents the total number of transactions of all nodes in the blockchain, represents the online duration of node i, represents the current time of the system, represents the time when node i registers to the blockchain; , is a weight coefficient, and ;

[0019] The computing power reputation is calculated as follows:

[0020]

[0021] wherein, represents the computing power reputation of node i, represents the cpu performance value of the node, represents the network delay performance value of the node, represents the network bandwidth performance value of the node, represents the storage speed performance value of the node; total cpu performance value of all nodes, represents the total network delay performance value of all nodes, represents the total network bandwidth performance value of all nodes, represents the total storage speed performance value of all nodes. , , , is a weight coefficient, and .

[0022] The formula for calculating the comprehensive reputation of a node is:

[0023]

[0024] wherein, the comprehensive reputation of node i, and is a weight coefficient, and

[0025] In some embodiments, the preset mechanism includes: a proof-of-work mechanism, a proof-of-stake mechanism, or a reputation mechanism; the first actor network, the first evaluation network, the second actor network, and the second evaluation network all adopt a deep neural network structure.

[0026] In some embodiments, the formula for calculating the transaction throughput reward is:

[0027]

[0028] wherein, represents the transaction throughput reward, represents the number of shards, represents the epoch length, represents the block interval, represents the block size, represents the average number of redundant transactions.

[0029] In some embodiments, the formula for calculating the cross-shard transaction optimization reward is:

[0030] ;

[0031] wherein, represents the cross-shard transaction optimization reward, represents the number of intra-shard transactions, represents the number of cross-shard transactions, A represents the total number of transactions, and U represents the load imbalance degree; 、 and are parameters.

[0032] In some embodiments, the first target value introduces the prediction of the first target network on the transaction throughput reward as a stable reference, and the formula is:

[0033] ;

[0034] wherein, represents the first target value, represents the current value of the transaction throughput reward, is a first discount factor, represents a predicted value of the first target network on the transaction throughput reward; represents a state of the i+1 epoch, represents a parameter of a first target actor network in the first target network, represents a parameter of a first target critic network in the first target network;

[0035] The second target value introduces a predicted value of the second target network on the cross-shard transaction optimization reward as a stable reference, and the calculation formula is:

[0036] ;

[0037] wherein, represents the second target value, represents a current value of the cross-shard transaction optimization reward, is a second discount factor, represents a predicted value of the second target network on the cross-shard transaction optimization reward; represents a state of the i+1 epoch, represents a parameter of a second target actor network in the second target network, represents a parameter of a second target critic network in the second target network.

[0038] In some embodiments, the parameter update of the first actor network and the first critic network according to the first target value comprises:

[0039] updating the parameter of the first critic network based on the first target value, and the expression is:

[0040]

[0041] wherein, represents a weight parameter of the first critic network, represents the first target value, represents a learning rate, represents a weight gradient of the first critic network, represents the first target value estimated by the first critic network at present; represents a state of a first state space at the i epoch, represents an action of a first action space at the i epoch;

[0042] updating the parameter of the first actor network, and the expression is:

[0043] ;

[0044] wherein, denotes a weight parameter of the first actor network, denotes a learning rate, denotes a weight gradient of the first actor network, denotes the first target value currently estimated by the first critic network, denotes a policy function of the first actor network;

[0045] performing parameter update on the second actor network and the second critic network according to the second target value, comprising:

[0046] updating parameters of the second critic network based on the second target value, expressed as:

[0047]

[0048] wherein, denotes a weight parameter of the second critic network, denotes the second target value, denotes a learning rate, denotes a weight gradient of the second critic network, denotes the second target value currently estimated by the second critic network; denotes a state of the second state space at the i-th epoch, denotes an action of the second action space at the i-th epoch;

[0049] updating parameters of the second actor network, expressed as:

[0050] ;

[0051] wherein, denotes a weight parameter of the second actor network, denotes a learning rate, denotes a weight gradient of the second actor network, denotes the second target value currently estimated by the second critic network, denotes a policy function of the second actor network.

[0052] In some embodiments, performing soft update on the first target network and the second target network, comprising:

[0053] the weight update expression of the first target actor network is:

[0054] ;

[0055] wherein, a weight parameter of the first target actor network, a soft update coefficient, a weight parameter of the first actor network currently,

[0056] a weight update expression of the first target evaluation network is:

[0057] ;

[0058] wherein, a weight parameter of the first target evaluation network, a soft update coefficient, a weight parameter of the first evaluation network currently,

[0059] a weight update expression of the second target actor network is:

[0060] ;

[0061] wherein, a weight parameter of the second target actor network, a soft update coefficient, a weight parameter of the second actor network currently,

[0062] a weight update expression of the second target evaluation network is:

[0063] ;

[0064] wherein, a weight parameter of the second target evaluation network, a soft update coefficient, a weight parameter of the second evaluation network currently.

[0065] In another aspect, the present application also provides a computer readable storage medium having stored thereon a computer program or instructions, which, when executed by a processor, implement the steps of the above method.

[0066] In another aspect, the present application also provides a computer program product comprising a computer program or instructions, which, when executed by a processor, implement the steps of the above method.

[0067] The present application has at least the following advantages:

[0068] The Internet of Things blockchain sharding method and device, after the nodes are selected into the Internet of Things blockchain network through the preset mechanism, a node-based computing capability and credit reputation screening policy committee is used to supervise and execute transaction control and blockchain sharding.

[0069] Further, when screening the policy committee, a credit grantor score table is introduced to maintain a global account score, preventing the influence of malicious nodes.

[0070] The additional advantages, objects, and features of the application will be better understood from the following description accompanied by the appended drawings.

[0071] Those skilled in the art will understand that the objects and advantages of the present application are not limited to the above specifically described and that the above and other objects and advantages of the present application can be more clearly understood from the following detailed description when read in conjunction with the following drawings. BRIEF DESCRIPTION OF DRAWINGS

[0072] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments of the application and serve to explain the principles of the application.

[0073] Figure 1 The flowchart of the Internet of Things blockchain sharding method according to an embodiment of the present application.

[0074] Figure 2 The deep reinforcement neural network structure diagram in the Internet of Things blockchain network sharding scheme according to an embodiment of the present application. DETAILED DESCRIPTION

[0075] To make the objects, technical solutions and advantages of the present application clearer, the present application is further described in detail below in conjunction with the embodiments and drawings. Herein, the illustrative embodiments of the present application and their descriptions are used to explain the present application, but are not intended to limit the present application.

[0076] It is also noted herein that, while the above describes example embodiments, no limitations are intended to the scope of the inventive core as defined by the appended claims depending on the basic inventive concept, wherein the illustrated embodiments do not limit the scope of the claims accordingly.

[0077] It should be emphasized that the term "comprises / comprising" when used in this text is taken to mean the presence of stated features, elements, steps or components, but does not preclude the presence or addition of one or more other features, elements, steps or components.

[0078] It is also noted herein that, if not otherwise specified, the term "connected" in this text can not only mean direct connection, but also indirect connection in the presence of intermediates.

[0079] When traditional blockchain applications are applied in the Internet of Things, they face scalability problems, especially when dealing with large-scale Internet of Things data, traditional blockchains face storage and computing bottlenecks, and dynamic sharding is a method to improve the performance of blockchains. The core purpose of the present application is to overcome the challenges brought by the explosive growth of Internet of Things devices and the explosive growth of data exchange to the traditional blockchain architecture. Through innovative dynamic sharding technology based on reputation and deep reinforcement learning (DRL) driven, the efficient expansion of the blockchain network is realized. This technology focuses on improving transaction processing speed, reducing network latency, optimizing storage solutions, and ensuring data security and privacy protection, and realizing the rational allocation and utilization of resources. In addition, the present application also aims to dynamically adjust the number and size of shards to adapt to changing network conditions, and integrate lightweight consensus mechanisms to reduce network energy consumption and computing costs, thereby providing a flexible and efficient Internet of Things blockchain network sharding solution for Internet of Things devices.

[0080] Specifically, one aspect of the present application provides an Internet of Things blockchain sharding method, which is executed by an agent of an Internet of Things blockchain network; when executing cross-shard transactions, the Internet of Things blockchain network introduces a relay block to relay the cross-shard transactions from the source shard to the target shard, and completes the subsequent operation of the transaction in the target shard; in each epoch, the method comprises the following steps S101-S104:

[0081] Step S101: adopt a preset mechanism to select nodes to join the Internet of Things blockchain network, and select a plurality of nodes to form a policy committee based on the computing resources of each node and the score in the trust scorer table.

[0082] Step S102: The policy committee, as an agent, constructs the number of transaction pool transactions, the number of completed transactions, and the total number of transaction transactions as a first state space, and constructs the epoch length, the block size, the number of shards, and the block interval as a first action space; constructs the shard state of each node in the Internet of Things blockchain network as a second state space, and constructs the shard adjustment of each node as a second action space.

[0083] Step S103: The policy committee, as an agent, based on the first actor network of deep reinforcement learning, takes the parameters of the first state space as input and outputs the prediction of the first action in the first action space; the first evaluation network predicts the first target value of the transaction throughput reward according to the first action, and the first target network introduces the prediction of the transaction throughput reward as a stable reference; the structure of the first target network is consistent with the first actor network and the first evaluation network.

[0084] Step S104: The policy committee, as an agent, based on the second actor network of deep reinforcement learning, takes the parameters of the second state space as input and outputs the prediction of the second action in the second action space; the second evaluation network predicts the second target value of the cross-shard transaction optimization reward according to the second action, and the second target network introduces the prediction of the cross-shard transaction optimization reward as a stable reference; the structure of the second target network is consistent with the second actor network and the second evaluation network.

[0085] Step S105: Update the parameters of the first actor network and the first evaluation network according to the first target value, update the parameters of the second actor network and the second evaluation network according to the second target value, and perform soft update on the first target network and the second target network.

[0086] Step S106: The policy committee executes the first action, and when the number of cross-shard transactions in the Internet of Things blockchain network is higher than the set number of the total number of transaction transactions, executes the second action.

[0087] In step S101, the nodes constituting the Internet of Things blockchain are first selected, and the preset mechanism adopted includes: Proof of Work (POW), Proof of Stake (Pos), or reputation mechanism.

[0088] In the PoW mechanism, nodes need to solve a complex mathematical problem, called the PoW problem, to prove the right to add a new block to the blockchain. The PoW problem is usually a computationally complex but relatively simple verification problem. Nodes need to find a solution that meets the conditions through a large number of computational attempts, while verifying whether the solution is correct is relatively easy.

[0089] In proof-of-stake, the selection of nodes is based on the amount of cryptocurrency they hold in the network (i.e.'stake') rather than by solving complex mathematical problems (as in PoW). Nodes holding more tokens have a higher probability of being selected as validators and are responsible for creating new blocks.

[0090] The reputation mechanism is based on the historical behavior and contribution of nodes, and the system selects the validation nodes of the epoch through a reputation or credit score mechanism. The higher the reputation of a node, the greater the chance of being selected to participate in validation.

[0091] After the construction of the Internet of Things blockchain network is completed, the application also introduces a lender role in the blockchain for maintaining the global trust score of each node based on historical data. Specifically, each node will save user accounts, which is equivalent to the node corresponding to a part of the user, and then the user account has a lender score, and the aggregation of the user's lender score obtains the lender score of the node. After the node has the lender score, the global information sharing synchronization is the global trust degree, which can be used as the historical behavior credit.

[0092] At the same time, combined with the computing power resources of the nodes, the reputation of each node is comprehensively evaluated, and based on this, members of the strategy committee are selected.

[0093] In some embodiments, based on the computing resources of each node and the score in the lender score table, a plurality of nodes are selected to constitute a strategy committee, including steps S1011-S1013:

[0094] Step S1011: Calculate the historical behavior credit, the calculation formula is:

[0095]

[0096] Among them, represents the historical behavior credit of node i, represents the number of transactions handled by node i in the t time period, represents the total number of transactions of all nodes in the blockchain, represents the online duration of node i, represents the current time of the system, represents the time when node i registers to the blockchain; 、 is a weight coefficient, and .

[0097] Step S1012: Calculate the computing power credit, the calculation formula is:

[0098]

[0099] Among them, a computing power reputation of a node i, a CPU performance value of a node, a network delay performance value of a node, a network bandwidth performance value of a node, a storage speed performance value of a node; a total CPU performance value of all nodes, a total network delay performance value of all nodes, a total network bandwidth performance value of all nodes, a total storage speed performance value of all nodes. , , , is a weight coefficient, and .

[0100] Step S1013: the formula for calculating the comprehensive reputation of a node is:

[0101]

[0102] wherein, a comprehensive reputation of a node i, and is a weight coefficient, and

[0103] In the implementation process, nodes with a comprehensive reputation higher than a set value can be included in the strategy committee, or a set proportion of nodes with a high comprehensive reputation can be included in the strategy committee.

[0104] In steps S102-S105, a double deep reinforcement learning (DRL) model is introduced. On the one hand, the first DRL constructs the number of transaction pool transactions, the number of completed transactions, and the total number of transaction transactions as a first state space, and constructs the epoch length, block size, number of shards, and block interval as a first action space, and adopts a transaction throughput reward. The first DRL includes a first actor network and a first evaluation network. The first actor network performs prediction on actions according to the state, and the first evaluation network is used to predict the transaction throughput reward to obtain a first target value. The first target value introduces a first target network to stabilize the prediction of the transaction throughput reward. The structure of the first target network is consistent with that of the first actor network and the first evaluation network, and adopts a soft update form, and its prediction changes less with the environment.

[0105] In another aspect, a second DRL constructs the shard state of each node in the Internet of Things blockchain network as a second state space, and constructs the shard adjustment of each node as a second action space, and adopts a cross-shard transaction optimization reward. The second DRL includes a second actor network and a second critic network. The second actor network performs a prediction of an action according to a state. The second critic network is used to predict that the cross-shard transaction optimization reward reaches a second target value. The second target value introduces a second target network to stabilize the prediction of the cross-shard transaction optimization reward. The structure of the second target network is consistent with the second actor network and the second critic network, and the second target network adopts a soft update form, and the prediction of the second target network changes less with the environment.

[0106] In some embodiments, the first actor network, the first critic network, the second actor network, and the second critic network all adopt a deep neural network structure. The structure of the first target network is consistent with the first actor network and the first critic network. The structure of the second target network is consistent with the second actor network and the second critic network.

[0107] In some embodiments, the calculation formula of the transaction throughput reward is:

[0108]

[0109] In some embodiments, the calculation formula of the transaction throughput reward is: represents the transaction throughput reward, represents the number of shards, represents the epoch length, represents the block interval, represents the average number of redundant transactions.

[0110] In some embodiments, the calculation formula of the cross-shard transaction optimization reward is:

[0111] ;

[0112] In some embodiments, the calculation formula of the cross-shard transaction optimization reward is: represents the cross-shard transaction optimization reward, represents the number of intra-shard transactions, represents the number of cross-shard transactions, A represents the total number of transactions, and U represents the load imbalance degree. and are parameters.

[0113] In some embodiments, the first target value introduces a first target network to predict the transaction throughput reward as a stable reference, and the calculation formula is:

[0114] ;

[0115] In some embodiments, the first target value introduces a first target network to predict the transaction throughput reward as a stable reference, and the calculation formula is: ​​denotes the first target value, denotes a current value of the transaction throughput reward, is a first discount factor, denotes a predicted value of the transaction throughput reward by the first target network; denotes a state of the i+1 epoch, denotes a parameter of a first target actor network in the first target network, denotes a parameter of a first target critic network in the first target network;

[0116] The second target value introduces a prediction of the cross-shard transaction optimization reward by the second target network as a stable reference, and the calculation formula is:

[0117] ;

[0118] wherein, denotes the second target value, denotes a current value of the cross-shard transaction optimization reward, is a second discount factor, denotes a predicted value of the cross-shard transaction optimization reward by the second target network; denotes a state of the i+1 epoch, denotes a parameter of a second target actor network in the second target network, denotes a parameter of a second target critic network in the second target network.

[0119] In some embodiments, the parameter updating of the first actor network and the first critic network according to the first target value comprises steps S201-S202:

[0120] Step S201: updating the parameter of the first critic network based on the first target value, and the expression is:

[0121]

[0122] wherein, denotes a weight parameter of the first critic network, denotes the first target value, denotes a learning rate, denotes a weight gradient of the first critic network, denotes a first target value estimated by the current first critic network; denotes a state of the first state space at the i epoch, denotes an action of the first action space at the i epoch.

[0123] Step S202: updating the parameter of the first actor network, and the expression is:

[0124] ;

[0125] wherein, denotes the weight parameter of the first actor network, denotes the learning rate, denotes the weight gradient of the first actor network, denotes the first target value currently estimated by the first critic network, denotes the policy function of the first actor network.

[0126] According to the second target value, the parameters of the second actor network and the second critic network are updated, including steps S301-S302:

[0127] Step S301: The parameters of the second critic network are updated based on the second target value, and the expression is:

[0128]

[0129] wherein, denotes the weight parameter of the second critic network, denotes the second target value, denotes the learning rate, denotes the weight gradient of the second critic network, denotes the second target value currently estimated by the second critic network. denotes the state of the second state space at the i-th epoch, denotes the action of the second action space at the i-th epoch.

[0130] Step S302: The parameters of the second actor network are updated, and the expression is:

[0131] ;

[0132] wherein, denotes the weight parameter of the second actor network, denotes the learning rate, denotes the weight gradient of the second actor network, denotes the second target value currently estimated by the second critic network, denotes the policy function of the second actor network.

[0133] In some embodiments, the first target network and the second target network are soft updated, including steps S401-S404:

[0134] Step S401: The weight update expression of the first target actor network is:

[0135] ;

[0136] wherein, denotes the weight parameter of the first target actor network, denotes the soft update coefficient, denotes the current weight parameter of the first actor network.

[0137] Step S402: the weight update expression of the first target evaluation network is:

[0138] ;

[0139] wherein, denotes the weight parameter of the first target evaluation network, denotes the soft update coefficient, denotes the current weight parameter of the first evaluation network.

[0140] Step S403: the weight update expression of the second target actor network is:

[0141] ;

[0142] wherein, denotes the weight parameter of the second target actor network, denotes the soft update coefficient, denotes the current weight parameter of the second actor network.

[0143] Step S404: the weight update expression of the second target evaluation network is:

[0144] ;

[0145] wherein, denotes the weight parameter of the second target evaluation network, denotes the soft update coefficient, denotes the current weight parameter of the second evaluation network.

[0146] In step S106, at the end of each epoch, the policy committee initiates re-sharding when the number of cross-shard transactions is higher than a set proportion, for example, when the proportion of cross-shard transactions in transaction transactions is 50%, re-sharding is performed once, that is, the second action is performed. The tolerance of cross-shard in different environments is different, and this proportion can be set according to needs. Re-sharding is equivalent to restarting the entire model once.

[0147] On the other hand, the application also provides a computer readable storage medium, which stores a computer program or instructions, and the computer program or instructions are executed by a processor to realize the steps of the above method.

[0148] In another aspect, the present application also provides a computer program product comprising computer programs or instructions which, when executed by a processor, implement the steps of the above method.

[0149] The present application will be described below in conjunction with a specific embodiment:

[0150] The present embodiment provides an Internet of Things blockchain network sharding scheme, and the system model is as follows:

[0151] The present embodiment proposes the architecture of the Credit and Dual DRL Verification (CDDV) model, as shown in Figure 2 At the beginning of each epoch, new nodes that complete the pow problem within a specified time can join the blockchain network. In each shard, any node can propose a node as a leader. A strategy committee (SC) is introduced to reset the shards according to the epoch strategy, which is composed of members elected by network users democratically. Inspired by Elastico, the SC ensures decentralized and reliable supervision. This design prevents single point of failure and minimizes the risk from centralized agents. As a decentralized and trusted coordinator, the SC supervises the user list of the entire network and allocates shards to nodes.

[0152] The most important work for SC is to determine and give the new sharding strategy when starting a new epoch at the end of each epoch, and to redistribute the nodes among shards according to the changes in the environment. This adaptability ensures that the system remains scalable and secure, and adapts to higher transaction volumes without compromising security. A DRL model is used to assist the blockchain reconfiguration process. The agent is composed of a committee elected by all peers in the blockchain network. The agent obtains the state from the node allocation, obtains the reward by virtual resharding, decides the optimal allocation strategy and executes the new node allocation action. The permissioned blockchain system restricts network access of specific nodes. In order to make the training process be considered as credible, it is important that the agent is reliable and has transparent records when training the DRL model. In addition, it is assumed that each shard involves at least a certain number of members, and it is claimed that the number of dishonest nodes in each shard does not exceed a certain number of the total number of nodes based on the Byzantine Fault Tolerance (BFT) principle. In the case of node failure or attack by dishonest nodes, the system continues to run with the remaining nodes. Each node maintains a lender ledger of the node, where the lender score is assigned to other network nodes. It aggregates the lender scores of individual users in the node to derive a global trust metric for each node, thereby protecting trust and node distribution. Specifically, each node will save the user's account, which is equivalent to a part of the user corresponding to the node, and then the user account will have a lender score. The aggregation of the lender scores of users is the lender score of the node, and after the node has the lender score, the global information sharing synchronization is the global trust degree.

[0153] In order to be able to find nodes with good reputation and rich computing resources and form a shard strategy committee SC (Strategy Committee), a temporary directory committee is elected in the same way as elastico when the blockchain is started for the first time. The directory committee calculates and stores the reputation of each node, and randomly selects 10 nodes from the high-reputation candidate node set as the shard strategy committee SC. After the shard strategy committee SC is selected, the SC will gradually replace the functions of the temporary directory committee, and the SC will synchronize and store the reputation of each node and handle the update of the SC. By comprehensively evaluating the reputation of the nodes, the invention designs a node reputation mechanism composed of historical behavior reputation and computing power reputation to accurately reflect the performance and historical behavior of the nodes. The historical behavior reputation represents the credibility of the historical behavior of the node in the recent t time period, and is defined as:

[0154]

[0155] wherein, represents the historical behavior reputation of node i, represents the number of transactions handled by node i in the t time period, represents the total number of transactions in the blockchain, represents the online duration of node i, represents the current time of the system, represents the time when node i registers to the blockchain; is a weight coefficient, and .

[0156] The computing power reputation reflects the computing power performance of the node hardware, and is defined as

[0157]

[0158] wherein, represents the computing power reputation of node i, represents the cpu performance value of the node, represents the network delay performance value of the node, represents the network bandwidth performance value of the node, represents the storage speed performance value of the node; total cpu performance value of all nodes, represents the total network delay performance value of all nodes, represents the total network bandwidth performance value of all nodes, represents the total storage speed performance value of all nodes. is a weight coefficient, and .

[0159] The comprehensive reputation of the node is defined as

[0160]

[0161] wherein, comprehensive reputation of node i, and is a weight coefficient, and

[0162] In the deep reinforcement learning model DRL, the SC (Strategy Committee) is an agent composed of a number of nodes. These nodes are selected from each shard according to computing resources and lender records, execute the PBFT protocol to reach consensus, and represent the collective action of the blockchain. The agent executes a learning process and decides shard allocation according to real-time network conditions.

[0163] ​​​​The environment is seen as a black box that executes the agent's actions and obtains the state. The embodiment takes the shard reconstruction process in the blockchain sharding network as the environment. Consider that the agent obtains a state s according to the node allocation of the current epoch t , s represents the distribution of all nodes in the current episode on different shards, and the matrix A accurately identifies the specific nodes that exist in a specific shard.

[0164] ;

[0165] ;

[0166] wherein, is the number of pool transactions, is the number of completed transactions, and n is the total number of transaction transactions. is the i-th shard to which the x-th node belongs, D is the number of nodes, and N is the number of shards.

[0167] The construction of the action space includes the epoch length T, the block size B, the number of shards K, the block interval and the node shard matrix , which plays an important role in solving the problem of dynamic sharding system.

[0168] When considering that the arrival of nodes obeys a certain distribution, the epoch length will determine the number of nodes in the system in the next epoch. In addition, the block interval and the block size can change the state of the transaction pool by affecting the rate of processing transactions. Therefore, it is necessary to adjust them to adapt to the dynamic environment. Considering the computation in the system, the action space a1 and a2 at time t are defined as:

[0169] ;

[0170]

[0171] is the i-th shard to which the x-th node is allocated, D is the number of nodes, and N is the number of shards.

[0172] The reward is set, and the optimization goal of the embodiment is to obtain a shard strategy that balances performance and security at each period. Since scalability can be easily quantified by transaction throughput, a transaction throughput reward is used as the reward function, and a cross-shard transaction optimization reward is introduced.

[0173] The calculation formula of the transaction throughput reward is:

[0174]

[0175] wherein, represents the transaction throughput reward, represents the number of shards, represents the epoch length, represents the block interval, represents the block size, represents the average number of redundant transactions.

[0176] The calculation formula of the cross-shard transaction optimization reward is:

[0177] ;

[0178] wherein, represents the cross-shard transaction optimization reward, represents the number of intra-shard transactions, represents the number of cross-shard transactions, A represents the total number of transactions, and U represents the load imbalance degree; 、 and are parameters.

[0179] The blockchain dynamic sharding framework based on the CDDV model, the input includes: the state of the blockchain network , the lender account record (the global trust degree of the node, the scoring situation), the experience replay buffer , the replay buffer size , the initial weight of the actor network , the initial weight of the evaluator network , the learning rate and the discount factor . The output includes: the sharding reconfiguration strategy, the number of shards , the epoch length , the block interval , the block size and the node allocation strategy .

[0180] Specifically, a double DRL is constructed, DRL 1 corresponds to the state space , the action space , and the transaction throughput reward is adopted; DRL 2 corresponds to the state space , the action space , and the cross-shard transaction optimization reward is adopted.

[0181] The specific operation steps are as follows:

[0182] S1. Initialize the network and data structure: initialize the actor network and the evaluator network .

[0183] S2. Initialize the target network and .

[0184] S3. Initialize experience replay buffer R and global lender scorebook.

[0185] S4. Initialize control variables for sharding reconfiguration strategy .

[0186] S5. Node selection and allocation at the beginning of each epoch: adopt PoW mechanism for node selection. The selected new nodes can join the blockchain network and be evaluated for their contribution value and trust level based on the lender score table. The strategy committee (SC) elected by democracy supervises the allocation of nodes.

[0187] S6. Use actor network According to the current state of the blockchain Select a sharding reconfiguration action , including epoch length T, block size B, number of shards K and block interval .

[0188] S7. Execute action and feedback reward: execute the sharding reconfiguration action, adjust node allocation and shard size, and receive rewards based on the execution results .

[0189] S8. Store state transition information: store state transitions into the experience replay buffer R, and randomly draw a small batch N of state transitions from the buffer for model training.

[0190] S9. Calculate target value using evaluation network: , update the weights of the evaluation network : , update the weights of the actor network to maximize the value of the evaluation network: .

[0191] S10. Soft update of target network: update the weights of the target network and : , .

[0192] S11. Sharding reconfiguration and cross-shard transaction optimization: at the end of each Epoch, the strategy committee (SC) decides whether to re-shard according to the results of the DRL model. If there are many cross-shard transactions, trigger re-sharding and node reconfiguration, and use the lender score mechanism to identify malicious or low-reputation nodes.

[0193] The execution flow of the DRL-based dynamic sharding framework in the algorithm. First, the network state, lender ledger, and sharding reconfiguration strategy are initialized. In each epoch, the algorithm selects a new node through the PoW mechanism, and the allocation of the node is supervised by the policy committee. The agent maintains the evaluation of the sharding strategy and the value function , where and are the parameters of the participant network and the critic network. The agent uses the target network modified for actor-critic (initialized as a copy of the actor-critic network) to calculate the target value of the network update. To ensure that the training samples are independent and identically distributed, a replay buffer of limited size is used for sampling. At each time step t, the agent selects and executes a sharding action A according to the current blockchain state s, and then applies noise N for exploration. After that, the blockchain environment will give the agent a reward measured by system security and throughput, and enter the next state . The agent will convert and store it in R, and then batch a constant number of previous conversions from the replay buffer R to calculate the loss function and the policy gradient for updating the critic network parameters and the actor network parameters . Finally, the algorithm uses soft target updates to slowly change the target network: and .

[0194] Further, when a cross-shard transaction needs to be performed, the system introduces a relay block to relay the cross-shard transaction from the source shard to the target shard, and the target shard completes the subsequent operation of the transaction. In the traditional two-phase commit process, the transaction is divided into multiple sub-transactions, and after the last sub-transaction is completed, the confirmation and submission process of the transaction is performed. Since the source shard needs to wait for the confirmation and submission message of the target shard after starting the transaction, the confirmation process prolongs the cross-shard transaction time of the two-phase commit, which makes the cross-shard transaction time much longer than the intra-shard transaction time.

[0195] After introducing the final atomicity of cross-shard, the confirmation time of cross-shard transaction can be reduced, thereby improving the efficiency of cross-shard transaction. In this embodiment, the cross-shard transaction is divided into two operations, withdrawal operation and deposit operation. The withdrawal operation is completed in the source shard, and a relay block is generated. After the relay block is generated, it is forwarded to the target shard, and the target shard verifies the relay block to complete the deposit operation. After the relay block is forwarded to the target shard, it will ensure its final execution of the deposit operation, i.e., final atomicity. The payment transaction involving withdrawal and deposit operations should be atomic, and the final atomicity ensures the correctness of the sharding system.

[0196] The embodiment introduces a lender role in the blockchain, the lender role maintains a global account score and a withdrawal record ledger, evaluates each account, when the account meets the requirements, the account performs an early withdrawal operation when the account performs a cross-shard transaction, the withdrawal amount is temporarily deducted from the lender and recorded, and does not need to wait for the cross-shard shard confirmation, when the lender receives the transaction confirmation submission of the target shard, the lender will deduct the account early withdrawal operation amount, the account can quickly change the account balance when performing a cross-shard transaction, and does not need to wait for transaction confirmation, which greatly reduces the confirmation time of the cross-shard transaction.

[0197] The lender will record the information of each account performing a cross-shard transaction, and update the score of the account after the account initiates early withdrawal. The score includes the user's initiative and safety, the higher the score of the account, the higher the amount of early withdrawal initiated, and the account with a substandard score will not enjoy the early withdrawal service. After introducing the lender role, the account performing a large number of cross-shard transactions will not reduce the efficiency due to the long cross-shard transaction time, and the malicious account will be quickly identified due to the low score.

[0198] The simulation experiment results show that the CCDV blockchain sharding model proposed in the application has obvious improvement in the number of transactions (TPS) of the blockchain compared with the random sharding, community-based sharding and reputation-based sharding model, and reduces the cross-shard transaction while ensuring a low proportion of corrupt nodes in the shard. The dynamic sharding based on reputation and double DRL verification proposed in the application can adjust the suitable blockchain sharding strategy according to the change of the blockchain environment and improve the cross-shard efficiency through the improved relay transaction processing, so that the dynamic sharding model proposed in the application can ensure the safety of the blockchain system while reducing the number of cross-shard transactions, thereby realizing a high number of transactions per second.

[0199] The embodiment realizes adaptive reconfiguration of blockchain sharding by introducing a double deep reinforcement learning model, dynamically adjusts the sharding strategy according to network load and transaction demand. The agent node monitors and learns the current state of the system in real time through the DRL algorithm, including the load of the node, the number of transactions, the resource usage of the shard, etc. The DRL model continuously adjusts the node allocation and shard structure in the blockchain through virtual resharding and strategy evaluation to adapt to the dynamic changes of the network environment. The lender role is introduced to maintain the global account score and record the ledger. In cross-shard transactions, the lender allows high credit score nodes to complete cross-shard transactions quickly without waiting for the target shard to confirm, thereby significantly reducing the transaction confirmation time and improving system efficiency. The simulation experiment results show that the scheme proposed by the application can realize higher system transaction per second while ensuring higher security.

[0200] Corresponding to the above method, the application also provides a device or system, which comprises a computer device including a processor and a memory, the memory storing computer instructions, and the processor is configured to execute the computer instructions stored in the memory, and the device or system implements the steps of the method as described above when the computer instructions are executed by the processor.

[0201] The embodiment of the application also provides a computer readable storage medium having a computer program stored thereon, which is executed by a processor to implement the steps of the aforementioned edge computing server deployment method. The computer readable storage medium can be a tangible storage medium, such as a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a floppy disk, a hard disk, a removable storage disk, a CD-ROM, or any other form of storage medium known in the technical field.

[0202] In summary, the Internet of Things blockchain sharding method and device of the application, after selecting nodes into the Internet of Things blockchain network through a preset mechanism, filters a policy committee based on the computing power and credit reputation of the nodes to supervise and execute transaction control and blockchain sharding. The deep reinforcement learning DRL actor network selects shard actions based on the environment state, establishes two state spaces and their corresponding action spaces for transaction transactions and blockchain sharding respectively, constructs a double deep reinforcement learning model, introduces a first target network to stabilize the first target value predicted by the first evaluation network, introduces a second target network to stabilize the second target value predicted by the second estimation network, and simultaneously controls transaction transactions and blockchain sharding to adapt to the dynamic changes of the network environment and realize efficient network sharding.

[0203] Further, in the screening policy committee, a credit score table is introduced to maintain a global account score to prevent the impact of malicious nodes.

[0204] Those of ordinary skill in the art will appreciate that the various illustrative components, systems and methods described in connection with the embodiments disclosed herein can be implemented as hardware, software, or hardware and software in combination. The various illustrative components, systems and methods can be implemented in hardware, software or a combination of both, depending on the particular application and design constraints imposed on the overall system. Skilled persons can use various methods to implement the described functions for each particular application, but such implementation should not be considered to be beyond the scope of the present application. When implemented in hardware, the hardware can be, for example, an electronic circuit, an application specific integrated circuit (ASIC), appropriate firmware, a plug-in, a functional card, etc. When implemented in software, the elements of the present application are the program or code segments to perform a particular task. The program or code segments can be stored in a machine-readable medium, or transmitted by a carrier wave in a transmission medium or communication link.

[0205] It should be understood that the present application is not limited to the particular configurations and processes described herein and shown in the drawings. For the sake of brevity, conventional methods and systems will not be described in detail. In the above-described embodiments, several specific steps are described and illustrated in order to provide a thorough understanding of the present application. However, the process of the present application can be practiced with less than all of the described and illustrated steps, or in a different order than that described and illustrated. Numerous modifications and adaptations will be apparent to those skilled in the art in view of the foregoing description.

[0206] In the present application, features described and / or illustrated in connection with one embodiment can be used in the same or a similar way or in conjunction with or in place of features of another embodiment.

[0207] The above description is merely illustrative of the application, and is not intended to limit the scope of the application. Various modifications and changes can be made by persons of ordinary skill in the art, which should be interpreted as falling within the scope of the application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included in the scope of the present application.

Claims

1. An IoT blockchain sharding method, characterized in that, The method is executed by an agent of an Internet of Things blockchain network; when performing cross-shard transactions, the Internet of Things blockchain network introduces a relay block to relay the cross-shard transactions from a source shard to a target shard, and completes subsequent operations of the transactions in the target shard; in each epoch, the method comprises the following steps: The nodes are selected to join the Internet of Things blockchain network by a preset mechanism, and based on the computing resources of each node and the scores in the trust scorer table, a plurality of nodes are selected to constitute a policy committee; The policy committee acts as an agent to construct the number of pool transactions, the number of completed transactions and the total number of transaction transactions as a first state space, and construct the epoch length, the block size, the number of shards and the block interval as a first action space; the shard state of each node in the Internet of Things blockchain network is constructed as a second state space, and the shard adjustment of each node is constructed as a second action space; The policy committee acts as an agent, and based on a first actor network of deep reinforcement learning, the parameters of the first state space are input and the prediction of the first action in the first action space is output; a first evaluation network is used to predict a first target value of transaction throughput reward according to the first action, and the first target value introduces a first target network to predict the transaction throughput reward as a stable reference; the structure of the first target network is consistent with the first actor network and the first evaluation network; The policy committee acts as an agent, and based on a second actor network of deep reinforcement learning, the parameters of the second state space are input and the prediction of the second action in the second action space is output; a second evaluation network is used to predict a second target value of cross-shard transaction optimization reward according to the second action, and the second target value introduces a second target network to predict the cross-shard transaction optimization reward as a stable reference; the structure of the second target network is consistent with the second actor network and the second evaluation network; The first target value is used to update the parameters of the first actor network and the first evaluation network, and the second target value is used to update the parameters of the second actor network and the second evaluation network; and the first target network and the second target network are soft updated; The policy committee executes the first action, and when the number of cross-shard transactions in the Internet of Things blockchain network is higher than a set number of the total number of transaction transactions, the second action is executed.

2. The IoT blockchain sharding method of claim 1, wherein, Based on the computing resources of each node and the scores in the trust scorer table, a plurality of nodes are selected to constitute a policy committee, comprising: The historical behavior reputation is calculated, and the calculation formula is: wherein, represents the historical behavior reputation of node i, represents the number of transactions handled by node i in time period t, represents the total number of transactions in the blockchain by all nodes, represents the online duration of node i, represents the current time of the system, represents the time when node i registered to the blockchain; , is a weight coefficient, and ; The computing power reputation is calculated, and the calculation formula is: wherein, represents a hash power reputation of a node i, represents a cpu performance value of a node, represents a network latency performance value of a node, represents a network bandwidth performance value of a node, represents a storage speed performance value of a node; total cpu performance value of all nodes, represents a total network latency performance value of all nodes, represents a total network bandwidth performance value of all nodes, represents a total storage speed performance value of all nodes. , , , is a weight coefficient, and ; The calculation formula of the comprehensive reputation of the node is: wherein, the overall reputation of node i, and is a weight coefficient, and .

3. The IoT blockchain sharding method of claim 1, wherein, The preset mechanism comprises a proof of work mechanism, a proof of stake mechanism or a reputation mechanism; the first actor network, the first evaluation network, the second actor network and the second evaluation network all adopt a deep neural network structure.

4. The IoT blockchain sharding method of claim 1, wherein, The calculation formula of the transaction throughput reward is: wherein, represents the transaction throughput reward, represents the number of shards, represents the epoch length, represents the out-block interval, represents the block size, represents the average number of redundant transactions.

5. The IoT blockchain sharding method of claim 4, wherein, The calculation formula of the cross-shard transaction optimization reward is: ; wherein, represents the cross-shard transaction optimization reward, represents the number of intra-shard transactions, represents the number of cross-shard transactions, A represents the total number of transactions, and U represents the load imbalance degree; , and are parameters.

6. The IoT blockchain sharding method of claim 1, wherein, The first target value introduces a prediction of a first target network rewarding the transaction throughput as a stable reference, and the calculation formula is: ; wherein, represents the first target value, represents a current value of the transaction throughput reward, is a first discount factor, represents a predicted value of the transaction throughput reward by the first target network; represents a state of the i+1 epoch, represents parameters of a first target actor network in the first target network, represents parameters of a first target critic network in the first target network; represents an action output by the first target actor network in parameters under conditions for the state The second target value introduces a prediction of a second target network rewarding the cross-shard transaction optimization as a stable reference, and the calculation formula is: ; wherein, represents the second target value, represents a current value of the cross- shard transaction optimization reward, is a second discount factor, represents a predicted value of the cross- shard transaction optimization reward by the second target network; represents a state of the i+1 epoch, represents parameters of a second target actor network in the second target network, represents parameters of a second target critic network in the second target network; represents an action output by the first target actor network in parameters under the condition of a state .

7. The IoT blockchain sharding method of claim 6, wherein, According to the first target value, the first actor network and the first evaluation network are updated, including: The parameters of the first evaluation network are updated based on the first target value, and the expression is: wherein, denotes a weight parameter of the first evaluation network, denotes the first target value, denotes a learning rate, denotes a weight gradient of the first evaluation network, denotes the first target value currently estimated by the first evaluation network; denotes a state of the first state space at the i-th epoch, denotes an action of the first action space at the i-th epoch; The parameters of the first actor network are updated, and the expression is: ; wherein, denotes a weight parameter of the first actor network, denotes a learning rate, denotes a weight gradient of the first actor network, denotes the first target value currently estimated by the first critic network, denotes a policy function of the first actor network; According to the second target value, the second actor network and the second evaluation network are updated, including: The parameters of the second evaluation network are updated based on the second target value, and the expression is: wherein, denotes a weight parameter of the second evaluation network, denotes the second target value, denotes a learning rate, denotes a weight gradient of the second evaluation network, denotes the second target value currently estimated by the second evaluation network; denotes a state of the second state space at the i-th epoch, denotes an action of the second action space at the i-th epoch; The parameters of the second actor network are updated, and the expression is: ; wherein, denotes a weight parameter of the second actor network, denotes a learning rate, denotes a weight gradient of the second actor network, denotes the second target value currently estimated by the second critic network, denotes a policy function of the second actor network.

8. The IoT blockchain sharding method of claim 7, wherein, The first target network and the second target network are soft-updated, including: The weight update expression of the first target actor network is: ; wherein, denotes a weight parameter of the first target actor network, denotes a soft update coefficient, denotes the current weight parameter of the first actor network; The weight update expression of the first target evaluation network is: ; wherein, denote the weight parameters of the first target evaluation network, denote the soft update coefficients, denote the current weight parameters of the first evaluation network; The weight update expression of the second target actor network is: ; wherein, denotes a weight parameter of the second target actor network, denotes the soft update coefficient, denotes the current weight parameter of the second actor network; The weight update expression of the second target evaluation network is: ; wherein, denotes the weight parameters of the second target evaluation network, denotes the soft update coefficient, denotes the current weight parameters of the second evaluation network.

9. A computer readable storage medium having stored thereon a computer program or instructions, characterized in that, The computer program or instructions are executed by the processor to realize the steps of the method in any one of claims 1 to 7.

10. A computer program product comprising computer programs or instructions, characterized in that, The computer program or instructions are executed by the processor to realize the steps of the method in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Block chain stable fragmentation method based on deep reinforcement learning and reputation mechanism

    CN116506444A

  • Block chain fragmentation method and system based on multi-modal behavior information

    CN117082077A