Method and node for storing state data in block chain system
By using the erasure coding algorithm in the blockchain system to slice and store the state data in multiple nodes, the problem of state data occupying a large amount of storage space in the blockchain system is solved, and the effect of reducing storage resource requirements and storage costs is achieved.
Patent Information
- Application Number
- CN202510125342.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-26
- Publication Date
- 2025-05-27
AI Technical Summary
State data on the chain in the blockchain system occupies a lot of storage space, resulting in a linear increase in storage costs.
By obtaining the status data packet corresponding to the target block, the state data is sharded and stored in multiple nodes using an erasure coding algorithm, and each node stores one shard, thereby reducing the need for storage resources.
It effectively reduces the demand for storage resources in blockchain systems, reduces storage costs, and improves storage efficiency.
Smart Images

Figure CN120045624A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of this specification belong to the technical field of blockchain, and particularly relate to a method and a node for storing state data in a blockchain system. Background Art
[0002] Blockchain is a new application mode of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanism, and encryption algorithms. In a blockchain system, data blocks are combined into a chained data structure in a sequential connection manner according to the time sequence, and it is a distributed ledger that is guaranteed to be immutable and unforgeable in a cryptographic manner. Due to the characteristics of blockchain such as decentralization, immutability of information, and autonomy, blockchain has received more and more attention and applications.
[0003] In a blockchain system, the on-chain state data occupies a relatively large storage space. Through the optimization of the underlying blockchain storage engine system and by means of cheap media, etc., the storage costs of each node in the blockchain are reduced to a certain extent. However, since each participating party stores a complete copy of the on-chain state data respectively, the data content stored by each participating party is exactly the same. Although this simple copy method easily meets the requirements of Byzantine fault tolerance, the storage cost increases linearly as the number of participating parties in the cluster increases. Summary of the Invention
[0004] The purpose of the present invention is to provide a method for storing state data in a blockchain system to reduce the demand for storage resources in the blockchain system.
[0005] The first aspect of this specification provides a method for storing state data in a blockchain system. The blockchain system includes n nodes, and at least m of the n nodes are valid nodes. The method is executed by any node and includes:
[0006] Obtain a state data packet corresponding to a target block. The state data packet includes at least the state values of multiple first accounts updated by the target block, and the state values of the multiple first accounts are arranged in sequence based on the strings included in the multiple first accounts;
[0007] According to the erasure code algorithm, obtain a target slice corresponding to itself from among n slices based on at least part of the data in the state data packet, where the n slices include t data slices and s parity slices, the n slices respectively correspond to the n nodes, and t is less than or equal to m;
[0008] Store the target slice.
[0009] The second aspect of this specification provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed on a computer, the computer is made to execute the method described in the first aspect.
[0010] The third aspect of this specification provides a blockchain node, including a memory and a processor. An executable code is stored in the memory. When the processor executes the executable code, the method described in the first aspect is implemented.
[0011] In the solution provided in the embodiments of this specification, multiple shards of a stripe are obtained based on the state data corresponding to a block, and each node stores one shard. In this way, the storage resources required in the blockchain system are greatly saved. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] To more clearly illustrate the technical solutions of the embodiments of this specification, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments recorded in this specification. For those of ordinary skill in the art, other drawings can also be obtained based on these drawings without creative efforts.
[0013] Figure 1 It is an architecture diagram of a blockchain system in an embodiment;
[0014] Figure 2 It is a schematic diagram of the block storage structure in an embodiment;
[0015] Figure 3 It is a schematic diagram of the block storage structure in an embodiment;
[0016] Figure 4 It is a schematic diagram of an MPT tree in an embodiment;
[0017] Figure 5 It is a schematic diagram of the block storage structure in an embodiment;
[0018] Figure 6 It is a schematic diagram of the MPT tree in the embodiments of this specification;
[0019] Figure 7 It is a flowchart of a method for storing state data in a blockchain system in the embodiments of this specification;
[0020] Figure 8 It is a schematic diagram of the process of storing state data in the embodiments of this specification;
[0021] Figure 9 It is a schematic diagram of splitting a state data packet in the embodiments of this specification;
[0022] Figure 10Schematic diagram of the process of storing status data in the embodiments of this specification;
[0023] Figure 11 Flowchart of the method for reading status data in the blockchain system in the embodiments of this specification. Detailed implementation manners
[0024] In order to enable those skilled in the art of this technology to better understand the technical solutions in this specification, the following will clearly and completely describe the technical solutions in the embodiments of this specification in conjunction with the accompanying drawings in the embodiments of this specification. Obviously, the described embodiments are only a part of the embodiments of this specification, rather than all the embodiments. Based on the embodiments in this specification, all other embodiments obtained by those of ordinary skill in the art without making creative efforts shall fall within the scope of protection of this specification.
[0025] Figure 1 Shows the architecture diagram of the blockchain system in one embodiment. As Figure 1 shown, the blockchain system includes N nodes, Figure 1 schematically shows nodes 1 - 8. The connections between the nodes schematically represent the connections between the nodes and are used to transmit data between the nodes. Among them, each node in the blockchain system can generate the same state in the blockchain system by executing the same transaction, and each node in the blockchain system can store the same state database.
[0026] A transaction in the blockchain field can refer to a task unit that is executed and recorded in the blockchain system. A transaction usually includes a sending field (From), a receiving field (To), and a data field (Data). Among them, in the case where the transaction is a transfer transaction, the From field represents the account address that initiates the transaction (i.e., initiates the transfer task to another account), the To field represents the account address that receives the transaction (i.e., receives the transfer), and the Data field includes the transfer amount.
[0027] The blockchain system can provide the function of smart contracts. A smart contract on the blockchain system is a contract that can be triggered and executed by a transaction on the blockchain system. A smart contract can be defined in the form of code. Invoking a smart contract in the blockchain system is to initiate a transaction pointing to the smart contract address, so that each node in the blockchain system runs the smart contract code distributively.
[0028] One of the decentralized features that distinguish blockchain technology from traditional technologies is that accounting is performed on each node, or distributed accounting, rather than traditional centralized accounting. For a blockchain system to become a decentralized, honest, and trustworthy system that is difficult to break, publicly accessible, and has immutable data records, it is necessary to achieve the security, clarity, and irreversibility of distributed data records in the shortest possible time. In different types of blockchain networks, in order to maintain the consistency of the ledger among the nodes that record the ledger, a consensus algorithm is usually used to ensure this, that is, the consensus mechanism. For example, a consensus mechanism at the block granularity can be implemented among blockchain nodes. For instance, after a block is generated at a node (such as a certain unique node), if this generated block is recognized by other nodes, it means that other nodes have recorded the same block. The consensus mechanism is a mechanism by which blockchain nodes reach a consensus across the network on block information (or block data), which can ensure that the latest block is accurately added to the blockchain. Currently, the mainstream consensus mechanisms include: Proof of Work (POW), Proof of Stake (POS), Delegated Proof of Stake (DPOS), Practical Byzantine Fault Tolerance (PBFT) algorithm, etc. Among them, in various consensus algorithms, usually after a preset number of nodes reach an agreement on the data to be consensus (i.e., the consensus proposal), the consensus on the consensus proposal is determined to be successful. Specifically, in the PBFT algorithm, assuming that when at most f nodes fail, if there are a total of at least 3f + 1 nodes, that is, at least 2f + 1 valid nodes, it can ensure security and liveness in the system.
[0029] In a blockchain system, accounts usually can be of two types:
[0030] Contract account: Stores the executed smart contract code and the values of the state in the smart contract code, and usually can only be activated by being called by an external account;
[0031] Externally owned account: The account of the user, such as the account of an Ether owner.
[0032] The design of external accounts and contract accounts is actually a mapping from the account address to the account state. The state of an account usually includes fields such as Nonce, Balance, Storage root, CodeHash, etc. Nonce and Balance exist in both external accounts and contract accounts. The CodeHash and Storage root attributes are generally only valid for contract accounts.
[0033] Nonce: Counter. For an external account, this number can represent the number of transactions sent from the account address; for a contract account, it can be the number of contracts created by the account.
[0034] Balance: The amount of Ether owned by this address.
[0035] Storage root: The hash of the root node of an MPT tree that organizes the storage of the state variables of a contract account.
[0036] CodeHash: The hash value of the smart contract code. For a contract account, this is the hash value of the smart contract; for an external account, since it does not include a smart contract, the CodeHash field can generally be an empty string / a string of all zeros.
[0037] MPT stands for Merkle Patricia Tree, which is a tree structure that combines Merkle Tree and Patricia Tree (a more space-saving Trie tree). The Merkle Tree algorithm calculates a Hash value for each transaction, then connects them in pairs and calculates the Hash again until the top-level Merkle root. In Ethereum, an improved MPT tree, such as a 16-way tree structure, is usually simply referred to as the MPT tree.
[0038] The data structure of the Ethereum MPT tree includes the state trie. The state trie contains key-value pairs (also written as key-value, abbreviated as k-v or kv) of the storage content corresponding to each account in the Ethereum network. The "key" in the state trie can be a 160-bit identifier (such as the address of an Ethereum account or a part of the hash value of the address, hereinafter collectively referred to as the account address), and this account address is distributed in the storage from the root node to the leaf nodes of the state trie. The "value" in the state trie is generated by encoding the information of the Ethereum account (using the Recursive-Length Prefix encoding (RLP) method). As mentioned above, for an external account, the value includes nonce and balance; for a contract account, the value includes nonce, balance, codehash, and storageroot.
[0039] The contract account is used to store the state related to the smart contract. After the smart contract is deployed on the blockchain, a corresponding contract account will be generated. This contract account generally has some states, which are defined by the state variables in the smart contract and new values are generated when the smart contract is created and executed. The so-called smart contract usually refers to a contract that can automatically execute terms defined in the form of code in the blockchain environment. Once an event triggers the terms in the contract (meeting the execution conditions), the code can be automatically executed. In the blockchain, the relevant state of the contract is saved in the storage trie, and the hash value of the root node of the storage trie is stored in the above-mentioned storage root, so that all the states of the contract are locked under the contract account through the hash. The storage trie is also an MPT tree structure, which stores the key-value mapping from the state address to the state value. The partial information from the root node to the leaf node of the storage trie tree is arranged in order to store the address of a state, and the value of the state is stored in this leaf node.
[0040] As Figure 2 In some blockchain data storages as shown, the block header of each block includes several fields, such as the previous block hash previous_Hash (Prev Hash in the figure), nonce (in some blockchain systems this nonce is not random, or in some blockchain systems the nonce in the block header is not enabled), timestamp, block number Block Num, state root hash State_Root, transaction root hash Transaction_Root, receipt root hash Receipt_Root, etc. Among them, the Prev Hash in the block header of the next block (such as block N + 1) points to the previous block (such as block N), that is, the hash value of the previous block. In this way, the next block is locked to the previous block through the block header on the blockchain. Among them, State_Root, Transaction_Root, and Receipt_Root lock the state set, transaction set, and receipt set respectively. The state set, transaction set, and receipt set organize states, transactions, and receipts in the form of trees respectively. Generally, they can be the same tree structure or different tree structures. For example, in Ethereum, the same MPT structure is adopted. In some tree structures including the state set of smart contracts such as Ethereum, it includes a two-level MPT structure: the leaf nodes of the upper-level MPT structure include two types, external accounts and contract accounts; each contract account among them includes the lower-level MPT structure, and the leaf nodes of the lower level include the values of the states in the contract account.
[0041] Figure 3It is a schematic diagram of the structure for blockchain data storage. Taking Ethereum as an example, it can be combined with Figure 2 As shown, state_root is the hash value of the root of the MPT tree composed of the states of all accounts in the current block. That is, what points to state_root is a state tree state trie in the form of an MPT. The root node of this MPT tree is generally an Extension Node or a Branch Node, and what is stored in state_root is generally the hash value of this root node. The root node can be connected to one or more layers of Extension Node / Branch Node below, and these multi-layer tree nodes can be collectively referred to as Internal Node. A part of the values in each node from the root node of this MPT to the leaf node can be concatenated in sequence to form an account address and serve as the key, and the account information stored in the leaf node is the value corresponding to this account address. In this way, a key-value pair is formed. This key can also be a part after sha3(Address), that is, a part of the hash value of the account address (the hash algorithm uses the sha3 algorithm for example), and the value it stores can be rlp(Account), that is, the rlp encoding of the account information. Among them, the account information is
[0042] a quadruple composed of [nonce, balance, storageRoot, codeHash]. As mentioned above, for an external account, generally there are only two items, nonce and balance, while the storageRoot and codeHash fields default to storing an empty string / all 0 strings. That is to say, an external account does not store a contract, nor does it store the state variables generated after the contract is executed. A contract account generally includes Nonce, Balance, Storage root, CodeHash. Among them, Nonce is the transaction counter of this contract account; Balance is the account balance; Storage root corresponds to another MPT, and through Storage root, the information related to the contract state can be linked; CodeHash is the hash value of the contract code. Whether it is an external account or a contract account, its account information is generally located in a separate Leaf Node. From the Extension Node / Branch Node of the root node to the Leaf Node of each account, there may be several Branch Nodes and Extension Nodes in the middle.
[0043] The state trie can be a tree in the form of a MPT, generally a 16-way tree, that is, each level can have at most 16 child nodes. For an Extension Node, which is used to store a common prefix, it generally has 1 child node, and this child node can be a Branch Node. For a Branch Node, it can have at most 16 child nodes, which may include Extension Nodes and / or Leaf Nodes.
[0044] Among them, for a contract account in the state trie, its storage_Root points to another tree in the form of a MPT, which stores the data of the state variables involved in the contract execution. The MPT-form tree pointed to by this storage_Root is the Storage Trie, that is, the hash value of the root node of the Storage Trie. Generally, the Storage Trie tree also stores key-value pairs. The key indicates the address of the state variable, and its value can be the result obtained after processing the position where the state variable is declared in the contract (the value starting from 0) according to certain rules. For example, it can be sha3(the position where the state variable is declared), or sha3(the contract name + the position where the state variable is declared). The value is used to store the value of the state variable (for example, the value encoded by RLP). The part of the data stored on the path from the root node through the intermediate nodes to the leaf node is concatenated to form the key, and the value is stored in the leaf node. As mentioned above, this Storage trie can also be a tree in the form of a MPT, generally also a 16-way tree, that is, for a Branch Node, it can have at most 16 child nodes, and these child nodes may include Extension Nodes and / or Leaf Nodes. And for an Extension Node, it generally can have 1 child node, and this child node can be a Branch Node or a Leaf Node.
[0045] For example Figure 3In the Leaf Node Account P of the state Trie, this account is a contract account, and its Storage Root locks all the states in the contract storage. These states are organized as an MPT tree, and the tree structure is like the Storage trie linked by this Storage Root. In this linked Storage trie, taking Leaf Node StateVariable N as an example, if it is the value of storedData in the aforementioned contract code example, then its key is sha3(the declared position of storedData, that is, line 2 of the code), and its value is s (for simplicity, the encoding format of the value is omitted here, such as RLP, and the same will be omitted subsequently). Among them, the values of the keys are distributed in order from the root node to the leaf node (that is, Leaf Node Variable N) of the storage Trie.
[0046] For another example, Figure 3 In the Leaf Node Account C in the state Trie, this account is an external account, and its key is sha3(Address C), that is, the hash value of the address of account C (the hash algorithm uses the sha3 algorithm for example), and the value it stores can be (Account), where the account information Account is a binary tuple composed of [nonce, balance]. As mentioned above, since Account C is an external account, its account information is nonce and balance (the codehash and storage root are omitted here, and the same applies hereinafter). For example, for an external account, its nonce is 20 and the Balance is 4550, then in the leaf node Leaf Node State Variable C, nonce = 20 and balance = 4550 are stored. And with the address of Account C as the key, its values are distributed in order from the root node to the leaf node (that is, Leaf NodeVariable C) of the state Trie.
[0047] These states, including the k-v of external accounts and the k-v of contract accounts, are ultimately stored in the database. The storage in the database does not directly store the states of these accounts, that is, it does not directly store the k-v of these accounts, but stores the k-v values of each tree node itself.
[0048] Such as Figure 4As shown in the example, in the MPT structure of the upper level, for the leaf node A1, the key of this leaf node is composed by sequentially combining a7 in the shared nibble of the root node A8 (Extension Node) - slot 1 of the intermediate node A7 (Branch Node) - 1335 at the key - end in the leaf node A1, that is, a711335. In this leaf node, Balance = 45.0 ETH and Nonce = n1 are stored. For the leaf node A2, the key of this leaf node is composed by sequentially combining a7 in the shared nibble of the root node A8 (Extension Node) - slot 7 of the intermediate node A7 (Branch Node) - d3 in the shared nibbles of the node A6 (Extension Node) - slot 3 in the intermediate node A5 (Branch Node) - 7 at the key - end in the leaf node A2, that is, a77d337. In this leaf node, Balance = 1.00 WEI and Nonce = n2 are stored. For the leaf node A3, the key of this leaf node is composed by sequentially combining a7 in the shared nibble of the root node A8 (Extension Node) - slot f of the intermediate node A7 (Branch Node) - 9365 at the key - end in the leaf node A3, that is, a7f9365. In this leaf node, Balance = 1.1 ETH and Nonce = n3 are stored. For the leaf node A4, the key of this leaf node is composed by sequentially combining a7 in the shared nibble of the root node A8 (Extension Node) - slot 7 of the intermediate node A7 (Branch Node) - d3 in the shared nibbles of the node A6 (Extension Node) - slot 9 in the intermediate node A5 (Branch Node) - 7 at the key - end in the leaf node A4, that is, a77d397. In this leaf node, Balance = 0.12 ETH, Nonce = n4, CodeHash = c1, and Storage root = s1 are stored. s1 can be H(A10), that is, the hash value of the root node A10 of the next - level tree. Among them, the leaf nodes of A1, A2, and A3 store information of external accounts, and the leaf node of A4 stores information of contract accounts. For a contract account, it contains the next - level MPT, forming a Storage Trie for storing the state variables in this contract account.
[0049] As Figure 4As shown in the example, in the MPT structure of the next level, for leaf node A11, through slot 3 in root node A10 (Branch Node) - key-end 35b2e4 in leaf node A11, they are sequentially combined to form the key of this leaf node, which is 335b2e4. In this leaf node, "Zhang San_A = 20" is stored. For example, it means that the share of type A digital assets defined in the contract belonging to Zhang San is 20, that is, the balance of Zhang San's type A assets is 20. For leaf node A12, through slot 7 in root node A10 (Branch Node) - key-end c25988 in leaf node A12, they are sequentially combined to form the key of this leaf node, which is 7c25988. In this leaf node, "Li Si_B = 20" is stored. For example, it means that the share of type B digital assets defined in the contract belonging to Li Si is 50, that is, the balance of Li Si's type B assets is 50. For leaf node A15, through slot f in root node A10 (Branch Node) - shared nibble a in intermediate node A13 (Extension Node) - slot 6 in intermediate node A14 (Branch Node) - key-end be33 in leaf node A15, they are sequentially combined to form the key of this leaf node, which is fa6be33. In this leaf node, "storedData = s" is stored. For leaf node A16, through slot f in root node A10 (Branch Node) - shared nibble a in intermediate node A13 (Extension Node) - slot 9 in intermediate node A14 (Branch Node) - key-end 9365 in leaf node A16, they are sequentially combined to form the key of this leaf node, which is fa99365. In this leaf node, "Wang Wu_A = 35" is stored. For example, it means that the share of type A digital assets defined in the contract belonging to Wang Wu is 35, that is, the balance of Wang Wu's type A assets is 35.
[0050] In the composition of the nodes of the above MPT tree, the prefix "prefix" is used to represent the type of tree node. For example, 0 represents an Extension Node containing an even number of shared nibbles (shared half-bytes), 1 represents an Extension Node containing an odd number of shared nibble(s), 2 represents a Leaf Node containing an even number of nibbles, and 3 represents a Leaf Node containing an odd number of nibble(s).
[0051] In the above node composition, the hash value of the overall content of the next tree node is filled into the corresponding position of the previous tree node. In the database, the key-value mapping of each tree node is actually stored, where the value includes the content stored in this tree node, and the corresponding key is the hash value of the overall content of this tree node. In this way, the tree node k-v actually stored in the database is as follows in the table:
[0052]
[0053] Table 1 The tree node k-v actually stored in the database
[0054] In Table 1 above, H() is used to represent the hash calculation. In this way, the hash value of the next tree node is anchored in the previous tree node. Through such layer-by-layer hashing, the root hash of the entire state trie tree is obtained, and this root hash is locked into the state root field of the block header.
[0055] Persons in this field know that the essence of a blockchain is a chain structure composed of a series of data blocks (blocks) associated by hashes. Each block has the corresponding world state of that block, but each block does not necessarily contain the complete world state, and in most cases, it contains a reference to the previous world state. The following explains this point:
[0056] In the block structure of Ethereum, when the value in a key-value changes, the content of the leaf node of the corresponding MPT tree changes. Correspondingly, the relevant hash values stored in a series of intermediate nodes and the root node above this leaf node will also change. In the structure shown in Table 1, the changed leaf node, intermediate nodes, and root node will also be added to the end of Table 1 in an appended manner. At the same time, it is possible that the values of many state variables in several state variables stored in a contract do not change.
[0057] As Figure 5 shown in the Ethereum block structure diagram, this Figure 5 and Figure 2The difference lies in that the root nodes and intermediate nodes of the two-level Merkle tree under State Root in the (N + 1)-th block reference the state variables that have not changed in the N-th block. For example, in the contract storage of a contract, the value of the same state variable changes from 29 in the N-th block to 45 in the (N + 1)-th block. Assuming that other state variables have not changed, in the (N + 1)-th block, except for the leaf node storing the value changed to 45, and the associated intermediate nodes and intermediate nodes, the other unchanged leaf nodes and intermediate nodes in the (N + 1)-th block directly reference the same leaf nodes and intermediate nodes in the N-th block. Here, two adjacent blocks, the N-th and the (N + 1)-th, are shown. It can be understood that for the block structures of multiple consecutive blocks on the blockchain, the leaf nodes and intermediate nodes in the (N + 1)-th block that have not changed compared to the previous blocks directly reference the same leaf nodes and intermediate nodes that were last updated to this value. The same leaf nodes and intermediate nodes that were last updated to this value are, for example, the same leaf nodes and associated intermediate nodes that were updated to this value only in the (N - 1)-th block, the same leaf nodes and associated intermediate nodes that were updated to this value only in the 5-th block, and so on. Here, different blocks can also be called different versions, and the block number can be the version number. This reference is achieved by the upper-level nodes recording the hash values of the lower-level nodes. The advantage of this is obvious, that is, the unchanged leaf nodes and intermediate nodes are not repeatedly stored in the underlying database.
[0058] Here, still taking the change in the state in the same contract storage in the (N + 1)-th version compared to the N-th version as a simple illustration. For example, if the content of a certain tree node changes, its hash will accordingly change, and the MPT uses the hash as an index in the database, thus achieving that for each value, there is a definite record in the database. And the MPT associates parent and child nodes based on the node hashes. Therefore, whenever the content of a node changes, ultimately for the parent node, only a hash index value changes; the content of the parent node also changes accordingly, generating a new and higher-level parent node, and recursively passing this influence to the root node. Eventually, one change corresponds to creating a new path from the modified node to the root node, while the old node can still be accessed through the old path based on the old root node.
[0059] As Figure 5As shown, when the content of a node in the MPT changes from 29 in version N to 45 in version N+1, assuming that other states in the contract state remain unchanged between the two versions, the changed 45 creates a new path to the contract root node (Storage root) in the world state of version N+1. And for the other leaf nodes and intermediate nodes that remain unchanged, the leaf nodes and / or intermediate nodes in the previous version (i.e., the version with block number N) are reused through hash pointers to construct a new MPT tree for the contract storage in version N+1, while still retaining its old path. As Figure 5 shown in Figure 5 , for the leaf nodes / intermediate nodes in the same contract storage where the values in version N+1 have not changed compared to version N, including V, U, T, S, etc., in the storage of this contract in version N+1, in addition to creating a new path from 45 to Storage root, the Storage root in version N+1 is hashed to V, U, T, S in version N, but does not include the leaf node with the value of 29 and the related intermediate nodes. Assuming that this contract was deployed in version N, the CodeHash is similar. The contract code corresponding to the CodeHash exists in the state database (StateDB) corresponding to version N, and the CodeHash in version N+1 also points to the contract code in version N.
[0060] In addition to the contract storage, similarly, other account storages (including the storages of external accounts and contract accounts) can also adopt a similar approach, such as Figure 5 shown in Figure 5 . For the account storages in version N+1 that have not changed compared to version N, the State Root in the block header of version N+1 points to the storages of external accounts or contract accounts such as L, M, N, P, Q, R in version N, and only the changed leaf nodes and related intermediate nodes are stored in the account storage corresponding to version N+1. Details are not elaborated one by one.
[0061] In fact, many of the states under the current block are likely to have been updated in a previous block and have remained unchanged until the current block. This is similar to Ethereum and is also called the current state, which is stored in the current database (currentDB). The currentDB is used to store data of the latest "current state", such as the latest account state, contract code, and contract storage. Moreover, as the blocks grow, the data in the currentDB can be continuously modified.
[0062] Regarding the problem that the on-chain status data occupies a large amount of storage space, considering that in the above various consensus algorithms, it is not necessary for each node to store a complete copy of the status data. Therefore, based on the erasure coding algorithm (EC), the status data corresponding to the target block can be scattered and stored on multiple nodes. The status data at least includes the status values of multiple accounts updated by the target block, or the status data can include all the status data corresponding to the target block. Wherein, when the target block is the current latest block, all the status data corresponding to the target block is the above-mentioned current state. Each node in the blockchain still stores the Merkle tree (such as the MPT tree) corresponding to the target block. Figure 6 It is a schematic diagram of the MPT tree stored in the embodiments of this specification, as Figure 6 shown. The difference from the MPT tree shown in Figure 4 is that in the Figure 6 MPT tree shown, each leaf node only includes the hash value H(value) of the value of the account, rather than the value itself. This hash value H(value) can be used to verify the read account status value. Here, each leaf node can be multiple leaf nodes corresponding to multiple accounts updated by the target block.
[0063] Among them, in the EC (Erasure Coding) algorithm, n shards can be generated based on the target data to be stored. The n shards include t data shards obtained by evenly slicing the target data and s parity shards generated based on the t data shards. The n shards can be stored in different geographical locations. In the case where some shards are damaged, the target data can be recovered through any t of the n shards. Among them, a set of data shards and parity shards of the EC shards form an EC stripe. That is to say, in the EC algorithm, the larger s is, the greater the fault tolerance rate of the system is, and at the same time, the higher the redundancy of the system storing data is.
[0064] Combined with the consensus algorithm in the blockchain system, multiple shards of the target data can be distributed to each blockchain node, and n and t are made to meet the limitations of the consensus algorithm. For example, for the PBFT algorithm, n = 3f + 1 and t is equal to or less than 2f + 1. Thus, when there are at most f failed nodes in the PBFT algorithm, the correct recovery of the target data can be guaranteed. It can be understood that the solution in the embodiments of this specification is not limited to being used in the PBFT algorithm, but can be used in other consensus algorithms as long as t is equal to or less than the number m of valid nodes in the consensus algorithm.
[0065] Figure 7The following is a flowchart of a method for storing state data in a blockchain system according to an embodiment of the present specification. Assume that the blockchain system includes n nodes, and at least m of the n nodes are valid nodes. The method is executed by any one of the nodes and includes:
[0066] In step S701, obtain a state data packet corresponding to a target block. The state data packet includes at least state values of a plurality of first accounts updated by the target block. The state values of the plurality of first accounts are arranged in sequence based on the characters included in the plurality of first accounts.
[0067] In step S703, according to the erasure code algorithm, obtain the shard corresponding to itself from among n shards based on at least part of the data in the state data packet. Among the n shards, there are t data shards and s parity shards, and the n shards respectively correspond to the n nodes, where t is less than or equal to m.
[0068] In step S705, store the obtained shard corresponding to itself.
[0069] In the following, with reference to Figure 8 and Figure 10 the schematic diagram of the process of storing state data in the embodiment of the present specification shown in Figure 7 describe the
[0070] In one implementation, as Figure 8 shown, assume that the blockchain system includes 4 nodes: node 1 to node 4. According to the PBFT algorithm, in the case of successful consensus, it can be determined that there is at most 1 malicious node or faulty node among node 1 to node 4, that is, there are at least 3 valid nodes among node 1 to node 4. Therefore, by splitting at least part of the data in the state data packet into 3 or 2 data shards and storing them in different nodes, the recovery of the at least part of the data can be guaranteed. Among them, in the case of splitting the at least part of the data into 3 data shards, for this stripe, a parity shard can be generated based on the 3 data shards; in the case of splitting the at least part of the data into 2 data shards, for this stripe, 2 parity shards can be generated based on the 2 data shards.
[0071] Specifically, after each node generates the state data of, for example, block B1, node 1 to node 4 respectively store as Figure 6The MPT tree shown. Specifically, the MPT tree includes the hash values (H(value)) of the status values of multiple accounts updated in block B1, and step S701 is executed to obtain the status data packet corresponding to block B1. Specifically, the status values of multiple accounts corresponding to block B1 can be obtained first. These multiple accounts can be the status values of all accounts corresponding to block B1, or can also be the status values of multiple accounts updated in block B1. Then, the status values of these multiple accounts can be arranged according to the strings included in these multiple accounts, so as to obtain the status data packet. For example, if the strings included in each account are hexadecimal numbers, the status values of these multiple accounts can be arranged according to the magnitude order of the values corresponding to each account.
[0072] Then, each node executes step S703 and step S705, and according to the EC algorithm, based on at least part of the data in the status data packet, a shard is obtained respectively and the shard is stored. That is to say, in this process, a stripe is obtained based on at least part of the data in the status data packet, and the stripe includes three data shards D1 to D3 and one parity shard C1.
[0073] In one implementation manner, it is assumed that the status data packet includes status values Va1 to Va4, Vb1 to Vb3, Vc1 to Vc4, where Vai is the status value of each account whose first character is character a, Vbi is the status value of each account whose first character is character b, and so on. Each node can Figure 9 perform slicing on the status data packet as shown. In one implementation manner, the data in the status data packet can be sliced according to a fixed data length, and data shards D1 to D3 are obtained in sequence. For example, as Figure 9 shown, the first half of the status value Va1 to the status value Vb1 is sliced into data shard D1, the second half of the status value Vb1 to the first half of the status value Vc1 is sliced into data shard D2, and the second half of the status value Vc1 to the status value Vc4 is sliced into data shard D3. Among them, when the data volume from the second half of the status value Vc1 to the status value Vc4 does not reach the above fixed length, padding data can be supplemented after the status value Vc4 to make it up to the above fixed length.
[0074] Each node can determine a corresponding shard through negotiation. For example, node 1 corresponds to shard D1, node 2 corresponds to shard D2, node 3 corresponds to shard D3, and node 4 corresponds to shard C. Based on this corresponding relationship, when node 1 executes step S703, it can Figure 9As shown in the figure, the status data packet is sliced into data shards D1 to D3, and a correspondence table of each node and the status values in its corresponding data shard is stored to facilitate subsequent reading of the status data; and in step S705, data shard D1 is stored. Among them, node 1 can store the association between data shard D1 and the identifier of the status data corresponding to block B1 for subsequent reading of the data in shard D1.
[0075] In one implementation, node 1 can record Tables 2 and 3 shown below:
[0076] Data Overall offset (unit: KB) Length (unit: KB) Va1 0 80 Va2 80 100 Va3 180 70 Va4 250 90 Vb1 340 60 … … …
[0077] Table 2
[0078] Node Position 1 Va1,0 2 Vb1,40 … …
[0079] Table 3
[0080] Among them, Table 2 records the overall offset values of each status value in the status data packet, and this overall offset value indicates the arrangement position of each status value in the status data packet. Among them, the identifier in the "Data" column is the account corresponding to the status value.
[0081] Table 3 records the identifier of the starting status value in the data shard stored by each node, and the offset value of the starting position of this shard in the starting status value. By recording Tables 2 and 3, when node 1 queries the status value subsequently, it can determine which node the status value is stored in and the specific storage position of the status value in that node by combining Tables 2 and 3, which is convenient for querying the status value. It can be understood that the correspondence table in the embodiments of this specification is not limited to Tables 2 and 3 shown, as long as the correspondence table can indicate the status values in the data shards corresponding to each node.
[0082] Node 2 and node 3 can slice the status data packet, record the correspondence table, and store the corresponding shard in a similar manner to node 1, which will not be elaborated here.
[0083] After node 4 slices the status data packet into data shards D1 to D3 and stores the above-mentioned correspondence table, according to the EC algorithm, check shard C is calculated based on data shards D1 to D3 and stored in association with the identifier of the status data packet corresponding to block B1.
[0084] In addition to recording the correspondence table as described above, each node can also record the slicing rules of the EC algorithm and the size limit of the blocks, the list of all nodes participating in building the EC storage network, the identifier of the check block, etc. Among them, the identifier of the check block is, for example, the hash value of this check block. Each node can also store in association the block identifier of the block it stores and the storage position of this block (such as logical address, etc.).
[0085] In another implementation, each node can divide the accounts corresponding to the status data packet into multiple account partitions, and divide the status data packet into multiple data shards according to a fixed number of account partitions. For example, each node can be divided into 16 account partitions according to 16 forks under the root node in the status tree, and then the 16 account partitions can be divided into 3 data shards according to a fixed partition number such as "6". Specifically, referring to Figure 8 , assuming that the accounts corresponding to the status data packet can be divided into partition a, partition b, and partition c, where partition a includes accounts a1 to a4, partition b includes accounts b1 to b3, and partition c includes accounts c1 to c4, then the status values (Va1 to Va4) of the accounts included in partition a can be sliced into data shard D1, the status values (Vb1 to Vb3) of the accounts included in partition b can be sliced into data shard D2, and the status values (Vc1 to Vc4) of the accounts included in partition c can be sliced into data shard D3. After this slicing, the other two parts of data can be filled according to the maximum data volume of the three sliced parts of data, so that the data volumes of the three data shards are equal. In this slicing method, when each node queries the status value, it can deduce which node stores the status value to be queried according to this slicing method.
[0086] In one implementation, in the case of slicing the status data packet according to a fixed number of account partitions as described above, node 1 can obtain Va1 to Va4 from the status data packet and store them as data shard D1. Node 2 and node 3 can similarly obtain their corresponding status values from the status data packet and store the status values as data shards. There may be a situation where the data volumes of the stored data shards are not equal. For this situation, when it is necessary to restore data through the EC algorithm, the data shards can be filled first based on the data volume of the parity shard, and then the data can be restored based on the filled data shards and the parity shard. In addition, node 4 can slice the status data packet according to this partitioning method. After this slicing, the other two parts of data can be filled according to the maximum data volume of the three sliced parts of data, so that the data volumes of the three data shards are equal. The parity shard C is calculated based on the three filled data shards and the parity shard C is stored.
[0087] The following describes Figure 10 the method shown in the Figure 7 process shown.
[0088] Each node may execute step S701 as described above. After that, when each node executes step S703, it may first cut the status data packet to obtain six data shards D11 to D31 and D12 to D32, and generate a parity shard C1 based on the data shards D11 to D31. Among them, the data shards D11 to D31 and the parity shard C1 form stripe s1. Generate a parity shard C2 based on the data shards D12 to D32. Among them, the data shards D12 to D32 and the parity shard C2 form stripe s2.
[0089] Specifically, the block body may be sliced according to a fixed data volume to obtain the above six data shards, or the status data packet may also be sliced according to a fixed number of account partitions as described above. There is no limitation on this. In one implementation, first, as described above, according to the fixed number of account partitions (for example, the number is 1), Va1 to Va4 are assigned to node 1, Vb1 to Vb3 are assigned to the node, and Vc1 to Vc4 are assigned to node 3. Then, the data assigned to each node is further sliced according to a fixed data volume. For example, according to a fixed data volume, Va1 to Va4, Vb1 to Vb3, and Vc1 to Vc4 are respectively sliced into two data shards. Among them, when the data volume of the latter data shard is less than the fixed data volume, the latter data shard can be filled with data, so that six data shards can be obtained.
[0090] Each node may obtain the shards of two stripes corresponding to itself according to the above division of the status data packet. For example, node 1 obtains the data shard D11 of stripe s1 and the data shard D12 of stripe s2, node 2 obtains the data shard D21 of stripe s1 and the data shard D22 of stripe s2, node 3 obtains the data shard D31 of stripe s1 and the data shard D32 of stripe s2, and node 4 can calculate the parity shard C1 of stripe s1 based on the data shards D11, D21, and D31, and calculate the parity shard C2 of stripe s2 based on the data shards D12, D22, and D32.
[0091] After each node slices the status data packet as Figure 10 shown, it may store Table 2 above and Table 4 shown below:
[0092]
[0093]
[0094] Table 4
[0095] In addition to recording the corresponding relationship table as described above, each node can also record the splitting rules of the EC algorithm and the size limit of each block, the list of all nodes participating in building the EC storage network, the identifier of the parity block, etc. Among them, the identifier of the parity block is, for example, the hash value of the parity block. Each node can also associate and store the block identifier of the block it stores and the storage location of the block (such as the logical address, etc.).
[0096] In step S705, each node stores the obtained shards. Specifically, each node can store the obtained shards in an associated manner with the identifier of the status data corresponding to block B1 and the stripe identifier. For example, node 1 stores data shard D11 in an associated manner with the identifier of the status data corresponding to block B1 and the identifier of stripe s1, stores data shard D12 in an associated manner with the identifier of the status data corresponding to block B1 and the identifier of stripe s2, node 2 stores data shard D21 in an associated manner with the identifier of the status data corresponding to block B1 and the identifier of stripe s1, stores data shard D22 in an associated manner with the identifier of the status data corresponding to block B1 and the identifier of stripe s2, node 3 stores data shard D31 in an associated manner with the identifier of the status data corresponding to block B1 and the identifier of stripe s1, stores data shard D32 in an associated manner with the identifier of the status data corresponding to block B1 and the identifier of stripe s2, node 4 stores parity shard C1 in an associated manner with the identifier of the status data corresponding to block B1 and the identifier of stripe s1, and stores parity shard C2 in an associated manner with the identifier of the status data corresponding to block B1 and the identifier of stripe s2.
[0097] Compared with storing by generating one stripe based on the status data packet, storing by generating multiple stripes based on the block body results in less computational effort and communication volume and lower recovery cost when performing data recovery.
[0098] Figure 11 This is a flowchart of the method for reading status data in the blockchain system in the embodiments of this specification. This method can be executed by any node in the blockchain system.
[0099] As Figure 11 shown, in step S1101, a read request for the status data is received.
[0100] The client can send a query request for the status data (i.e., the status value) of, for example, account A1 to any node in the blockchain system. The query request can include the account address of account A1. Assume that the blockchain system Figure 8 or Figure 10 shown includes 4 nodes, and node 1 receives the read request. This will be used as an example for description hereinafter.
[0101] In step S1103, the first node corresponding to the status data is determined.
[0102] In one embodiment, in the case of slicing the status data packet by a fixed number of account partitions as described above, Node 1 can determine the first node corresponding to the status data according to the account partition to which Account A1 belongs.
[0103] In another embodiment, in the case of slicing the status data packet by a fixed data volume as described above, Node 1 can determine the node corresponding to the status data to be queried according to Table 2 and Table 3, or Table 2 and Table 4 in the above text. Specifically, referring to Figure 9 Table 2 and Table 3, Node 1 can first retrieve the position of the target status data in the status data packet based on Account A1 according to Table 2. For example, according to Table 2, determine the offset position of the status data Vb2 to be queried in the status data packet, and then can determine that the status data Vb2 corresponds to Node 2 according to Table 3. In this case, this Node 2 is the first node corresponding to the status data.
[0104] In one embodiment, Node 1 can determine that the status data to be read is stored in itself (i.e., Node 1). For example, referring to Figure 9 , in the case where the status data to be read is Va2, Node 1 can determine that Va2 is stored in itself according to Table 2 and Table 3. In this case, this Node 1 is the first node corresponding to the status data.
[0105] In one embodiment, Node 1 can determine that the status data to be read is stored in multiple nodes. For example, referring to Figure 8 and Figure 9 , in the case where the status data to be read is Vb1, Node 1 can determine that Vb1 is stored in Node 1 and Node 2 according to Table 2 and Table 3. In this case, Node 1 and Node 2 are the two first nodes corresponding to the status data Vb1.
[0106] In one embodiment, Node 1 can determine which node and which shard the target status data corresponds to only according to Table 3 or Table 4. Specifically, Node 1 determines which interval of the account intervals shown in Table 3 or Table 4 the Account A1 is in, so as to determine which node and which shard the target status data corresponds to.
[0107] In step S1105, read the status data from the first node.
[0108] After Node 1 determines the first node corresponding to the status data, when the first node is itself, Node 1 can directly read the status data from the shard stored in itself. For example, as shown in Figure 9 , in the case where the status data to be read is Va2, Node 1 can obtain the reading position in shard D1 corresponding to Va2 (i.e., the position from 80KB to 179KB in shard D1) according to Table 2 and Table 3, and read Va2 based on this reading position.
[0109] When the first node includes other nodes, node 1 may request to read status data from the other nodes. For example, referring to Figure 9 the strip shown, when the status data to be read is Vb1, node 1 may determine according to Table 2 and Table 3 that the first half of Vb1 is at node 1, the second half of Vb2 is at node 2, and at the position of offset 0 to 19 in shard D2. Node 1 may read the first half of Vb1 from node 1, and this reading process may refer to the description above and will not be elaborated here. Node 1 also requests node 2 to read the position of offset 0 to 19 of shard D2. After node 2 successfully reads the second half of Vb1 and returns it to node 1, node 1 thus reads the second half of Vb1.
[0110] In step S1107, the status data is returned.
[0111] After node 1 reads the status data from the first node as described above, it returns the status data to the client.
[0112] In one implementation, after reading the status data, node 1 may verify the status data using the hash value of the status value of account A1 in the status tree. If the verification passes, the status data is returned to the client.
[0113] In step S1109, reading the status data from the first node fails.
[0114] When the first node is not node 1 itself, when the first node fails, or the first node is a malicious node, or the shard stored by the first node is damaged, the first node cannot return the status data to node 1, so that node 1 fails to read the status data from the first node.
[0115] When reading the status data from the first node fails, step S1111 is executed to read at least t shards other than the shards of the first node, and the status data is restored based on the at least t shards.
[0116] Referring to Figure 8 and Figure 9In the example described above where node 1 requests to read the second half of Vb1 from node 2, if node 2 does not return status data to node 1, according to the EC algorithm, node 1 reads a slice of the status data packet from itself (i.e., data slice D1), requests node 3 to read a slice of the status data packet (i.e., data slice D3), and requests node 4 to read a slice of the status data packet (i.e., verification slice C), and performs calculations based on slice data D1, data slice D3, and verification slice C to obtain the restored data slice D2. Afterwards, node 1 can read the second half of Vb1 from the offset 0 to 19 of the restored data slice D2, and then execute step S1107 to return the status data Vb1 to the client. In one embodiment, after obtaining the status data, node 1 can use the hash value of the status value of account A1 in the status tree to verify the status data, and if the verification passes, the status data is returned to the client.
[0117] Refer to Figure 10 In the example where node 1 requests to read Vb1 from node 2, and node 2 does not return status data to node 1, since Vb1 is located in stripe s1, according to the EC algorithm, node 1 reads the shard of stripe s1 (i.e., data shard D11) of the status data of block B1 from itself, requests node 3 to read the shard of stripe s1 (i.e., data shard D31) of the status data of block B1, and requests node 4 to read the shard of stripe s1 (i.e., check shard C1) of the status data of block B1, and restores the shard of stripe s1 corresponding to node 2 (i.e., data shard D21) based on the three shards of stripe s1 (i.e., data shard D11, data shard D31, and check shard C1). Afterwards, node 1 can read Vb1 from the offset position of data shard D21 corresponding to Vb1, and execute step S1107 to return Vb1 to the client.
[0118] In the case where the client initiates a status data read request to node 1 and node 1 fails, the client can re-initiate a status data read request to any other node in the blockchain system (e.g., node 2) and indicate that node 1 has failed. Node 2 can perform similar operations as node 1. Figure 11 In the method flow shown, when it is determined that the state data is stored in node 1, step S1111 is executed to restore the state data and return it to the client.
[0119] The embodiment of the present specification also provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed in a computer, the computer is caused to execute the following Figure 7 or Figure 11 The method shown.
[0120] An embodiment of this specification also provides a blockchain node, including a memory and a processor. An executable code is stored in the memory. When the processor executes the executable code, the method as shown in Figure 7 or Figure 11 is implemented.
[0121] In the 1990s, improvements to a technology could be clearly distinguished as either hardware improvements (e.g., improvements to circuit structures such as diodes, transistors, switches, etc.) or software improvements (improvements to method flows). However, with the development of technology, many method flow improvements today can be regarded as direct improvements to hardware circuit structures. Designers almost always obtain the corresponding hardware circuit structure by programming the improved method flow into the hardware circuit. Therefore, it cannot be said that an improvement to a method flow cannot be implemented using a hardware entity module. For example, a Programmable Logic Device (PLD) (such as a Field Programmable Gate Array (FPGA)) is an integrated circuit whose logic function is determined by the user programming the device. Designers can program themselves to "integrate" a digital system onto a single PLD, without having to ask a chip manufacturer to design and fabricate a dedicated integrated circuit chip. Moreover, today, instead of manually fabricating integrated circuit chips, this programming is mostly implemented using "logic compiler" software, which is similar to the software compilers used in program development and writing. The original code before compilation also has to be written in a specific programming language, which is called a Hardware Description Language (HDL), and there is not just one type of HDL, but many, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. The most commonly used ones currently are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also be aware that by simply performing a little logical programming on the method flow using the above-mentioned several hardware description languages and programming it into an integrated circuit, it is easy to obtain the hardware circuit that implements the logical method flow.
[0122] The controller can be implemented in any suitable manner. For example, the controller can take the form of, for example, a microprocessor or a processor and a computer-readable medium storing computer-readable program code (such as software or firmware) executable by the (micro)processor, logic gates, switches, an application specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller. Examples of the controller include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art also know that in addition to implementing the controller in the form of pure computer-readable program code, it is entirely possible to logically program the method steps to enable the controller to be implemented in the form of logic gates, switches, application specific integrated circuits, programmable logic controllers, embedded microcontrollers, etc. to achieve the same function. Therefore, such a controller can be considered a hardware component, and the devices included therein for implementing various functions can also be regarded as the structures within the hardware component. Or even, the devices for implementing various functions can be regarded as either software modules for implementing the method or the structures within the hardware component.
[0123] The systems, devices, modules, or units illustrated in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a server system. Of course, this application does not exclude that with the development of future computer technologies, the computers for implementing the functions of the above embodiments can be, for example, personal computers, laptop computers, in-vehicle human-machine interaction devices, cellular phones, camera phones, smart phones, personal digital assistants, media players, navigation devices, email devices, game consoles, tablet computers, wearable devices, or any combination of these devices.
[0124] Although one or more embodiments of this specification provide method operation steps as described in the embodiments or flowcharts, more or fewer operation steps may be included based on conventional or non-creative means. The order of steps listed in the embodiments is only one way among the execution orders of numerous steps and does not represent the only execution order. When the actual device or terminal product is executed, it may be executed in the order of the method shown in the embodiments or the drawings or in parallel (for example, in a parallel processor or multi-threaded processing environment, or even in a distributed data processing environment). The term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, product or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or also includes elements inherent to such process, method, product or device. Without further limitation, there is no exclusion of additional identical or equivalent elements in the process, method, product or device comprising the said elements. For example, if terms such as first and second are used to denote names, they do not denote any particular order.
[0125] For the convenience of description, when describing the above device, it is described by dividing it into various modules according to functions. Of course, when implementing one or more of this specification, the functions of each module can be implemented in the same or multiple software and / or hardware, or the modules implementing the same function can be realized by a combination of multiple sub-modules or sub-units, etc. The device embodiments described above are only illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the couplings or direct couplings or communication connections shown or discussed with each other may be through some interfaces, and the indirect couplings or communication connections of the devices or units may be in electrical, mechanical or other forms.
[0126] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0127] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction device that implements the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 specified in one block or multiple blocks.
[0128] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 specified in one block or multiple blocks.
[0129] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.
[0130] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM), and / or non-volatile memory such as read-only memory (ROM) or flash memory (flash RAM). Memory is an example of computer-readable media.
[0131] Computer-readable media includes permanent and non-permanent, removable and non-removable media and can be implemented by any method or technology for information storage. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic disk storage, graphene storage, or other magnetic storage devices, or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media such as modulated data signals and carrier waves.
[0132] Those skilled in the art should understand that one or more embodiments of this specification can be provided as a method, a system, or a computer program product. Therefore, one or more embodiments of this specification can take the form of an all-hardware embodiment, an all-software embodiment, or an embodiment combining software and hardware aspects. Moreover, one or more embodiments of this specification can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0133] One or more embodiments of this specification can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. One or more embodiments of this specification can also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.
[0134] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for system embodiments, since they are basically similar to method embodiments, the description is relatively simple, and reference can be made to the relevant parts of the method embodiments for the related content. In the description of this specification, the description of reference terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of this specification. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0135] The above description is only for the embodiments of one or more embodiments of this specification and is not intended to limit one or more embodiments of this specification. For those skilled in the art, one or more embodiments of this specification can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of this specification shall be included within the scope of the claims.
Claims
1. A method for storing state data in a blockchain system, wherein the blockchain system includes n nodes, wherein the n nodes include at least m valid nodes, and the method is executed by any node, comprising: Obtaining a status data packet corresponding to the target block, the status data packet at least including status values of multiple first accounts updated by the target block, the status values of the multiple first accounts being arranged in sequence based on the character strings included in the multiple first accounts; According to the erasure code algorithm, based on at least part of the data in the status data packet, obtain a target shard corresponding to itself from among the n shards, wherein the n shards include t data shards and s check shards, the n shards correspond to the n nodes respectively, and t is less than or equal to m; The target shard is stored.
2. The method according to claim 1, further comprising: A corresponding relationship table is recorded, where the corresponding relationship table is used to indicate information about status data in the shards stored at each node.
3. According to the method of claim 2, the t data shards are obtained by sequentially splitting the status data packet, and the correspondence table includes the identifier of the first account corresponding to the starting position of the shard stored in each node, and the offset position of the starting position in the status value of the first account.
4. According to the method of claim 3, the status data packet is divided into multiple stripes, and at least part of the data belongs to the first stripe of the multiple stripes, and the correspondence table also includes the identifier of the stripe corresponding to the shards stored in each node.
5. According to the method of claim 4, the storing of the target shard comprises storing the target shard in association with an identifier of status data corresponding to a target block and an identifier of the first stripe.
6. The method according to claim 3, further comprising: Based on the multiple account partitions corresponding to the status data packet, the status data packet is segmented to obtain the t data shards.
7. The method according to claim 1, wherein the plurality of first accounts include the second account, and the method further comprises: receiving a read request for a status value of a second account from a client; Determining one or more first nodes corresponding to the second account; Reading a state value of the second account from the one or more first nodes; Return the status value of the second account to the client.
8. The method according to claim 7, further comprising: In the case that the one or more first nodes are faulty nodes, at least t shards are read from the remaining at least t nodes among the n nodes, and the status value of the second account is obtained based on the at least t shards.
9. The method according to claim 7 or 8, further comprising: After obtaining a status data packet corresponding to the target block, storing a status tree, the status tree including hash values of status values of each of the first accounts; After obtaining the status value of the second account, the status value of the second account is verified based on the hash value of the status value of the second account in the status tree.
10. A blockchain node, comprising a memory and a processor, wherein the memory stores an executable code, and when the processor executes the executable code, the method according to any one of claims 1 to 9 is implemented.