Management method of state data in block chain system and block chain node

By adopting index and data separation in the blockchain system, the key-value pairs are stored using primary index blocks and secondary index blocks, the index file update problem caused by data governance is solved, and query efficiency and system performance are improved.

CN120256522APending Publication Date: 2025-07-04ANT BLOCKCHAIN TECHNOLOGY (SHANGHAI) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510401565.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

In blockchain systems, when data governance causes changes in state data position information, the existing technology needs to traverse and update the index file, affecting system performance.

Method used

Using index and data separation, key-value pairs are stored through primary index blocks and multiple secondary index blocks, the data file is determined using file change information, and the location information of the target key-value pair is directly read to avoid traversal updates of the index file.

Benefits of technology

It realizes that there is no need to traverse the index file when the data location information changes, which improves query efficiency and system performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120256522A_ABST
    Figure CN120256522A_ABST
Patent Text Reader

Abstract

According to the state data management method in the block chain system, the block chain system stores a plurality of data files through a persistent storage medium, meta-information of the data files comprises a first-level index block and a plurality of second-level index blocks, the first-level index block stores a plurality of first index entries, and the second-level index blocks store a plurality of second index entries. The first index entry is used for accessing a third index entry in a secondary index block, the third index entry and a key-value pair accessed by a previous second index entry correspond to different block numbers or the third index entry is the first index entry in the secondary index block, and the method comprises the following steps: obtaining position information of a target key-value pair, comprising a first block number and a first file identifier of the first data file; if the first data file does not exist, determining a second data file according to the first file identifier and the file change information; determining a second target index block according to the first block number and a first-level index block corresponding to the second data file; and reading the target key-value pair from the second data file according to the second target index block.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of this specification belong to the field of computer technology, and particularly relate to a method for managing state data in a blockchain system and a blockchain node. Background Art

[0002] A blockchain system is a new application mode of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanism, and encryption algorithms. In a blockchain system, data blocks are combined into a chain data structure in a sequential connection manner according to the time sequence, and a distributed ledger that is tamper-proof and forgery-proof is guaranteed by cryptographic means. Due to the characteristics of decentralization, information immutability, and autonomy of the blockchain system, the blockchain system has received more and more attention and applications. Summary of the Invention

[0003] The purpose of the present invention is to provide a method for managing state data in a blockchain system and a blockchain node.

[0004] In a first aspect, a method for managing state data in a blockchain system is provided. The blockchain system stores multiple data files through a persistent storage medium. Multiple key-value pairs in the data files are stored in an orderly manner according to the block number and the key. The meta-information of the data files includes a first-level index block and multiple second-level index blocks. Multiple first index entries arranged in an orderly manner according to the block number and the key are stored in the first-level index block. Multiple corresponding second index entries are stored in the multiple second-level index blocks in the storage order of the multiple key-value pairs. The first index entry is used to access a third index entry. The third index entry corresponds to a different block number from the key-value pair accessed by the previous second index entry or is the first index entry in the second-level index block. The method includes: obtaining the position information of a target key-value pair to be queried, including a first block number of a first block and a first file identifier of a first data file, where the target key-value pair is generated by executing a transaction sequence belonging to the first block; if the first data file does not exist, determining a second data file for storing the target key-value pair at the current moment according to the first file identifier and file change information; determining a second target index block to which a second target index entry for accessing the target key-value pair belongs from the multiple second-level index blocks corresponding to the second data file according to the first block number, the target key in the target key-value pair, and the first-level index block corresponding to the second data file; and reading the target key-value pair from the second data file according to the index entry in the second target index block.

[0005] Second aspect, a blockchain node in a blockchain system and a method for managing state data in the blockchain system are provided. The blockchain system stores multiple data files through a persistent storage medium. Multiple key-value pairs in the data files are stored in an orderly manner according to block numbers and keys. The meta-information of the data files includes a first-level index block and multiple second-level index blocks. The first-level index block stores multiple first index entries arranged in an orderly manner according to block numbers and keys. The multiple second-level index blocks store their corresponding multiple second index entries in the storage order of the multiple key-value pairs. The first index entry is used to access a third index entry, and the third index entry corresponds to a different block number from the key-value pair accessed by the previous second index entry or is the first index entry in the second-level index block. The blockchain node includes: a location acquisition unit configured to acquire location information of a target key-value pair to be queried, where the location information includes a first block number of a first block and a first file identifier of a first data file, and the target key-value pair is generated by executing a transaction sequence belonging to the first block; a file determination unit configured to, if the first data file does not exist, determine a second data file for storing the target key-value pair at the current moment according to the first file identifier and file change information; an entry retrieval unit configured to, according to the first block number, the target key in the target key-value pair, and the first-level index block corresponding to the second data file, determine a second target index block to which a second target index entry for accessing the target key-value pair belongs from the multiple second-level index blocks corresponding to the second data file; and a query processing unit configured to read the target key-value pair from the second data file according to the index entry in the second target index block.

[0006] Third aspect, a computing device is provided, including a memory and a processor. A computer program is stored in the memory. When the processor executes the computer program, the method described in the first aspect is implemented.

[0007] Fourth aspect, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed in a computing device, the computing device executes the method described in the first aspect.

[0008] The technical solution provided by the embodiments of this specification can achieve the following: On the basis of storing the state data of a blockchain system in a manner of separating indexes and data, when data governance is performed on the state data, causing the key-values of the state variables stored in some old data files to be migrated to new data files, thereby changing the location information of these key-value pairs in the persistent storage medium, it is not necessary to traverse and update the index entries in the index file that could originally be used to support querying these key-value pairs in the persistent storage medium, and these key-value pairs can be queried correctly and efficiently. Brief Description of the Drawings

[0009] To more clearly illustrate the technical solutions of the embodiments of this specification, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments recorded in this specification. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0010] Figure 1 It is an architecture diagram of a blockchain system exemplarily provided in the embodiments of this specification;

[0011] Figure 2 It is a schematic diagram of a tree structure exemplarily provided in the embodiments of this specification;

[0012] Figure 3 It is a schematic diagram of the structure of a data file provided in the embodiments of this specification;

[0013] Figure 4 It is a flowchart of a method for managing state data in a blockchain system provided in the embodiments of this specification;

[0014] Figure 5 It is a schematic diagram of the structure of a blockchain node provided in the embodiments of this specification. Detailed Embodiments

[0015] To enable those skilled in the art to better understand the technical solutions in this specification, the following will clearly and completely describe the technical solutions in the embodiments of this specification with reference to the drawings. Obviously, the described embodiments are only some embodiments of this specification, rather than all embodiments. Based on the embodiments in this specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this specification.

[0016] Figure 1 It is an architecture diagram of a blockchain system exemplarily provided in the embodiments of this specification. The blockchain system can include N blockchain nodes, whereFigure 1 Exemplarily shown are 8 blockchain nodes such as Node 1 - Node 8. The connections between the nodes schematically represent the connections between the nodes, and the aforementioned connections are used to support data transmission between different nodes.

[0017] The blockchain system can provide the function of smart contracts. The smart contracts in the blockchain system are contracts that can be triggered and executed by transactions. Smart contracts can be defined in the form of contract code. Invoking a smart contract in the blockchain system is to initiate a transaction pointing to the contract address of the smart contract, so that each node in the blockchain system runs the corresponding contract code distributively.

[0018] In various blockchain systems with the introduction of smart contracts, accounts can generally be divided into two types:

[0019] Contract account (CA): mainly used to store the contract code of the corresponding smart contract and the values of the state variables defined in the smart contract, and usually can only be activated by being called by an external account;

[0020] Externally owned account (EOA): an account registered by an external user in the blockchain system.

[0021] The design of external accounts and contract accounts is actually a mapping from account addresses to account states. The account state of any account usually includes fields such as nonce, balance, storageRoot, and codeHash. Among them, nonce and balance exist in both external accounts and contract accounts, and the codeHash and storageRoot attributes are generally only valid for contract accounts.

[0022] More specifically, for an external account, the value of nonce represents the number of transactions sent from the relevant account address; for a contract account, the value of nonce can represent the number of smart contracts created by the relevant account address. The value of balance represents the number of a certain digital resource / token owned by the relevant account address. The value of storage root is the hash value of the root node of a tree structure, such as an MPT tree, which is used to organize / manage the storage of the state variables of the relevant contract account. The value of codeHash represents the hash value of the contract code of the relevant smart contract. For an external account, since it does not include a smart contract, the values of the storageRoot and CodeHash fields can generally be an empty string / all - 0 string.

[0023] It should be noted that MPT, short for Merkle Patricia Tree, is a tree structure that combines Merkle Tree and Patricia Tree (a more space-saving trie). Among them, the Merkle tree algorithm can calculate a Hash value for multiple transactions respectively, and then connect them in pairs and calculate the Hash again until the top-level Merkle root. In some blockchain systems, an improved MPT tree is usually adopted, such as a 16-ary tree structure, which is usually simply referred to as the MPT tree.

[0024] The system data that needs to be persistently stored in the blockchain system can be divided into two parts: block data and state data.

[0025] The block data includes one or more blocks in ascending order of block height (or block number). A single block can include a block header and a block body. The block header can include the block hash previous_Hash (or parent hash) of the previous block, timestamp Timestamp, block number BlockNum, state root hash State_Root, transaction root hash Transaction_Root, receipt root hash Receipt_Root, and nonce, etc. The block body can include a transaction set and a receipt set.

[0026] A transaction in the blockchain system refers to a task unit that is executed and recorded in the blockchain system. A single transaction usually includes a sending field (From), a receiving field (To), and a data field (Data). The From field includes the account that initiates the transaction (i.e., the sender account), and the To field may include another account involved / pointed to by the transaction.

[0027] For any Nth block, based on the state data with block number (or version number) N - 1, multiple transactions included in the transaction set belonging to the Nth block can be executed in sequence to obtain the execution results of the multiple transactions. Then, based on the execution results of the multiple transactions, the state data with version number N - 1 is updated to obtain the state data with version number k.

[0028] In a blockchain system, state data can be managed through a tree structure, and different versions of state data will correspond to different tree structures. The location information of the value of a state variable in the persistent storage medium is stored in a leaf node of the tree structure, and the key of the state variable is stored in the directed path from the root node to the leaf node of the tree structure; the aforementioned state variable can be the account address of a contract account / external account, or can be a state variable in a smart contract. This tree structure can include, for example, MPT (Merkle Patricia Tree) or SMT (Sparse Merkle Tree), etc.

[0029] The tree structure for managing state data can include a state trie, and the hash value of the root node of the state trie is stored in State_Root in the block header. The location information of the account state of an external account / contract account in the persistent storage medium is stored in a leaf node of the state trie. In the directed path from the root node of the state trie to a leaf node, the account address of an external account / contract account, or part or all of the hash value calculated based on the account address is stored. As mentioned above, the account state of a single account usually can include fields such as Nonce, Balance, Storageroot, CodeHash, etc. Nonce and Balance exist in both external accounts and contract accounts, and CodeHash and Storage root are generally only valid for contract accounts.

[0030] The tree structure for managing state data can also include a storage trie. The hash value of the root node of the storage trie is stored in the storageroot field of the contract account corresponding to the relevant smart contract, so as to lock the contract state of the smart contract to the relevant contract account through the hash value. Similarly, the location information of the value of a state variable defined in the smart contract in the persistent storage medium is stored in a leaf node of the storage trie. In the directed path from the root node of the storage trie to a leaf node, the state key of a state variable defined in the relevant smart contract is stored. That is, part of the information in the directed path from the root node to the leaf node of the storage trie can be arranged in sequence to form the key of a state variable defined in the relevant smart contract, and the location information of the value of this state variable in the persistent storage medium is stored in this leaf node.

[0031] Based on the foregoing tree structure, in a blockchain system, it is possible to separate data and indexes, which helps improve the flexibility and performance of the system. The basic concepts involved here include data files and index files. Among them, the data file stores the variable values of state variables (including the state variables in the account address or smart contract); the index file stores the index entries pointing to the data file. The implementation principle is as follows: Write the data content (the value of the state variable or the key-value pair of the state variable) into the data file, and record the position information of the data content in the data file (such as file identifier, address offset, data length, etc.); furthermore, create an index entry in the index file, and the index entry contains the retrieval key related to the data content and the foregoing position information. In this way, when querying the data content, after obtaining the retrieval key of the data content, an index entry containing the retrieval key can be found in the index file, and the position information of the data content can be obtained from this index entry. Furthermore, the actual data content can be read from the corresponding data file according to the position information.

[0032] Exemplarily, referring to Figure 2As shown. For the tree structure corresponding to the state data with version number N, in the upper-level MPT, that is, in the state trie, for the leaf node A1, in the root node A8 (Extension Node), through the a7 of the shared nibble - slot 1 of the intermediate node A7 (Branch Node) - the 1335 of the key-end in the leaf node A1, they can be combined in sequence to form the key of a certain state variable: a711335. The value of this state variable is "Nonce = n1, Balance = 45.0ETH". The data file for storing "Nonce = n1, Balance = 45.0ETH" is file 5. The file identifier of file 5 is "5". The address offset of "Nonce = n1, Balance = 45.0ETH" in file 5 is 600. The data length of "Nonce = n1, Balance = 45.0ETH" is 100. Then, in the leaf node A1, for example, through the locatione field, the location information (5, 600, 100) of the value "Nonce = n1, Balance = 45.0ETH" in the persistent storage medium can be stored. Similar to the principle of the leaf node A1, in the leaf node A2, through the location field, the location information (3, 805, 150) of the value "Nonce = n2, Balance = 1.00WEI" of the state variable with key a77d337 in the persistent storage medium can be stored; in the leaf node A3, through the location field, the location information (2, 100, 180) of the value "Nonce = n3, Balance = 1.1ETH" of the state variable with key a7f9365 in the persistent storage medium can be stored; in the leaf node A4, through the location field, the location information P1 of the value "Nonce = n4, Balance = 0.12ETH, CodeHash = c1, Storage root = s1" of the state variable with key a77d397 in the persistent storage medium can be stored. s1 can be the hash value H(A10) of the tree node A10, that is, the hash value of the root node A10 of the next-level tree. It should be noted that Figure 1 To illustrate the relationship between the next-level MPT and the upper-level MPT, in the leaf node A4, the variable value "Nonce = n4, Balance = 0.12ETH, CodeHash = c1, Storage root = s1" is exemplified. In fact, in the leaf node A4, the location information P1 should be stored through the locaotion field. However Figure 2It is not shown in []. Among them, the leaf nodes A1, A2, and A3 correspond to external accounts, and the leaf node A4 corresponds to a contract account. For the contract account, it contains the next-level MPT, forming a Storage Trie, which is used to store the status of the state variables in the smart contract corresponding to the contract account.

[0033] As Figure 2 shown in the example of [], in the next-level MPT (i.e., the storage trie), for the leaf node A11, through the slot 3 in the root node A10 (Branch Node) - the 35b2e4 of key-end in the leaf node A11, they are sequentially combined to form the key of a certain state variable: 335b2e4. The value of this state variable is "Zhang San_A = 20". "Zhang San_A = 20" means, for example, that the share of digital assets of type A defined in the contract belonging to Zhang San is 20, that is, the balance of Zhang San's type A assets is 20. Among them, the data file for storing the value "Zhang San_A = 20" is file 2, the file identifier of file 2 is "2", the address offset of "Zhang San_A = 20" in file 2 is 7500, the data length of the value "Zhang San_A = 20" is 210, and in the leaf node A11, for example, through the locatione field, the location information (2, 750, 210) of "Zhang San_A = 20" in the persistent storage medium can be stored. Similar to the principle of the leaf node A11, in the leaf node A12, through the location field, the location information (3, 350, 210) of the value "Li Si_B = 50" of the state variable with the key 7c25988 in the persistent storage medium can be stored. "Li Si_B = 50" means, for example, that the share of digital assets of type B defined in the contract belonging to Li Si is 50, that is, the balance of Li Si's type B assets is 50; in the leaf node A15, through the location field, the location information (5, 760, 140) of the value "storedData = s" of the state variable with the key fa6be33 in the persistent storage medium can be stored; in the leaf node A16, through the location field, the location information (5, 170, 210) of the value "Wang Wu_A = 35" of the state variable with the key fa99365 in the persistent storage medium can be stored.

[0034] In the node composition of the aforementioned MPT, a prefix is used to represent the type of tree node. For example, 0 represents an Extension Node containing an even number of shared nibbles, 1 represents an Extension Node containing an odd number of shared nibbles, 2 represents a Leaf Node containing an even number of nibbles, and 3 represents a Leaf Node containing an odd number of nibbles.

[0035] In the above node composition, the hash value of the overall content of the next tree node is filled into the corresponding position of the previous tree node.

[0036] Based on the tree structure of the above example, the key-value pairs of the tree nodes in the tree structure can be obtained. The key of the tree node can be the result obtained by performing a hash operation on the overall content of the tree node (i.e., the value of the tree node). In this way, the key-value pairs of the tree nodes can be used as index entries and stored in the corresponding index file. For example, for the example in the previous text Figure 2 of Figure 2 the example, the key-value pairs of the tree nodes of the tree structure in Table 1 below can be stored as index entries.

[0037]

[0038]

[0039] Table 1

[0040] In Table 1 above, H() represents the hash calculation. In this way, the hash value of the next tree node is anchored in the previous tree node. Through such layer-by-layer hashing, the root hash of the entire state trie is obtained, and this root hash is locked into the state root field of the block header. Assume that the k-v pairs in Table 1 are saved on disk as index entries in the index file and the LSM structure is adopted. In this way, after querying the root node of the state trie and matching the key of the state variable to be searched with the shared nibble(s) field (for Extension Node) or slot (for Branch Node) of the root node from the beginning, the hash of the next layer of the tree node can be read from the matching position, and then the next InternalNode pointed to by this hash value can be found; after unlocking the Internal Node, continue to match the remaining part of the key of the state variable to be read from the front to the back. If there is a match, after reading the hash value from the matching position, jump to the next-level tree node pointed to by this hash value. This process is repeated continuously, unlocking the Internal Node level by level and matching the remaining part of the key of the state variable to be read from the front to the back. The matched hash value is used as the basis for the next search for the intermediate node or leaf node until the Leafnode is matched, so as to read the location information of the value of the state variable in the persistent storage medium from the Leaf node. Exemplarily, for example, finally, the content in the value of leaf node A11, that is, "prefix:2,Key-end:35b2e4,location:(2,750,210)", can be loaded into the memory, so as to obtain the location information "2,750,210" of the value "Zhang San_A = 20" of the state variable with key 335b2e4 in the persistent storage medium. Then, according to the location information "2,750,210", the data content between the 750KB and 960KB in file 2 can be read, and finally the value "Zhang San_A = 20" of the state variable with key 335b2e4 can be obtained.

[0041] When the key-value pairs of a tree node are directly used as index entries, the key of the tree node is used as the retrieval key. Alternatively, the key-value pairs of the tree node may not be directly used as index entries. For example, for any tree node, the key of the state variable stored in the directed path from the root node to this tree node or a component of the key of the state variable, that is, the key of the lexicographical content distribution from the root node through intermediate nodes to this tree node (hereinafter referred to as node ID or nodeID), is combined with the version number / block number of the state data to form a retrieval key. The position information of the value of the state variable included in the value of this tree node in the persistent storage medium is combined with this retrieval key to obtain an index entry corresponding to this tree node.

[0042] With a large increase in the version of state data, it will cause a significant increase in the amount of state data stored in the persistent storage medium. It is necessary to continuously prune / data govern the state data stored in the data storage system to delete the state data of earlier versions. During this process, some or all of the data content (the value of the state variable or the key-value pair) in some data files will be migrated from the deleted old data files to the new data files, resulting in a change in the position information of this part of the data content in the persistent storage medium. For example, the position information of a certain key-value pair corresponding to a certain state variable in the persistent storage medium will change from the old position information Location1 to the new position information Location2.

[0043] In some embodiments, when the position information of a certain data content changes, for example, when a certain key-value in an old data file is migrated to a new data file, it is necessary to traverse the index entries that could originally be used to access this key-value pair in the index file, and in the index entries that could originally be used to access this key-value pair, update the old position information of this key-value pair in the persistent storage medium, such as Location1, to the new position information, such as Location2.

[0044] In the foregoing embodiments, it takes a lot of time and computing resources to traverse and update all the index entries corresponding to all the data content whose position information has changed in the index file, which will have a negative impact on the performance of the blockchain system.

[0045] An embodiment of this specification provides a method for managing state data in a blockchain system, a blockchain node, a computing device, and a computer-readable storage medium. The blockchain system stores multiple data files through a persistent storage medium. Multiple key-value pairs in the data files are stored in an orderly manner according to the block number and the key. The meta-information of the data files includes a first-level index block and multiple second-level index blocks. The first-level index block stores multiple first index entries arranged in an orderly manner according to the block number and the key. Multiple second-level index blocks store their corresponding multiple second index entries in the storage order of the multiple key-value pairs. The first index entry is used to access a third index entry, and the third index entry corresponds to a different block number from the key-value pair accessed by the previous second index entry or is the first index entry in the second-level index block. First, the location information of a certain target key-value pair to be queried can be queried from a certain index entry in the index file, including the first block number of the first block and the first file identifier of the first data file. The target key-value pair is generated by executing a transaction sequence belonging to the first block. If the first data file does not exist in the persistent storage medium, the second data file for storing the target key-value pair at the current moment can be determined according to the first file identifier and the file change information. Then, according to the first block number, the target key in the target key-value pair, and the first-level index block corresponding to the second data file, the second target index block to which the second target index entry for accessing the target key-value pair belongs is determined from the multiple second-level index blocks corresponding to the second data file. Finally, the target key-value pair is read from the second data file according to the index entry in the second target index block.

[0046] In this way, on the basis of storing the state data of the blockchain system in a way of separating the index and the data, when data governance is performed on the state data, which causes the key-value of the state variables stored in some old data files to be migrated to new data files, resulting in a change in the location information of this part of the key-value pairs in the persistent storage medium, it is not necessary to traverse and update the index entries in the index file that could originally be used to support querying this part of the key-value pairs in the persistent storage medium, and the correct and efficient query of this part of the key-value pairs can be supported.

[0047] Figure 4 It is a schematic diagram of the structure of a data file provided in an embodiment of this specification.

[0048] For Figure 4The data file Data file DF2 provided exemplarily in [document]. Multiple key-value pairs in Data file DF2 can usually be stored in order according to the block number. Exemplarily, for any two key-value pairs kv1 and kv2 stored in Data file DF2, kv1 is generated due to the execution of the transaction sequence included in block N1, and kv2 is generated due to the execution of the transaction sequence included in block N2, where N1 is less than N2, then the storage position of kv1 in Data file DF2 should be before kv2.

[0049] For Figure 4 The data file Data file DF2 provided exemplarily in [document]. Multiple key-value pairs in Data file DF2 can usually be stored in order according to the key, for example, in alphabetical order, or can also be stored in order according to other custom sorting rules, such as in order according to the tree identifier combining alphabetical order and tree structure. Exemplarily, taking storage in alphabetical order as an example, for any two key-value pairs kv1 and kv3 stored in Data file DF2, both kv1 and kv3 are generated due to the execution of the transaction sequence included in block N1. If the first character of kv1 is before the first character of kv3 in alphabetical order, the storage position of kv1 in Data file DF2 should be before kv3; if the first P characters of kv1 are the same as the first P characters of kv3, then the order of kv1 and kv3 in Data file DF2 is determined according to the order relationship between the (P + 1)-th characters of kv1 and kv3.

[0050] For Figure 4 The data file Data file DF2 provided exemplarily in [document]. The metadata of Data file DF2 can include a first-level index block and multiple second-level index blocks arranged in order, such as including second-level index block 0, second-level index block 1, and second-level index block 2; the first-level index block corresponding to Data file DF2 and multiple second-level index blocks can be directly stored in Data file DF2, or can also be independently stored in a certain metadata file outside Data file DF2. Usually, for a single second-level index block among multiple second-level index blocks, the size of the second-level index block can be the same as the amount of data read from the persistent storage medium to the memory at one time, that is, the size of a single index block can be the same as the PageSize of the memory page, improving the access efficiency to the second-level index block.

[0051] In the second-level index blocks 0, 1, and 2, multiple second-index entries for accessing the multiple key-value pairs are stored in the storage order of the multiple key-value pairs stored in the Data file DF2. The second-index entry includes a retrieval key (different from the retrieval key of the index file) and an address offset. The retrieval key can be a key prefix with a predetermined length extracted from the key included in the corresponding key-value pair, or it can be determined in other ways. For example, the retrieval key in the second-index entry can be the first H bits of the hash value calculated for the key included in the corresponding key-value pair.

[0052] In the first-level index block corresponding to the Data file DF2, multiple first-index entries can be stored in an orderly manner according to the block number and the key. As described above, the first-index entry is used to access the third-index entry in multiple second-level index blocks. Among them, the block number corresponding to the key-value pair accessed by the third-index entry is different from that of the previous second-index entry of the third-index entry, or the third-index entry is the first index entry in the second-level index block. Refer to Figure 4 As shown, the first second-index entries in the second-level index blocks 0, 1, and 2 are all used as the third-index entry; in addition, the block number corresponding to the key-value pair accessed by the first second-index entry in the second-level index block 0 is 9, and the block number corresponding to the key-value pair accessed by the next second-index entry is 10, then the next second-index entry will also be used as the third-index entry.

[0053] Refer to Figure 4 As shown, the first-index entry includes a block number, a retrieval key, the block identifier of the second-level index block to which the third-index entry corresponding to the first-index entry belongs, and the address offset of the third-index entry. It can be understood that the retrieval key in the first-index entry should be the same as the retrieval key in the third-index entry accessed by the first-index entry. For example, refer to Figure 4 the second first-index entry in the exemplary first-level index block. This first-index entry is used to access the second index entry in the second-level index block 0. The first-index entry and the second-index entry include the same retrieval key (key prefix), such as "1b41f64e". It can be understood that the block number in the first-index entry can be determined based on the key-value pair accessed by the relevant third-index entry. For example, if the key-value pair that the second index entry in the second-level index block 0 can access is generated by executing the transaction sequence in block 9, then the block number in the first-index entry can also be block number 9.

[0054] For Figure 4The exemplary data file DF2 provided herein may be a new data file generated by data governance of status data, or it may not be a new data file generated by data governance of status data. In short, regardless of whether Data file DF2 is a new data file generated by data governance of status data, the meta-information of Data file DF2 may include the first-level index block and multiple second-level index blocks of the foregoing examples.

[0055] By setting the first-level index block and second-level index block of the foregoing examples for the data file, in the process of data governance of the status data of the blockchain system, more specifically, in the process of garbage collection of the data file already stored in the persistent storage medium, when the garbage ratio of a certain first data file reaches a preset threshold, that is, when a certain first data file is determined to be a data file to be deleted, valid key-value pairs can be selected from the first data file and added to the newly added second data file, and based on the first-level index block and multiple second-level index blocks corresponding to the first data file, the first-level index block and multiple second-level index blocks corresponding to the second data file are generated, and a change record is newly added to the file change information, and this change record is used to record the mapping relationship between the file identifier of the first data file and the file identifier of the second data file.

[0056] Referring to the data structures of the first-level index block and second-level index block described exemplary above, the first-level index block and multiple second-level index blocks corresponding to the first data file can reflect the block numbers corresponding to the valid key-values selected from the first data file. Therefore, when the key-values selected from the first data file are added to the second data file, based on the first-level index block and multiple second-level index blocks corresponding to the first data file, it can be ensured that the multiple key-value pairs stored in the second data file are also stored in an orderly manner according to the block number and key, and at the same time, the first-level index block and multiple second-level index blocks corresponding to the second data file are constructed.

[0057] Exemplarily, if kv1 stored in data file Data file DF1 is migrated as a valid key-value pair from data file Data file DF1 to data file Data file DF2, a change record can be generated in the file change information. This update record includes, for example, the mapping relationship between file identifiers DF1 and DF2. If in a subsequent process, kv1 stored in data file Dta file DF2 is migrated as a valid key-value pair from data file Datafile DF2 to data file Data file DF3, a change record can be generated in the file change information. This update record includes, for example, the mapping relationship between file identifiers DF2 and DF3. Thus, through the change records based on file identifiers recorded in the file change information, the file identifier of the data file storing kv1 at the current moment can be determined.

[0058] Combined with the foregoing Figure 3 , an exemplary method for managing state data in a blockchain system is described.

[0059] Figure 4 The figure is a flowchart of a method for managing state data in a blockchain system provided in an embodiment of this specification. The blockchain system stores multiple data files through a persistent storage medium. Multiple key-value pairs in the data files are stored in order according to the block number and the key. The meta-information of the data file includes a first-level index block and multiple second-level index blocks. The first-level index block stores multiple first index entries arranged in order according to the block number and the key. Multiple second-level index blocks store their corresponding multiple second index entries in the storage order of multiple key-value pairs. The first index entry is used to access the third index entry. The third index entry corresponds to a different block number from the key-value pair accessed by the previous second index entry or is the first index entry in the second-level index block.

[0060] This method can be executed by a blockchain node in the blockchain system, for example, by the storage engine of the blockchain node.

[0061] Referring to Figure 4 shown, this method may include, but is not limited to, some or all of the following steps S401 to step S407.

[0062] Step S401, obtain the location information of the target key-value pair to be queried, including the first block number of the first block and the first file identifier of the first data file. The target key-value pair is generated by executing a transaction sequence belonging to the first block.

[0063] As described above, the location information in the index entries included in the index file usually may include the file identifier of the data file to which the data content belongs, the address offset of the data content in the data file to which it belongs, and the data length. In the embodiments of this specification, the block number corresponding to the data content may also be added to the location information of the data content, and the data content is generated by executing the transaction sequence in the block indicated by the block number. For example, if the target key-value pair is generated by a blockchain node in a blockchain system by executing the transaction sequence included in the block with block number 11, after the blockchain node writes the target key-value pair into a data file such as Data file DF1, it updates the target tree node corresponding to the target key in the target key-value pair in the tree structure. More specifically, when writing the location information of the target key-value pair in the persistent storage medium, such as Locatione1, into the target tree node, the location information Locatione1 should include the block number 11, the file identifier DF1 of Data file DF1, the address offset of the target key-value pair in Data file DF1 (denoted as a1), and the data length (denoted as S1); when writing the target index entry used to access the target key-value pair in the persistent storage medium into the index file in the persistent storage medium, the location information Location1 should be included, and in addition, the retrieval key corresponding to the target key-value pair may also be included. The retrieval key is, for example, the key of the target tree node corresponding to the target key-value pair (such as the hash value calculated for all the data content of the target tree node), or the retrieval key may include the block number 11 and the key in the target key-value pair (i.e., the target key).

[0064] When it is necessary to query the target key-value pair in the persistent storage medium, it is first necessary to query the location information of the target key-value pair, such as Location1, from the index file according to the retrieval key corresponding to the target key-value pair. In this way, the block number 11 in Location1 is the first block number, and the file identifier DF1 in Location1 is the first file identifier.

[0065] Step S403, if the first data file does not exist, determine the second data file used to store the target key-value pair at the current moment according to the first file identifier and the file change information.

[0066] If there is a first data file indicated by the first file identifier in the persistent storage medium, for example, there is a data file Data file DF1 indicated by the file identifier DF1 in the persistent storage medium, directly read the target key-value pair from Data file DF1 according to the location information, such as the address offset a1 and data length s1 in Location1.

[0067] If there is no first data file indicated by the first file in the persistent storage medium, it means that the target key-value pair may have been migrated from the first data file originally used to store the target key-value pair to other data files due to data governance on the state data of the blockchain system. At this time, the second data file used to store the target key-value pair at the current moment can be determined according to the first file identifier, such as DF1, and the file change information.

[0068] Continuing with the previous example, assuming that the file change information records the mapping relationship between file identifiers DF1 and DF2, it means that the target key-value pair has been migrated from the data file indicated by file identifier DF1 to the data file indicated by file identifier DF2. Therefore, the data file indicated by file identifier DF2 can be determined as the second data file used to store the target key-value pair at the current moment. Assuming that the file change information records the mapping relationship between file identifiers DF1 and DF2, and in addition, records the mapping relationship between file identifiers DF2 and DF3, it means that the target key-value pair has been migrated from the data file indicated by file identifier DF1 to the data file indicated by file identifier DF2, and has been migrated from the data file indicated by file identifier DF2 to the data file indicated by file identifier DF3. Therefore, the data file indicated by file identifier DF3 can be determined as the second data file used to store the target key-value pair at the current moment.

[0069] Step S405: Determine the second target index block to which the second target index entry for accessing the target key-value pair belongs from the multiple secondary index blocks corresponding to the second data file according to the first block number, the target key in the target key-value pair, and the primary index block corresponding to the second data file.

[0070] In a possible implementation, the first target index entry may be first determined from the first-level index block corresponding to the second data file, the key arrangement position corresponding to the first index entry is not after the target key, and the next first index entry of the first target index entry does not include the first block number or the corresponding key arrangement position is after the target key; then, the second-level index block to which the third target index entry corresponding to the first target index entry belongs is determined as the second target index block.

[0071] Refer to Figure 3 As shown, the second index entry in the second-level index block corresponding to Data file DF2 includes the retrieval key of the key-value pair corresponding to the second index entry and the address offset of the key-value pair in Data file DF2. The retrieval key is, for example, a key prefix with a predetermined length extracted from the key included in the corresponding key-value pair. In addition, the first index entry in the first-level index block corresponding to Data file DF2 includes the block number, the retrieval key, the block identifier of the third index entry accessed in the second-level index block to which it belongs, and the address offset of the third index entry in the second-level index block to which it belongs. Here, it is assumed that the target key prefix of the target key included in the target key-value pair to be queried is 2c8b4ff3. Then, according to the first block number (i.e., block number 11) and the target key with 2c8b4ff3 as the key prefix, through binary search or sequential traversal, the "11, 0a3df2bc, 1, 0" serving as the first target index entry can be queried from the first-level index block corresponding to Data file DF2. Then, based on the block identifier 1 included in "11, 0a3df2bc, 1, 0", the second-level index block is determined as the second target index block.

[0072] Step S407, read the target key-value pair from the second data file according to the index entry in the second target index block.

[0073] In a possible implementation manner, the target key-value can be obtained from the second data file according to each index entry in the second target index block that is after the third target index entry. In a more specific example, each candidate index entry in the second target index block that is after the third target index entry can be traversed in sequence, and the retrieval key in the candidate index entry belongs to a component of the target key; the key-value pair is read from the second data file according to the candidate index entry, and when the key included in the read key-value pair is the same as the target key, the read key-value pair is determined as the target key-value pair. Wherein, when the corresponding data length of the key-value pair in the data file is not included in the second index entry, the data length to be read can be determined according to the address offset included in the next second index entry of the candidate index entry, and the key-value pair is read from the second data file according to the address offset and the data length included in the candidate index entry; when the corresponding data length of the key-value pair in the data file is included in the second index entry, the key-value pair can be directly read according to the address offset and the data length included in the candidate index entry.

[0074] The foregoing implementation manner is exemplary. In the case where the requirement for data access efficiency is low and the conflict probability is small, all the second index entries in the second target index block can also be traversed in the arranged order, and finally the target key-value pair including the target key can be read from the second data file.

[0075] The following is a specific analysis of the technical solution provided in the embodiments of the present specification for efficiently querying key-value pairs.

[0076] The impact on the query efficiency is mainly reflected in the caching of the first-level index block and the second-level index block.

[0077] Here, it is assumed that the total number of key-value pairs involved in the multi-version state data in the blockchain system is A. In the theoretical limit scenario, the location information of each key-value pair changes due to data governance. The total memory occupancy of the secondary index block is A * (kBytes * vBytes). Here, k refers to the data length of the key prefix included in the second index entry, and v refers to the data length of the location information included in the second index entry. The location information in the second index entry can include only the address offset or both the address offset and the data length of the corresponding key-value pair. If the size of the secondary index block is taken as 4KB, the key prefix is 4B, and the addressing range is 2^32, a single secondary index block can contain 4KB / 8B = 512 second index entries. The storage space required for the secondary index block in memory is relatively small and can usually be fully cached in memory. Here, it is further assumed that in a single first index entry in the primary index block, it includes the block number of kBytes, the key prefix of kBytes, the block identifier of iBytes, and the address offset of oBytes. The values of k, i, and o are relatively small. For each secondary index block, only a very small number of first index entries for this secondary index block need to be stored in the primary index block. The storage space required for the primary index block in memory is relatively small and can be fully cached in memory.

[0078] Based on the same concept as the foregoing method embodiments, an embodiment of this specification also provides a blockchain node 500 in a blockchain system, a method for managing state data in the blockchain system. The blockchain system stores multiple data files through a persistent storage medium. The multiple key-value pairs in the data files are stored in an orderly manner according to the block number and the key. The meta-information of the data files includes a primary index block and multiple secondary index blocks. The primary index block stores multiple first index entries arranged in an orderly manner according to the block number and the key. The multiple second index entries corresponding to the multiple key-value pairs are stored in the multiple secondary index blocks in the storage order of the multiple key-value pairs. The first index entry is used to access a third index entry, and the third index entry corresponds to a different block number from the key-value pair accessed by the previous second index entry or is the first index entry in the secondary index block. Refer to Figure 5, the blockchain node 500 includes: a location acquisition unit 501 configured to acquire the location information of a target key-value pair to be queried, where the location information includes the first block number of a first block and the first file identifier of a first data file, and the target key-value pair is generated by executing a transaction sequence belonging to the first block; a file determination unit 503 configured to, if the first data file does not exist, determine a second data file for storing the target key-value pair at the current moment according to the first file identifier and file change information; an entry retrieval unit 505 configured to determine, according to the first block number, the target key in the target key-value pair, and a first-level index block corresponding to the second data file, a second target index block to which a second target index entry for accessing the target key-value pair belongs from a plurality of second-level index blocks corresponding to the second data file; and a query processing unit 507 configured to read the target key-value pair from the second data file according to the index entry in the second target index block.

[0079] An embodiment of the present specification also provides a computer-readable storage medium, on which a computer program / instructions are stored. When the computer program / instructions are executed on a computer, the computer is made to execute a method for managing state data in a blockchain system provided in each of the foregoing embodiments.

[0080] An embodiment of the present specification also provides a computing device, including a memory and a processor. A computer program / instructions are stored in the memory, and when the processor executes the computer program / instructions, a method for managing state data in a blockchain system provided in each of the foregoing embodiments is implemented.

[0081] In the 1990s, it was obvious to distinguish whether an improvement in a technology was an improvement in hardware (e.g., improvement in circuit structures such as diodes, transistors, switches, etc.) or an improvement in software (improvement in method flows). However, with the development of technology, many improvements in method flows today can be regarded as direct improvements in hardware circuit structures. Almost all designers obtain the corresponding hardware circuit structures by programming the improved method flows into the hardware circuits. Therefore, it cannot be said that an improvement in a method flow cannot be implemented with a hardware entity module. For example, a Programmable Logic Device (PLD) (e.g., a Field Programmable Gate Array (FPGA)) is such an integrated circuit whose logical function is determined by the user's programming of the device. Designers can program themselves to "integrate" a digital system on a single PLD, without having to ask a chip manufacturer to design and fabricate a dedicated integrated circuit chip. Moreover, nowadays, instead of manually fabricating integrated circuit chips, this programming is mostly implemented using "logic compiler" software, which is similar to the software compilers used in program development and writing. The original code before compilation also has to be written in a specific programming language, which is called a Hardware Description Language (HDL), and there is not only one kind of HDL, but many kinds, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. The most commonly used ones currently are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also be aware that by simply performing a little logical programming on the method flow with the above-mentioned several hardware description languages and programming it into an integrated circuit, it is easy to obtain the hardware circuit that implements the logical method flow.

[0082] The controller can be implemented in any suitable manner. For example, the controller can take the form of, for example, a microprocessor or a processor and a computer-readable medium storing computer-readable program code (such as software or firmware) executable by the (micro)processor, logic gates, switches, an application specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller. Examples of the controller include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art also know that in addition to implementing the controller in the form of pure computer-readable program code, it is entirely possible to logically program the method steps to enable the controller to be implemented in the form of logic gates, switches, application specific integrated circuits, programmable logic controllers, embedded microcontrollers, etc. to achieve the same function. Therefore, such a controller can be considered a hardware component, and the devices included therein for implementing various functions can also be regarded as the structures within the hardware component. Or even, the devices for implementing various functions can be regarded as either software modules for implementing the method or the structures within the hardware component.

[0083] The systems, devices, modules, or units illustrated in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a server system. Of course, this application does not exclude that with the development of future computer technologies, the computers for implementing the functions of the above embodiments can be, for example, personal computers, laptop computers, in-vehicle human-machine interaction devices, cellular phones, camera phones, smart phones, personal digital assistants, media players, navigation devices, email devices, game consoles, tablet computers, wearable devices, or any combination of these devices.

[0084] Although one or more embodiments of this specification provide method operation steps as described in the embodiments or flowcharts, more or fewer operation steps may be included based on conventional or non-creative means. The order of steps listed in the embodiments is only one way among many execution orders of steps and does not represent the only execution order. When actually executed by a device or terminal product, it may be executed in the order of the method shown in the embodiments or the drawings or executed in parallel (for example, in an environment of parallel processors or multi-threaded processing, or even in a distributed data processing environment). The term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, product or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, product or device. Without further limitation, there is no exclusion of additional identical or equivalent elements in the process, method, product or device comprising the said elements. For example, if terms such as first and second are used to denote names, they do not denote any particular order.

[0085] For the convenience of description, the above device is described by dividing it into various modules according to functions. Of course, when implementing one or more of this specification, the functions of each module can be implemented in the same or multiple software and / or hardware, or the modules implementing the same function can be realized by a combination of multiple sub-modules or sub-units, etc. The device embodiments described above are only illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the couplings or direct couplings or communication connections shown or discussed with each other can be through some interfaces, and the indirect couplings or communication connections of the devices or units can be in electrical, mechanical or other forms.

[0086] The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be realized by computer program instructions. These computer program instructions can be provided to the processor of a general computer, a special computer, an embedded processor or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate a device for realizing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0087] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction device that implements the functions specified in one or more processes and / or boxes Figure 1 in one or more processes and / or boxes Figure 1 specified in one or more boxes.

[0088] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operational steps are performed on the computer or other programmable device to produce a computer-implemented process, thereby providing steps for implementing the functions specified in one or more processes and / or boxes Figure 1 in one or more processes and / or boxes Figure 1 specified in one or more boxes.

[0089] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.

[0090] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM) and / or non-volatile memory such as read-only memory (ROM) or flash memory (flash RAM). Memory is an example of computer-readable media.

[0091] Computer-readable media includes both permanent and non-permanent, removable and non-removable media implemented by any method or technology for storing information. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile discs (DVD) or other optical storage, magnetic cassettes, magnetic tape disk storage, graphene storage or other magnetic storage devices, or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory media such as modulated data signals and carrier waves.

[0092] Those skilled in the art should understand that one or more embodiments of this specification can be provided as a method, a system, or a computer program product. Therefore, one or more embodiments of this specification can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, one or more embodiments of this specification can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0093] One or more embodiments of this specification can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. One or more embodiments of this specification can also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.

[0094] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for system embodiments, since they are basically similar to method embodiments, the description is relatively simple. For related parts, reference can be made to the partial description of method embodiments. In the description of this specification, the description of reference terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of this specification. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine different embodiments or examples described in this specification and the features of different embodiments or examples.

[0095] The above description is only for the embodiments of one or more embodiments of this specification and is not intended to limit one or more embodiments of this specification. For those skilled in the art, one or more embodiments of this specification can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of this specification shall be included within the scope of the claims.

Claims

1. A method for managing state data in a blockchain system, where the blockchain system stores multiple data files through a persistent storage medium, multiple key-value pairs in the data files are stored in an orderly manner according to block numbers and keys, the meta-information of the data files includes a first-level index block and multiple second-level index blocks, the first-level index block stores multiple first index entries arranged in an orderly manner according to block numbers and keys, the multiple second-level index blocks store their corresponding multiple second index entries in the storage order of the multiple key-value pairs, the first index entry is used to access a third index entry, and the third index entry corresponds to a different block number from the key-value pair accessed by the previous second index entry or is the first index entry in the second-level index block. The method includes: Obtaining the location information of a target key-value pair to be queried, including the first block number of the first block and the first file identifier of the first data file, where the target key-value pair is generated by executing a transaction sequence belonging to the first block; If the first data file does not exist, determining, according to the first file identifier and file change information, the second data file used to store the target key-value pair at the current moment; According to the first block number, the target key in the target key-value pair, and the first-level index block corresponding to the second data file, determining, from the multiple second-level index blocks corresponding to the second data file, the second target index block to which the second target index entry for accessing the target key-value pair belongs; Reading the target key-value pair from the second data file according to the index entry in the second target index block.

2. The method according to claim 1, where the second index entry includes a retrieval key and an address offset, and the retrieval key is a key prefix with a predetermined length extracted from the key included in the corresponding key-value pair.

3. The method according to claim 2, where the first index entry includes a block number, a retrieval key, the block identifier of the second-level index block to which the third index entry corresponding to the first index entry belongs, and the address offset of the third index entry.

4. The method according to claim 1, where the step of determining, from the multiple second-level index blocks corresponding to the second data file, the second target index block to which the second target index entry for accessing the target key-value pair belongs according to the first block number, the target key in the target key-value pair, and the first-level index block corresponding to the second data file includes: Determining a first target index entry from the first-level index block corresponding to the second data file, where the key arrangement position corresponding to the first target index entry is not after the target key, and the next first index entry after the first target index entry does not include the first block number or the corresponding key arrangement position is after the target key; Determine the second target index block as the secondary index block to which the third target index entry corresponding to the first target index entry belongs.

5. The method according to claim 4, wherein reading the target key-value pair from the second data file according to the index entry in the second target index block includes: Obtain the target key-value from the second data file according to each index entry in the second target index block that is after the third target index entry.

6. The method according to claim 5, wherein the obtaining the key-value from the second data file according to each index entry in the second target index block that is after the third target index entry includes: Traverse in sequence each candidate index entry in the second target index block that is after the third target index entry, and the retrieval key in the candidate index entry is a component of the target key; Read the key-value pair from the second data file according to the candidate index entry. When the key included in the read key-value pair is the same as the target key, determine the read key-value pair as the target key-value pair.

7. The method according to claim 6, wherein reading the key-value pair from the second data file according to the candidate index entry includes: Determine the data length to be read according to the address offset included in the next second index entry of the candidate index entry, and read the key-value pair from the second data file according to the address offset and the data length included in the candidate index entry.

8. The method according to any one of claims 1-7, the method further includes: During the process of garbage collection of the data file, select valid key-value pairs from the first data file and add them to the second data file, and generate the first index block and multiple secondary index blocks corresponding to the second data file according to the first index block and multiple secondary index blocks corresponding to the first data file; Add a change record to the file change information, and the change record is used to record the mapping relationship between the file identifier of the first data file and the file identifier of the second data file.

9. A blockchain node in a blockchain system, a method for managing state data in the blockchain system, the blockchain system stores multiple data files through a persistent storage medium, multiple key-value pairs in the data files are stored in an orderly manner according to the block number and the key, the meta-information of the data file includes a first index block and multiple secondary index blocks, the first index block stores multiple first index entries arranged in an orderly manner according to the block number and the key, multiple second index entries corresponding to the multiple key-value pairs are stored in the multiple secondary index blocks in the storage order of the multiple key-value pairs, the first index entry is used to access a third index entry, the third index entry corresponds to a different block number from the key-value pair accessed by the previous second index entry or is the first index entry in the secondary index block, the blockchain node includes: A location acquisition unit, configured to acquire location information of a target key-value pair to be queried, where the location information includes a first block number of a first block and a first file identifier of a first data file, and the target key-value pair is generated by executing a transaction sequence belonging to the first block; A file determination unit, configured to, if the first data file does not exist, determine a second data file for storing the target key-value pair at the current moment according to the first file identifier and file change information; An entry retrieval unit, configured to determine a second target index block to which a second target index entry for accessing the target key-value pair belongs from a plurality of second index blocks corresponding to the second data file according to the first block number, a target key in the target key-value pair, and a first-level index block corresponding to the second data file; A query processing unit, configured to read the target key-value pair from the second data file according to the index entry in the second target index block.

10. A computer-readable storage medium, on which a computer program is stored, and when the computer program is executed in a computing device, the computing device executes the method according to any one of claims 1-8.