Index reconstruction method and device for state data in block chain system
By combining dredged index files and target index files to generate new index files, the problem of inefficient index reconstruction in the existing technology is solved, and the index reconstruction efficiency and resource utilization of the blockchain system are improved.
Patent Information
- Application Number
- CN202510349534.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-07-11
AI Technical Summary
When prior art is used to re-index in blockchain systems, old index entries may be located in multiple levels of index files, resulting in inefficient rewriting processes and waste of resources.
By combining the dredged index files with the target index files whose write time is behind the barrier point, a new index file is generated and written in the persistent storage medium, ensuring that the new index file is higher than the remaining index files whose creation time is before the barrier point, thereby efficiently accessing the value of the state variable.
It realizes that while ensuring the normal use of the index file, there is no need to rewrite multiple levels of index entries, which improves the efficiency and resource utilization of index reconstruction.
Smart Images

Figure CN120295973A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of this specification belong to the field of computer technology, and in particular, relate to a method and device for index reconstruction of state data in a blockchain system. Background Art
[0002] A blockchain system is a new application mode of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanism, and encryption algorithms. In a blockchain system, data blocks are combined into a chain data structure in a sequential connection manner according to the time sequence, and it is a distributed ledger that is guaranteed to be tamper-proof and unforgeable by cryptographic means. Due to the characteristics of decentralization, information immutability, and autonomy of the blockchain system, the blockchain system has received more and more attention and applications. Summary of the Invention
[0003] The purpose of the present invention is to provide a method, device, computing device, and computer-readable storage medium for index reconstruction of state data in a blockchain system, which can more efficiently complete the index reconstruction of the state data of the blockchain system.
[0004] In a first aspect, a method for index reconstruction of state data in a blockchain system is provided. The blockchain system stores data files and index files through a persistent storage medium. The data files are used to store the values of state variables, and the index files are used to store index entries. The index entries include the position information of the values of state variables in the persistent storage medium. The method includes: performing a merging process on a dredged index file and a target index file whose writing time is after a barrier point to obtain a new index file. The dredged index file is obtained by performing the current round of garbage collection on the state data of the blockchain system. The barrier point is the index file that was last written from the memory to the persistent storage medium when starting to perform the current round of garbage collection; writing the new index file into the persistent storage medium.
[0005] Second aspect, there is provided an apparatus for index reconstruction of state data in a blockchain system. The blockchain system stores data files and index files through a persistent storage medium. The data files are used to store the values of state variables, and the index files are used to store index entries. The index entries include the location information of the values of state variables in the persistent storage medium. The apparatus includes: a merging processing unit configured to perform a merging process on a dredged index file and a target index file whose writing time is after a barrier point to obtain a new index file. The dredged index file is obtained by performing the current round of garbage collection on the state data of the blockchain system. The barrier point is the index file that is latest written from the memory to the persistent storage medium when starting to perform the current round of garbage collection; a storage processing unit configured to write the new index file into the persistent storage medium.
[0006] Third aspect, there is provided a computing device including a memory and a processor. A computer program is stored in the memory. When the processor executes the computer program, the method described in the first aspect is implemented.
[0007] Fourth aspect, there is provided a computer-readable storage medium with a computer program stored thereon. When the computer program is executed in a computing device, the computing device executes the method described in the first aspect.
[0008] In the technical solution provided in the embodiments of this specification, the level where the new data file is located may be higher than the other index files whose generation time is before the barrier point. When an external party needs to access the value of a certain state variable, usually, it can start from the index file at the highest level and gradually search for index entries that can be used to access the value of the state variable in index files at lower levels. Even if the value of the state variable is rewritten to a certain new data file in the persistent storage medium due to the execution of garbage collection, resulting in the invalidation of the original old index entries that were used to access the value of the state variable in the index files before the new index file, new index entries that can be used to access the value of the state variable can still be queried from the new index file, so as to correctly access the value of the state variable using the new index entries; in other words, during the process of index reconstruction based on the dredged index file, it can be achieved that without the need to rewrite a large number of index entries in the index files at multiple levels while ensuring that each index file can be normally used to correctly access the value of the state variable, thereby enabling more efficient implementation of index reconstruction. Description of the Drawings
[0009] To more clearly illustrate the technical solutions of the embodiments of this specification, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments recorded in this specification. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0010] Figure 1 It is an architecture diagram of a blockchain system exemplarily provided in the embodiments of this specification;
[0011] Figure 2 It is a schematic diagram of a tree structure for managing state data exemplarily provided;
[0012] Figure 3 It is a schematic diagram of generating an incremental page and a base page for a logical page including multiple tree nodes exemplarily provided;
[0013] Figure 4 It is a schematic diagram of the technical scenario of the technical solutions provided in the embodiments of this specification;
[0014] Figure 5 It is a schematic diagram of the relationship between a data file and an index file exemplarily provided;
[0015] Figure 6 It is a flowchart of a method for index reconstruction of state data in a blockchain system provided in the embodiments of this specification;
[0016] Figure 7 It is a schematic diagram of merging a dredging index file and other index files to obtain a new index file exemplarily provided;
[0017] Figure 8 It is a schematic diagram of a device for index reconstruction of state data in a blockchain system provided in the embodiments of this specification. Detailed implementation manners
[0018] To enable those skilled in the art to better understand the technical solutions in this specification, the following will clearly and completely describe the technical solutions in the embodiments of this specification in conjunction with the drawings. Obviously, the described embodiments are only some embodiments of this specification, rather than all embodiments. Based on the embodiments in this specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of this specification.
[0019] Figure 1 It is an architecture diagram of a blockchain system exemplarily provided in the embodiments of this specification. The blockchain system may include N blockchain nodes, where Figure 1Exemplarily shown are eight blockchain nodes such as Node 1 - Node 8. The connections between the nodes schematically represent the connections between the nodes, and the foregoing connections are used to support data transmission between different nodes.
[0020] The blockchain system can provide the function of smart contracts. The smart contracts in the blockchain system are contracts that can be triggered and executed by transactions. The smart contracts can be defined in the form of contract codes. Invoking a smart contract in the blockchain system is to initiate a transaction pointing to the contract address of the smart contract, so that each node in the blockchain system runs the corresponding contract code distributively.
[0021] In various blockchain systems with smart contracts introduced, accounts can generally be divided into two types:
[0022] Contract account (CA): mainly used to store the contract code of the corresponding smart contract and the values of the state variables defined in the smart contract, and usually can only be activated by being called by an external account;
[0023] Externally owned account (EOA): an account registered by an external user in the blockchain system.
[0024] The design of external accounts and contract accounts is actually a mapping from account addresses to account states. The account state of any account usually includes fields such as nonce, balance, storageRoot, and codeHash. Among them, nonce and balance exist in both external accounts and contract accounts, and the codeHash and storageRoot attributes are generally only valid for contract accounts.
[0025] More specifically, for an external account, the value of nonce represents the number of transactions sent from the relevant account address; for a contract account, the value of nonce can represent the number of smart contracts created by the relevant account address. The value of balance represents the number of a certain digital resource / token owned by the relevant account address. The value of storage root is the hash value of the root node of a tree structure, such as an MPT tree, and this MPT tree is used to organize / manage the storage of the state variables of the relevant contract account. The value of codeHash represents the hash value of the contract code of the relevant smart contract. For an external account, since it does not include a smart contract, the values of the storageRoot and CodeHash fields can generally be an empty string / all-zero string.
[0026] It should be noted that MPT, whose full name is Merkle Patricia Tree, is a tree structure that combines Merkle Tree and Patricia Tree (a more space-saving trie). Among them, the Merkle tree algorithm can calculate a Hash value for multiple transactions respectively, and then connect them in pairs and calculate the Hash again until the top-level Merkle root. In some blockchain systems, an improved MPT tree is usually adopted, such as a 16-ary tree structure, which is usually simply referred to as the MPT tree.
[0027] The system data that needs to be persistently stored in the blockchain system can be divided into two parts: block data and state data.
[0028] The block data includes one or more blocks in increasing order of block height (or block number). A single block can include a block header and a block body. The block header can include the block hash previous_Hash (or parent hash) of the previous block, timestamp Timestamp, block number BlockNum, state root hash State_Root, transaction root hash Transaction_Root, receipt root hash Receipt_Root, and nonce, etc. The block body can include a transaction set and a receipt set.
[0029] A transaction in the blockchain system refers to a task unit that is executed and recorded in the blockchain system. A single transaction usually includes a sending field (From), a receiving field (To), and a data field (Data). The From field includes the account that initiates the transaction (i.e., the sender account), and the To field may include another account involved / pointed to by the transaction.
[0030] For any Nth block, multiple transactions included in the transaction set belonging to the Nth block can be executed in sequence according to the state data with the block number (or version number) of N - 1, and the execution results of the multiple transactions can be obtained. Then, the state data with the version number of N - 1 can be updated according to the execution results of the multiple transactions to obtain the state data with the version number of k.
[0031] In a blockchain system, state data can be managed through a tree structure, and different versions of state data will correspond to different tree structures. The location information of the value of a state variable in a persistent storage medium is stored in a leaf node of the tree structure, and the key of the state variable is stored in the directed path from the root node to the leaf node of the tree structure; the aforementioned state variable can be the account address of a contract account / external account, or can be a state variable in a smart contract. The tree structure can include, for example, MPT (Merkle Patricia Tree) or SMT (Sparse Merkle Tree), etc.
[0032] The tree structure for managing state data can include a state trie, and the hash value of the root node of the state trie is stored in State_Root in the block header. The location information of the account state of an external account / contract account in a persistent storage medium is stored in a leaf node of the state trie. In the directed path from the root node to a leaf node of the state trie, the account address of an external account / contract account, or part or all of the hash value calculated based on the account address is stored. As mentioned above, the account state of a single account usually can include fields such as Nonce, Balance, Storageroot, CodeHash, etc. Nonce and Balance exist in both external accounts and contract accounts, and CodeHash and Storage root are generally only valid for contract accounts.
[0033] The tree structure for managing state data can also include a storage trie. The hash value of the root node of the storage trie is stored in the storageroot field of the contract account corresponding to the relevant smart contract, so as to lock the contract state of the smart contract to the relevant contract account through the hash value. Similarly, the location information of the value of a state variable defined in the smart contract in a persistent storage medium is stored in a leaf node of the storage trie. In the directed path from the root node to a leaf node of the storage trie, the state key of a state variable defined in the relevant smart contract is stored. That is, part of the information in the directed path from the root node to the leaf node of the storage trie can be arranged in sequence to form the key of a state variable defined in the relevant smart contract, and the location information of the value of the state variable in the persistent storage medium is stored in this leaf node.
[0034] Based on the foregoing tree structure, in the blockchain system, it is possible to separate data and indexes, which helps to improve the flexibility and performance of the system. The basic concepts involved here include data files and index files. Among them, the data file stores the values of state variables (including the state variables in the account address or smart contract); the index file stores index entries pointing to the data file. The implementation principle is as follows: write the data content (the variable value of the state variable or the key-value pair of the state variable) into the data file, and record the location information of the data content in the data file (such as file identifier, address offset, data length, etc.); then create an index entry in the index file, and the index entry contains the retrieval key related to the data content and the foregoing location information. In this way, when querying the data content, after obtaining the retrieval key of the data content, an index entry containing the retrieval key can be found in the index file, and then the location information of the data content can be obtained from the index entry, and then the actual data content can be read from the corresponding data file according to the location information.
[0035] Exemplarily, with reference to Figure 2As shown. For the tree structure corresponding to the state data with version number N, in the upper-level MPT, that is, in the state trie, for the leaf node A1, in the root node A8 (Extension Node), through the a7 of the shared nibble - the slot 1 of the intermediate node A7 (Branch Node) - the 1335 of the key - end in the leaf node A1, they can be sequentially combined to form the key of a certain state variable: a711335. The value of this state variable is "Nonce = n1, Balance = 45.0ETH". The data file for storing "Nonce = n1, Balance = 45.0ETH" is file 5, the file identifier of file 5 is "5", the address offset of "Nonce = n1, Balance = 45.0ETH" in file 5 is 600, and the data length of "Nonce = n1, Balance = 45.0ETH" is 100. Then, in the leaf node A1, for example, through the locatione field, the location information (5, 600, 100) of the value "Nonce = n1, Balance = 45.0ETH" in the persistent storage medium can be stored. Similar to the principle of the leaf node A1, in the leaf node A2, through the location field, the location information (3, 805, 150) of the value "Nonce = n2, Balance = 1.00WEI" of the state variable with key a77d337 in the persistent storage medium can be stored; in the leaf node A3, through the location field, the location information (2, 100, 180) of the value "Nonce = n3, Balance = 1.1ETH" of the state variable with key a7f9365 in the persistent storage medium can be stored; in the leaf node A4, through the location field, the location information P1 of the value "Nonce = n4, Balance = 0.12ETH, CodeHash = c1, Storage root = s1" of the state variable with key a77d397 in the persistent storage medium can be stored. s1 can be the hash value H(A10) of the tree node A10, that is, the hash value of the root node A10 of the next-level tree. It should be noted that Figure 1 To illustrate the relationship between the next-level MPT and the upper-level MPT, in the leaf node A4, the variable value "Nonce = n4, Balance = 0.12ETH, CodeHash = c1, Storage root = s1" is exemplified. In fact, in the leaf node A4, the location information P1 should be stored through the locaotion field. However, Figure 2It is not shown in []. Among them, the leaf nodes A1, A2, and A3 correspond to external accounts, and the leaf node A4 corresponds to a contract account. For the contract account, it contains the next-level MPT, forming a Storage Trie, which is used to store the states of the state variables in the smart contract corresponding to the contract account.
[0036] As Figure 2 shown in the example of [], in the next-level MPT (i.e., storage trie), for the leaf node A11, through the slot 3 in the root node A10 (Branch Node) - the 35b2e4 of key-end in the leaf node A11, they are sequentially combined to form the key of a certain state variable: 335b2e4. The value of this state variable is "Zhang San_A = 20". "Zhang San_A = 20" means, for example, that the share of type A digital assets defined in the contract belonging to Zhang San is 20, that is, the balance of Zhang San's type A assets is 20. Among them, the data file for storing the value "Zhang San_A = 20" is file 2, the file identifier of file 2 is "2", the address offset of "Zhang San_A = 20" in file 2 is 7500, the data length of the value "Zhang San_A = 20" is 210. In the leaf node A11, for example, through the locatione field, the location information (2, 750, 210) of "Zhang San_A = 20" in the persistent storage medium can be stored. Similar to the principle of the leaf node A11, in the leaf node A12, through the location field, the location information (3, 350, 210) of the value "Li Si_B = 50" of the state variable with key 7c25988 in the persistent storage medium can be stored. "Li Si_B = 50" means, for example, that the share of type B digital assets defined in the contract belonging to Li Si is 50, that is, the balance of Li Si's type B assets is 50; in the leaf node A15, through the location field, the location information (5, 760, 140) of the value "storedData = s" of the state variable with key fa6be33 in the persistent storage medium can be stored; in the leaf node A16, through the location field, the location information (5, 170, 210) of the value "Wang Wu_A = 35" of the state variable with key fa99365 in the persistent storage medium can be stored.
[0037] In the node composition of the aforementioned MPT, a prefix is used to represent the type of tree node. For example, 0 represents an Extension Node containing an even number of shared nibbles, 1 represents an Extension Node containing an odd number of shared nibbles, 2 represents a Leaf Node containing an even number of nibbles, and 3 represents a Leaf Node containing an odd number of nibbles.
[0038] In the above node composition, the hash value of the overall content of the next tree node is filled into the corresponding position of the previous tree node.
[0039] Based on the tree structure of the above example, the key-value pairs of the tree nodes in the tree structure can be obtained. The key of the tree node can be the result obtained by performing a hash operation on the overall content of the tree node (i.e., the value of the tree node). In this way, the key-value pairs of the tree nodes can be used as index entries and stored in the corresponding index file. For example, for the example Figure 2 in the previous text, Figure 2 the key-value pairs of the tree nodes of the example of the tree structure in the index file can be stored as index entries as shown in Table 1 below.
[0040]
[0041]
[0042] Table 1
[0043] In Table 1 above, H() represents the hash calculation. In this way, the hash value of the next tree node is anchored in the previous tree node. Through such layer-by-layer hashing, the root hash of the entire state trie is obtained, and this root hash is locked into the state root field of the block header. Assume that the k-v pairs in Table 1 are saved on disk as index entries in the index file and the LSM structure is adopted. In this way, after querying the root node of the state trie and matching the key of the state variable to be searched with the shared nibble(s) field (for Extension Node) or slot (for Branch Node) of the root node from the beginning, the hash of the next layer of the tree node can be read from the matching position, and then the next InternalNode pointed to by this hash value can be found; after unravelling the Internal Node, continue to match the remaining part of the key of the state variable to be read from front to back in it. If there is a match, after reading the hash value from the matching position, jump to the next-level tree node pointed to by this hash value. Repeat this process continuously, unravelling the Internal Node level by level and matching the remaining part of the key of the state variable to be read from front to back. The matched hash value is used as the basis for the next search for the intermediate node or leaf node until the Leafnode is matched, so as to read the location information of the value of the state variable in the persistent storage medium from the Leaf node. Exemplarily, for example, finally, the content in the value of the leaf node A11, that is, "prefix:2,Key-end:35b2e4,location:(2,750,210)", can be loaded into the memory, so as to obtain the location information "2,750,210" of the value "Zhang San_A = 20" of the state variable with the key 335b2e4 in the persistent storage medium. Then, according to the location information "2,750,210", the data content between the 750KB and 960KB in File 2 can be read, and finally the value "Zhang San_A = 20" of the state variable with the key 335b2e4 can be obtained.
[0044] It should be noted that for the key of the tree node in the storage trie, such as in the foregoing example Figure 2The key of leaf node A16 in Table 1 can be obtained by combining the hash value calculated for leaf node A16 with a specific prefix. Only the hash value H(A16) of node A16 is shown in Table 2, and the specific prefix is not shown. The aforementioned specific prefix can be determined based on the contract address of the smart contract corresponding to the storage trie. For example, it can be the contract address of the smart contract corresponding to the storage trie or the hash value of the contract address. In this article, it is essentially only described directly for convenience as the key of a tree node can be the hash value of the content of the tree node, but it does not mean that the keys of all tree nodes are the hash values of the content of the tree nodes.
[0045] When directly using the key - value pair of a tree node as an index entry, the key of the tree node is used as the retrieval key. Or it is also possible not to directly use the key - value pair of the tree node as an index entry. For example, for any tree node, the key of the state variable stored in the directed path from the root node to this tree node or the component of the key of the state variable, that is, the key of the lexicographical content distribution from the root node through intermediate nodes until this tree node (hereinafter referred to as node ID or nodeID), can be combined with the version number / block number of the state data to form a retrieval key. The position information of the value of the state variable included in the value of this tree node in the persistent storage medium is combined with this retrieval key to obtain an index entry corresponding to this tree node.
[0046] Based on the tree - shaped structure of the foregoing example, the entire tree - shaped structure, that is, the Merkle trie, can also be divided into multiple logical pages (Logical Page). Specifically, according to the node association relationship of the tree - shaped structure, several adjacent tree nodes above and below can be aggregated into one LogicalPage. For example, in a 2 - level 16 - ary trie with a total of 256 child nodes, they can be aggregated into one LogicalPage. Among them, there are multiple layers of LogicalPages from the root node to the leaf node of the tree - shaped structure, and there are parent - child and sibling relationships between LogicalPages, thus forming a complete tree - shaped structure. In this way, a logical page contains at least one tree node, different logical pages contain different tree nodes, and the tree nodes corresponding to the currently latest version of the state data can be maintained in the logical page.
[0047] Such as Figure 3As shown, within each LogicalPage, a MemoryPage can be maintained to represent the content of all tree nodes corresponding to the current latest version of the logical page. Further, the content of the tree nodes maintained by the logical page can be updated according to consecutive versions, and a BasePage and a DeltaPage can be generated based on the generation / modification behavior. Specifically, the BasePage and the DeltaPage are generated based on the generation / modification behavior. For example, for consecutive state changes, such as in version N, the state variables are a = 5, b = 8, c = 3, and these values a = 5, b = 8, c = 3 can be the content in the BasePage. For version N + 1 where a = 6, and in version N + 2 where a = 8, assuming b and c remain unchanged, the values a = 6 in version N + 1 and a = 8 in version N + 2 can be used as the DeltaPage, such as DeltaPage1, on the basis of this BasePage, and the cases where b and c remain unchanged are not included in DeltaPage1. For version N + 3 where a = 11, and in version N + 4 where a = 15, the values a = 11 in version N + 3 and a = 15 in version N + 4 can be used as the DeltaPage, such as DeltaPage2, on the basis of this BasePage. Again, assuming b and c remain unchanged, b and c are not included in DeltaPage2, and so on. It should be particularly noted that the BasePage and the DeltaPage can be the smallest units of the memory structure in a tree structure. In addition, the BasePage can be the smallest unit of disk persistence in a tree structure. On this basis, the DeltaPage can also be the smallest unit of disk persistence in a tree structure.
[0048] For the BasePage, a new BasePage can be generated every time the state is modified a predetermined number of times (such as m times). In the previous example, in version N, the state variable a = 5, and this value a = 5 can be the content in the BasePage1. For version N + 1 where a = 6, and in version N + 2 where a = 8, the values a = 6 in version N + 1 and a = 8 in version N + 2 can be used as the DeltaPage1 on the basis of this BasePage. For version N + 3 where a = 11, and in version N + 4 where a = 15, the values a = 11 in version N + 3 and a = 15 in version N + 4 can be used as the DeltaPage2 on the basis of this BasePage. For version N + 5 where a = 20, assuming the modification reaches the predetermined number of times 5, the value a = 20, as well as b = 8 and c = 3, can be the content in the new BasePage2.
[0049] For DeltaPage, it describes several version modifications of the LogicalPage. These multiple modifications are aggregated into a set. For example, a DeltaPage is generated every M modifications, and a new DeltaPage is created to collect subsequent version modification operations. As mentioned above, for the value of a being 6 in version N + 1 and a being 8 in version N + 2, the values of a = 6 in version N + 1 and a = 8 in version N + 2 are used as the incremental page DeltaPage1 based on this BasePage. For a being 11 in version N + 3 and a being 15 in version N + 4, the values of a = 11 in version N + 3 and a = 15 in version N + 4 can be used as the incremental page DeltaPage2 based on this BasePage. It can be seen that a DeltaPage is generated every 2 modifications.
[0050] For each modification of the BasePage and DeltaPage, corresponding dirty data is generated in memory. Usually, when the memory occupied by the BasePage and DeltaPage reaches a predetermined share, the BasePage and DeltaPage are batch disk-persisted, thus avoiding persisting all dirty data for each modification and continuously occupying memory and CPU resources.
[0051] The previous text described DeltaPage and BasePage in terms of the changes in the values of the state variables in the LogicalPage. However, what is maintained in the logical page is the content of the tree nodes. According to the continuous version updates of the state data, DeltaPage and BasePage are generated according to the generation / modification behavior. What is described in the DeltaPage and BasePage are still the key-value pairs of the tree nodes associated with the version numbers.
[0052] Similar to the node ID, the logical page can have a page ID. The page ID can be the lexicographical content from the root node to the topmost tree node in this logical page (i.e., the memory page), that is, the page ID can be the key of the state variable or a component of the key of the state variable whose lexicographical content is distributed from the root node through the intermediate nodes to the topmost tree node in this LogicalPage.
[0053] DeltaPage and BasePage can also have versions. BasePage generally corresponds to a memory page and contains all the tree nodes in a memory page, that is, it contains global state variables. Therefore, it can have the same version as the corresponding memory page. DeltaPage can correspond to one or more memory pages and contains the tree nodes in one or more consecutive versions of the memory pages that have changed relative to the previous adjacent DeltaPage / BasePage, that is, it contains the changed state variables, which are generally not global state variables. Therefore, the version of the delta page can be the lowest or highest version among the one or more memory pages it corresponds to.
[0054] Exemplarily, a BasePage with version 4 contains state variables a = 1, b = 2, c = 3 and corresponding intermediate nodes and root nodes. In a DeltaPage generated after this BasePage, it can include a = 2 with version 5 and corresponding intermediate nodes and root nodes, a = 1 with version 6 and corresponding intermediate nodes and root nodes, b = 4 with version 7 and corresponding intermediate nodes and root nodes; in this case, although a = 1 with version 6 is the same as the state variable a = 1 contained in this BasePage, it is different from a = 2 with version 5 of the previous adjacent one. Therefore, a = 1 with version 6 and corresponding intermediate nodes and root nodes are also included in the DeltaPage. In addition, for the latter DeltaPage, that is, the DeltaPage after the DeltaPage corresponding to versions 5, 6, 7, for example, the DeltaPage containing tree nodes with versions 7, 8, and 9, it is actually relative to the immediately previous DeltaPage and the tree nodes of the versions. Thus, the tree nodes with versions 5, 6, 7 respectively correspond to the memory pages with versions 5, 6, 7. The DeltaPage containing the tree nodes with versions 5, 6, 7 can have its own version as 5, or it can also be 7, that is, it can be the lowest or highest version among the corresponding multiple memory pages.
[0055] Moreover, a page type can be set for DeltaPage and BasePage respectively to distinguish between DeltaPage and BasePage.
[0056] Combined with the determination of index entries based on tree nodes described above, for example, using the key-value pairs of tree nodes as index entries as described above. When persisting the aforementioned BasePage and DeltaPage pages, one or more BasePage / DeltaPage are stored as the smallest storage unit in the index file of the persistent storage medium. In other words, a single index file can include one or more BasePage / DeltaPage, and a single BasePage / DeltaPage can include one or more key-value pairs of tree nodes. Moreover, the page ID of the logical page corresponding to the BasePage / DeltaPage, together with the page type and version of the BasePage / DeltaPage, can form the page identifier of the BasePage / DeltaPage. In addition, based on the key-value pairs of tree nodes (i.e., index entries) in all BasePage / DeltaPage included in the index file, the data file usage information of the index file can be determined, and the data file usage information can be stored in the corresponding index file.
[0057] Generally, a Log-Structured Merge-Tree (LSM) can be adopted to manage the continuously appended index files in a blockchain system. In the persistent storage medium, the key-value pairs of tree nodes can be stored in Sorted String Table (SST) files at multiple levels. For example, BasePage / DeltaPage can be stored through multiple levels of SSTs. An SST file is an index file. Among them, BasePage / DeltaPage can be first stored in memory, such as first stored in the MemTable in memory. When the data volume in the MemTable reaches a certain threshold, such as 256MB, the BasePage / DeltaPage in the MemTable can be written (flushed) to the disk. To avoid the write operation of the MemTable from blocking the flush, this MemTable can be converted into an immutable Immutable Memtable, that is, the Immutable Memtable is set to read-only, and a new MemTable is generated to receive the newly incoming BasePage / DeltaPage, and then the Immutable MemTable is written to the SST at the highest level in the persistent storage medium. Generally, after the storage capacity of the upper level (such as Level 0) reaches or approaches the upper limit, a process called "compaction" can be used to merge the SST files in Level 0 into the next level, such as Level 1. In this way, the key-value pairs of tree nodes stored in the SST files at higher levels are more updated than those stored in the SST files at lower levels. When it is necessary to find the key-value pair of the tree node corresponding to a certain state variable, usually, the search can start from the SST file at the highest level and gradually search for the SST files at lower levels until the key-value pair of the tree node to be searched is found. It is not difficult to understand that if the logical pages are not divided for the tree structure, in each index file in the persistent storage medium, the key-value pairs of the tree nodes in BasePage / DeltaPage will be directly stored without taking BasePage / DeltaPage as the minimum storage unit.
[0058] With a large increase in the number of state data versions, the amount of state data stored in the persistent storage medium will increase significantly. It is necessary to continuously prune / data govern the state data stored in the data storage system to delete the state data of earlier versions. During this process, some index entries in the index file will be deleted according to the specified version number. For example, when the specified version number is 105, the BasePage / DeltaPage with a version number less than 105 will be deleted. After the deletion process of the index entries is completed, some or all of the data contents (the value of the state variable or key-value pair) in some data files become invalid; through the corresponding garbage collection process, some or all of the invalid data contents can be eliminated to save storage space.
[0059] The garbage ratio of multiple data files can be determined according to the data file usage information stored in each of the multiple index files; based on the garbage ratios of the multiple data files, the target data files to be recycled are determined from the multiple data files; valid data is selected from the target data files to generate a new data file, and a corresponding dredging index file is generated based on the valid data in the new data file.
[0060] For the new data file, it can be directly stored in the persistent storage medium. For the dredging index file, it is necessary to reconstruct the index entries corresponding to the valid data according to the dredging index file. Generally, it is necessary to traverse the index entries in the index file, find the old index entries originally used to access a certain valid data in the target data file, and rewrite them as new index entries that can be used to access the valid data in the new data file stored in the persistent storage medium in the dredging index file.
[0061] When using LSM to manage the continuously appended index files in the blockchain system, the old index entries may be located in the index files with lower levels, and it may be necessary to rewrite a large number of index files located in multiple levels, making it difficult to efficiently complete index reconstruction.
[0062] In view of this, an index reconstruction method, apparatus, computing device, and computer-readable storage medium for state data in a blockchain system are provided in an embodiment of this specification. For the dredged index file obtained by performing the current round of garbage collection, the dredged index file and the target index file whose writing time is after the barrier point can be merged to obtain a new index file. The foregoing barrier point is the index file that was last written from the memory to the persistent storage medium when the current round of garbage collection started; then the new index file is written to the persistent storage medium. In this way, the level of the new data file may be higher than that of the remaining index files whose generation time is before the barrier point. When an external party needs to access the value of a certain state variable, it can usually start from the index file at the highest level and gradually search for the index entry that can be used to access the value of the state variable in the index files at lower levels. Even if the value of the state variable is rewritten to a certain new data file in the persistent storage medium due to the execution of garbage collection, resulting in the invalidation of the original index entry for accessing the value of the state variable in the index file before the new index file in terms of generation time, a new index entry that can be used to access the value of the state variable can still be queried from the new index file, so as to correctly access the value of the state variable using the new index entry; in other words, during the process of index reconstruction based on the dredged index file, it can be realized that without rewriting the index entries in the index files at multiple levels, the index reconstruction can be more efficiently implemented while ensuring that each index file can be normally used to correctly access the value of the state variable.
[0063] Figure 4 This is a schematic diagram of the technical scenario of the technical solution provided in the embodiment of this specification. Refer to Figure 4As shown, multiple data files and multiple index files can be stored in the persistent storage medium; the multiple index files can be divided into multiple levels arranged from high to low, for example, divided into Level 0, Level 1, and Level 2 arranged in sequence. The main logic that can be used to access the index files and data files can be divided into two processes, such as the foreground management process and the background management process. For the foreground management process, it is mainly responsible for the access to the index files and data files required in the blockchain node due to the execution of transactions or in response to external data access requests; for the background management process, it is mainly responsible for pruning / data governance of the state data of the blockchain system. In addition, it can also be responsible for performing compaction operations such as Minor Comaction and Major Conmpaction on the index files. Among them, Minor Comaction can achieve the merging of multiple index files with relatively small data volumes belonging to the same level into one index file with a relatively large data volume belonging to the same level; Major Conmpaction can achieve the merging of multiple index files with relatively small data volumes belonging to the upper level into an index file with a relatively large data volume in the lower level.
[0064] With the large increase in the version of the state data, it will cause a substantial increase in the data volume of the state data stored in the persistent storage medium, and it is necessary to continuously perform pruning / data governance on the state data stored in the data storage system to delete the state data of earlier versions. During this process, some index entries in the index file will be deleted according to the specified version number. For example, when the specified version number is 105, the BasePage / DeltaPage with a version number less than 105 will be deleted. After the deletion process of the index entries is completed, some or all of the data content (the value of the state variable or the key-value pair) in some data files becomes invalid; through the corresponding garbage collection process, some or all of the invalid data content can be eliminated to save storage space.
[0065] Exemplarily, after the foreground management process writes the index file Index file IF6 to the persistent storage medium, the background management process may start data governance on the state data of the blockchain system. During this process, the background management process may delete some target index entries from each index entry whose write time is before Index file IF6. For example, by deleting the BasePage / DeltaPage with a version number less than 105 in the relevant index file, multiple target index entries can be deleted. After deleting these target index entries, it may cause the values of some state variables corresponding to the target index entries to become invalid.
[0066] Please refer to Figure 5As shown here, it is assumed that the data file Data file DF1 stores "the value V1 of the status variable M1 in the status data with version number 103", and the data file Data file DF2 stores "the value V2 of the status variable M2 in the status data with version number 104". There is an index entry T1 in the index files Index file IF1 to Index file IF6 that can be used to access "V1", and there is an index entry T2 that can be used to access "V2", and there are no other index entries that can be used to access V1 and V2. If the index entry T1 and the index entry T2 are deleted, the "V1" stored in the data file Data file DF1 becomes invalid, and the "V2" stored in the data file Data file DF2 becomes invalid. As invalid data, V1 and V2 will waste storage resources due to occupying additional storage space of the persistent storage medium.
[0067] After completing the deletion of the index entry from the index file, the background management process can then perform garbage collection on the data files in the persistent storage medium. Continuing with the previous example, it is further assumed here that when performing garbage collection on the data files Data file DF1 to Data file DF4, the amount of data occupied by the invalid data in the data files Data file DF1 and Data file DF2 is relatively large, that is, the garbage ratios of Data file DF1 and Data file DF2 are relatively high. Then Datafile DF1 and Data file DF2 may be deleted as target data files, and the valid data in Data file DF1 and Datafile DF2 is added to a new data file such as Data file DF5. At the same time, corresponding dredging index files will be generated for the valid data in Data fileDF5, and the dredging index files will include index entries that can be used to access the valid data in Datafile DF5.
[0068] Please refer to Figure 5As described above, it is assumed here that the valid data in Data file DF1 includes "the value V3 of the status variable M3 in the status data with version number 103", and the valid data in Data file DF2 includes "the value V4 of the status variable M4 in the status data with version number 104". In the status data of each version with a version number greater than 104, there may be multiple consecutive versions that do not update the values of the status variables M3 and M4. Therefore, the index files Index file IF1 to Index file IF5 may all include an index entry T3 for accessing V3 in the persistent storage medium, and an index entry T4 for accessing V4 in the persistent storage medium. It is further assumed here that the retrieval key included in the index entry T3 is H1(V3), and the location information is L1(V3), and it is assumed that the retrieval key included in the index entry T4 is H1(V4), and the location information is L1(V4); in addition, it is further assumed that after V3 and V4 are written into the data file Data file DF5 in the persistent storage medium, the location information of V3 changes to H2(V3), and the location information of V4 changes to H2(V4). Then, the dredged index file may include an index entry T5 that can be used to access V3 and an index entry T6 for accessing V4. The retrieval key included in the index entry T5 is H1(V3), and the location information is L2(V3). The retrieval key included in the index entry T6 is H1(V4), and the location information is L2(V4).
[0069] The index files Index file IF1 to Index file IF5 may all include an index entry T3 for accessing V3 in the persistent storage medium, and an index entry T4 for accessing V4 in the persistent storage medium. If the conventional index reconstruction method is followed, it is necessary to rewrite each index entry T3 included in the index files Index file IF1 to Index file IF5 as the index entry T5, and rewrite each index entry T4 included in the index files Index file IF1 to Index file IF5 as the index entry T6, in order to ensure that the foreground management process can find the rewritten index entry T5 and index entry T6 in the index file through H1(V3) and H1(V4) as the retrieval keys, and then correctly access V3 and V4 according to the location information in the index entries T5 and T6. This requires rewriting a large number of index entries in a large-scale index file, with extremely low efficiency and waste of resources.
[0070] Figure 6The figure is a flowchart of a method for index reconstruction of state data in a blockchain system provided in an embodiment of this specification. The blockchain system stores multiple data files and multiple index files through a persistent storage medium. The data files are used to store the values of state variables, and the index files are used to store index entries. An index entry includes a retrieval key and the location information of the value of the state variable in the persistent storage medium. As mentioned above, state data in the blockchain system can be managed through a tree structure. The location information of the value of a state variable in the persistent storage medium is stored in a leaf node of the tree structure, and the key of the state variable is stored in the directed path from the root node to the leaf node of the tree structure. An index entry can include the key-value pair of a tree node in the tree structure; that is, the key of the tree node can be used as the retrieval key, and the value of the tree node includes the location information of the value of the state variable in the persistent storage medium. This method exemplarily describes the process of index reconstruction of the state data of the blockchain system by using the dredged index file during the process of data governance of the state data of the blockchain system, and more specifically, after obtaining the dredged index file by performing garbage collection on the state data of the blockchain system.
[0071] Referring to Figure 6 as shown, this method may include, but is not limited to, some or all of the following steps S601 to S605.
[0072] Step S601: Merge the dredged index file and the target index file whose write time is after the barrier point to obtain a new index file. The dredged index file is obtained by performing this round of garbage collection on the state data of the blockchain system, and the barrier point is the index file that is newly written from the memory to the persistent storage medium when starting to perform this round of garbage collection.
[0073] Referring to Figure 4 as shown, assume that when starting to perform this round of garbage collection, the index file that the foreground management process newly writes from the memory to the persistent storage medium is Index file IF6, then Index file IF6 can be used as the barrier point. Here, continue to assume that the index file Index file IF7 is the next index file that the foreground management process writes from the memory to the persistent storage medium after Index file IF6, then the index file Index file IF7 can be used as the target index file.
[0074] In a possible implementation, index pairs existing between the dredging index file and the target index file can be determined. The index pairs include a first index entry belonging to the dredging index file and a second index entry belonging to the target index file. The first index entry and the second index entry include the same retrieval key. The first index entry in the index pair is added to the new index file, and each index entry that does not belong to the index pair in the dredging index file and the target index file is added to the new index file.
[0075] Exemplarily, please refer to Figure 7 As shown. In the latest generated status data, the value of the status variable M3 remains V3 and has not been modified. Therefore, the Index file IF7 serving as the target index file may include the index entry T3 that could originally be used to access V1 in the data file Data file DF1. As described above, the index entry T3 includes H1(V3) as the retrieval key and the old location information L1(V3) of V3. In addition, it may also include other index entries such as the index entry T7. According to the foregoing implementation, the dredging index file and the Index file IF7 are merged. The index entries T3 and T5 will be used as an index pair. The index entry T5 as the first index entry in the index pair, and the index entries T6 and T7 that do not belong to the index pair will all be added to a new index file such as the Index file IF8.
[0076] The foregoing implementation is exemplary, or the merging of the dredging index file and the target index file can be achieved in other ways. Exemplarily, the set of index entries deleted from the index file during this round of data governance can be recorded. For any index entry in the target index file, if the index entry belongs to the set of index entries, it is deleted; if it does not belong to the set of index entries, it is retained in the target index file. Finally, the union of the index entries generated in the target index file and each index entry in the dredging index file is taken to obtain all the index entries to be added to the new index file.
[0077] In some embodiments, the barrier point, the dredging index file, and the target index file whose write time is after the barrier point can also be merged to obtain a new index file. Exemplarily, the union of the index entries in the barrier point such as the Index file IF6 and the target index file such as the Index file IF7 can be taken first to obtain an intermediate file. Then, according to the implementation of merging the target index file and the dredging index file exemplified above, the intermediate file and the dredging index file are merged.
[0078] Next, in step S603, the new index file is written to the persistent storage medium.
[0079] For example, continue to refer to Figure 5 and Figure 7 As shown, after obtaining the new index file Index file IF8, the new index file Index file IF8 will be written into the persistent storage medium. In addition, the target index file, such as Index file IF7, will be deleted.
[0080] As mentioned above, LSM is usually used to manage multiple index files in a persistent storage medium. It is necessary to ensure that the foreground management process can query the index entries that can be used to correctly access the values of related state variables from multiple index files based on the corresponding retrieval key; at the same time, considering that the foreground management process will gradually query the index entries containing the retrieval key from the index files of higher levels to the index files of lower levels according to the retrieval key, and the embodiment of this specification will not change the old index entries corresponding to the values of state variables whose position information has changed in the index files whose generation time is before the barrier point, it is necessary to ensure that the index files whose generation time is after the barrier point cannot be merged into the same index file with the index files whose generation time is before the barrier point. After merging to obtain the new index file, the following step S605 can be performed.
[0081] Step S605 , setting a prohibition condition, where the prohibition condition is used to indicate that it is prohibited to merge the index file whose writing time is before the barrier point and the index file whose writing time is after the barrier point into the same index file by performing a compaction operation.
[0082] Exemplarily, by setting a prohibition condition. When the index file Index file IF6 is used as a barrier point, the index files whose generation time is after the barrier point, such as Index file IF8, will not be merged into the same index file by the background management process through Minor Comaction or Major Conmpaction. Obviously, the background management process can still perform Minor Comaction or Major Conmpaction on multiple index files whose generation time is before the barrier point, and can also perform Minor Comaction or Major Conmpaction on multiple index files whose generation time is after the barrier point.
[0083] In some embodiments, the aforementioned prohibition condition can also be used to indicate prohibiting compaction operations on barrier points. In this case, the index file of the barrier point, such as Index file IF6, can be stored in a dedicated file in the persistent storage medium that allows foreground management processes and background management processes to access, prohibiting any compaction operations on the barrier point.
[0084] Based on the same concept as the foregoing method embodiments, an index reconstruction apparatus 800 for state data in a blockchain system is further provided in the embodiments of this specification. The blockchain system stores data files and index files through a persistent storage medium. The data files are used to store the values of state variables, and the index files are used to store index entries. The index entries include the location information of the values of state variables in the persistent storage medium. Refer to Figure 8 As shown, the apparatus 800 includes: a merging processing unit 801 configured to perform a merging process on the dredging index file generated by executing the current round of garbage collection and the target index file whose writing time is after the barrier point to obtain a new index file, where the barrier point is the index file that was last written from the memory to the persistent storage medium when starting to execute the current round of garbage collection on the data files of the blockchain system; a storage processing unit 803 configured to write the new index file into the persistent storage medium.
[0085] In the embodiments of this specification, a computer-readable storage medium is further provided, on which computer programs / instructions are stored. When the computer programs / instructions are executed on a computer, the computer is made to execute an index reconstruction method for state data in a blockchain system described in each of the foregoing embodiments.
[0086] In the embodiments of this specification, a computing device is further provided, including a memory and a processor. Computer programs / instructions are stored in the memory. When the processor executes the computer programs / instructions, an index reconstruction method for state data in a blockchain system described in each of the foregoing embodiments is implemented.
[0087] In the 1990s, it was obvious to distinguish whether an improvement in a technology was a hardware improvement (e.g., improvement in circuit structures such as diodes, transistors, switches, etc.) or a software improvement (improvement in method processes). However, with the development of technology, many improvements in method processes today can be regarded as direct improvements in hardware circuit structures. Almost all designers obtain the corresponding hardware circuit structures by programming the improved method processes into the hardware circuits. Therefore, it cannot be said that an improvement in a method process cannot be implemented by a hardware entity module. For example, a Programmable Logic Device (PLD) (e.g., a Field Programmable Gate Array (FPGA)) is such an integrated circuit whose logical function is determined by the user's programming of the device. Designers can program by themselves to "integrate" a digital system on a piece of PLD without asking the chip manufacturer to design and fabricate a dedicated integrated circuit chip. Moreover, nowadays, instead of manually fabricating integrated circuit chips, this programming is mostly implemented using "logic compiler" software, which is similar to the software compiler used in program development and writing. The original code before compilation also has to be written in a specific programming language, which is called Hardware Description Language (HDL), and there is not only one kind of HDL, but many kinds, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. Currently, the most commonly used are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also be aware that as long as the method process is slightly logically programmed with the above-mentioned several hardware description languages and programmed into the integrated circuit, it is easy to obtain the hardware circuit that implements the logical method process.
[0088] The controller can be implemented in any suitable manner. For example, the controller can take the form of, for example, a microprocessor or a processor and a computer-readable medium storing computer-readable program code (such as software or firmware) executable by the (micro)processor, logic gates, switches, an application specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller. Examples of the controller include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art also know that in addition to implementing the controller in the form of pure computer-readable program code, it is entirely possible to logically program the method steps to enable the controller to be implemented in the form of logic gates, switches, application specific integrated circuits, programmable logic controllers, embedded microcontrollers, etc. to achieve the same function. Therefore, such a controller can be considered a hardware component, and the devices included therein for implementing various functions can also be regarded as the structures within the hardware component. Or even, the devices for implementing various functions can be regarded as either software modules for implementing the method or the structures within the hardware component.
[0089] The systems, devices, modules, or units illustrated in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a server system. Of course, this application does not exclude that with the development of future computer technologies, the computers for implementing the functions of the above embodiments can be, for example, personal computers, laptop computers, in-vehicle human-machine interaction devices, cellular phones, camera phones, smart phones, personal digital assistants, media players, navigation devices, email devices, game consoles, tablet computers, wearable devices, or any combination of these devices.
[0090] Although one or more embodiments of this specification provide method operation steps as described in the embodiments or flowcharts, more or fewer operation steps may be included based on conventional or non-creative means. The order of steps listed in the embodiments is only one way among many orders of step execution and does not represent the only execution order. When an actual device or terminal product is executed, it may be executed in the order of the method shown in the embodiments or the drawings or in parallel (for example, in a parallel processor or multi-threaded processing environment, or even in a distributed data processing environment). The term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, product or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, product or device. Without further limitation, there is no exclusion of additional identical or equivalent elements in the process, method, product or device comprising the said elements. For example, if terms such as first and second are used to denote names, they do not denote any particular order.
[0091] For convenience of description, when describing the above device, it is divided into various modules according to functions and described separately. Of course, when implementing one or more of this specification, the functions of each module may be implemented in the same or multiple software and / or hardware, or the modules implementing the same function may be implemented by a combination of multiple sub-modules or sub-units, etc. The device embodiments described above are only illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other may be through some interfaces, and the indirect coupling or communication connection of the device or unit may be in electrical, mechanical or other forms.
[0092] The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing device generate a device for implementing the specified functions in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0093] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to operate in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including instruction means that implement the function specified in one or more of the blocks and / or steps of the flowchart. Figure 1 one or more of the blocks and / or steps Figure 1 specified in one or more of the blocks and / or steps.
[0094] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, thereby providing steps for implementing the function specified in one or more of the blocks and / or steps of the flowchart. Figure 1 one or more of the blocks and / or steps Figure 1 specified in one or more of the blocks and / or steps.
[0095] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.
[0096] The memory may include non-permanent memory in the computer-readable medium, random access memory (RAM) and / or non-volatile memory such as read only memory (ROM) or flash memory (flash RAM). The memory is an example of a computer-readable medium.
[0097] Computer-readable media includes both permanent and non-permanent, removable and non-removable media implemented by any method or technology for storage of information such as computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read only memory (ROM), electrically erasable programmable read only memory (EEPROM), flash memory or other memory technologies, compact disc read only memory (CD-ROM), digital versatile discs (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage, graphene storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory media such as modulated data signals and carrier waves.
[0098] Those skilled in the art should understand that one or more embodiments of this specification can be provided as a method, a system, or a computer program product. Therefore, one or more embodiments of this specification can take the form of an all-hardware embodiment, an all-software embodiment, or an embodiment combining software and hardware aspects. Moreover, one or more embodiments of this specification can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) that contain computer-usable program code.
[0099] One or more embodiments of this specification can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. One or more embodiments of this specification can also be practiced in a distributed computing environment, where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.
[0100] The various embodiments in this specification are described in a progressive manner. For the same or similar parts among the various embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for system embodiments, since they are basically similar to method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiments. In the description of this specification, the description of reference terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of this specification. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0101] The above description is only for the embodiments of one or more embodiments of this specification and is not used to limit one or more embodiments of this specification. For those skilled in the art, one or more embodiments of this specification can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of this specification shall be included within the scope of the claims.
Claims
1. A method for index reconstruction of state data in a blockchain system, where the blockchain system stores data files and index files through a persistent storage medium. The data files are used to store the values of state variables, and the index files are used to store index entries. The index entries include the location information of the values of state variables in the persistent storage medium. The method includes: Performing a merging process on a dredged index file and a target index file whose writing time is after a barrier point to obtain a new index file. The dredged index file is obtained by performing the current round of garbage collection on the state data of the blockchain system. The barrier point is the index file that was most recently written from memory to the persistent storage medium when the current round of garbage collection started. Writing the new index file to the persistent storage medium.
2. The method according to claim 1, where the performing a merging process on a dredged index file and a target index file whose writing time is after a barrier point to obtain a new index file includes: Determining index pairs existing between the dredged index file and the target index file. The index pairs include a first index entry belonging to the dredged index file and a second index entry belonging to the target index file. The first index entry and the second index entry include the same retrieval key. Adding the first index entry in the index pair to the new index file, and adding each index entry that does not belong to the index pair in the dredged index file and the target index file to the new index file.
3. The method according to claim 2, where the state data in the blockchain system is managed through a tree structure. In a leaf node of the tree structure, the location information of the value of a state variable in the persistent storage medium is stored. In the directed path from the root node to the leaf node of the tree structure, the key of the state variable is stored. Where the index entry includes the key-value pair of the tree node in the tree structure, and the retrieval key is the tree node key.
4. According to the method described in claim 1, the index file is managed in the persistent storage medium through a Log-Structured Merge Tree (LSM); wherein, The method further includes: setting a prohibition condition, which is used to indicate that it is prohibited to merge an index file whose writing time is before the barrier point and an index file whose writing time is after the barrier point into the same index file by performing a compaction operation.
5. According to the method described in claim 4, the prohibition condition is further used to indicate that compaction operations are prohibited on the barrier points; the method further includes: Storing the barrier point in a dedicated file in the persistent storage medium.
6. The method according to claim 1, wherein the merging process for the dredging index file generated by performing the current round of garbage collection and the target index file whose writing time is after the barrier point includes: Performing a merging process on the barrier point, the dredged index file, and the target index file whose writing time is after the barrier point.
7. The method according to any one of claims 1-6, where the performing the current round of garbage collection on the state data of the blockchain system includes: Predicting the garbage ratio of each data file stored in the persistent storage medium according to the data file usage information stored in the index file whose writing time is before the barrier point. Determining the target data file to be recycled according to the garbage ratio of each data file. Selecting valid data from the target data file and generating a new data file using the valid data. Generate the dredging index file according to the new data file.
8. An index reconstruction device for state data in a blockchain system, where the blockchain system stores data files and index files through a persistent storage medium, the data files are used to store the values of state variables, the index files are used to store index entries, and the index entries include the location information of the values of state variables in the persistent storage medium. The device includes: A merging processing unit configured to perform a merging process on a dredging index file and a target index file whose writing time is after the barrier point to obtain a new index file. The dredging index file is obtained by performing the current round of garbage collection on the state data of the blockchain system, and the barrier point is the index file that was last written from the memory to the persistent storage medium when starting to perform the current round of garbage collection; A storage processing unit configured to write the new index file into the persistent storage medium.
9. A computing device, including a memory and a processor. A computer program is stored in the memory. When the processor executes the computer program, the method according to any one of claims 1-7 is implemented.
10. A computer-readable storage medium, on which a computer program is stored. When the computer program is executed in a computing device, the computing device executes the method according to any one of claims 1-7.