Garbage collection method for state data in block chain system and block chain node
By using the data file in the index file to predict the garbage ratio in the blockchain system and efficiently selecting the data files to be recycled, the problem of excessive storage data caused by the increase in the state data version in the blockchain system is solved, and fast and efficient garbage collection and storage efficiency are achieved.
Patent Information
- Application Number
- CN202510201390.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-05-16
AI Technical Summary
With the increase in the version of state data in the blockchain system, the amount of stored data has increased significantly. Continuous data governance is needed to delete earlier versions of state data, but it is difficult for the existing technology to efficiently collect garbage.
By storing multiple data files and index files in a persistent storage medium, using the data files in the index file to predict the junk ratio of the data file, determining the target data file to be recycled, selecting valid data to generate a new data file, and updating the index entry in the index file.
It realizes the garbage ratio of data files quickly and efficiently predicts, accurately selects the data files to be recycled, reducing the computing complexity and storage occupation in the garbage collection process, and improving the storage efficiency of the blockchain system.
Smart Images

Figure CN120011379A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of this specification belong to the field of computer technology, and more particularly to a method for garbage collection of state data in a blockchain system and a blockchain node. Background Art
[0002] The blockchain system is a new application model of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanism, and encryption algorithm. In the blockchain system, data blocks are combined into a chain data structure in a sequential manner according to time order, and a distributed ledger that cannot be tampered with or forged is guaranteed by cryptography. Due to the characteristics of decentralization, information cannot be tampered with, and autonomy, the blockchain system has received more and more attention and application.
[0003] In the blockchain system, state data can be managed through a tree structure, and the aforementioned tree structure may include but is not limited to MPT (Merkle Patricia Tree) and SMT (Sparse Merkle Tree), etc. A leaf node in the tree structure stores the value of a state variable, and part of the information in the directed path from the root node to the leaf node of the tree structure constitutes the key of the state variable. The data storage system can store the key-value pairs of the tree node itself in the tree structure; based on the tree structures corresponding to different versions of state data, different state root (State_Root) hashes can be calculated, and the state root hash can be used as the identifier of the corresponding version of the state data and anchored to the corresponding block header.
[0004] Based on the characteristics of the tree structure, the blockchain system can store multiple versions of state data corresponding to multiple blocks in an append-only manner with a relatively small amount of data storage. However, with the large increase in the number of state data versions, the amount of state data stored in the blockchain system will still increase significantly, and it is necessary to continuously prune / manage the state data stored in the data storage system to delete earlier versions of the state data. Summary of the invention
[0005] The object of the present invention is to provide a method for garbage collection of state data in a blockchain system and a blockchain node.
[0006] In a first aspect, a method for garbage collection of state data in a blockchain system is provided, wherein the blockchain system stores multiple data files and multiple index files through a persistent storage medium, wherein the data files store variable values of state variables, and the index files store multiple index entries and data file usage information determined based on the multiple index entries, wherein the index entries include location information of the variable values of the state variables in the persistent storage medium, and the method comprises: predicting garbage ratios of the multiple data files according to the data file usage information in the multiple index files; determining target data files to be recycled from the multiple data files according to the garbage ratios of the multiple data files; selecting valid data from the target data files and generating new data files using the valid data; deleting the target data files from the persistent storage medium, storing the new data files in the persistent storage medium, and updating the index entries corresponding to the variable values in the new data files in the multiple index files.
[0007] In a second aspect, a blockchain node in a blockchain system is provided, wherein the blockchain system stores multiple data files and multiple index files through a persistent storage medium, wherein the data files store variable values of state variables, wherein the index files store multiple index entries and data file usage information determined based on the multiple index entries, wherein the index entries include location information of the variable values of the state variables in the persistent storage medium, and wherein the blockchain node includes: a garbage ratio prediction unit, configured to predict the garbage ratios of the multiple data files according to the data file usage information in the multiple index files; a target determination unit, configured to determine target data files to be recycled from the multiple data files according to the garbage ratios of the multiple data files; a file generation unit, configured to select valid data from the target data files and generate new data files using the valid data; and an update processing unit, configured to delete the target data files from the persistent storage medium, store the new data files in the persistent storage medium, and update the index entries corresponding to the variable values in the new data files in the multiple index files.
[0008] According to a third aspect, a computing device is provided, including a memory and a processor, wherein a computer program is stored in the memory, and when the processor executes the computer program, the method described in the first aspect is implemented.
[0009] According to a fourth aspect, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed in a computing device, the computing device executes the method described in the first aspect.
[0010] In the technical solution provided by the embodiments of this specification, by directly maintaining the data file usage information of the index file itself in the index file stored in the persistent storage medium, when some index files in the persistent storage medium are deleted and garbage collection of the state data of the blockchain system is required, the garbage ratios of the multiple data files stored in the persistent storage medium at the current moment can be quickly and efficiently predicted based on the data file usage information maintained in the multiple index files that have not been deleted. There is no need to traverse the index entries in the multiple index files to accurately calculate the garbage ratios of the multiple data files each time garbage collection is performed, which is conducive to faster and more efficient implementation of garbage collection of the state data. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] In order to more clearly illustrate the technical solutions of the embodiments of this specification, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this specification. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.
[0012] Figure 1 This is an architectural diagram of a blockchain system exemplarily provided in the embodiments of this specification;
[0013] Figure 2 A schematic diagram of a tree structure for managing state data is provided as an example;
[0014] Figure 3 A schematic diagram of generating a delta page and a base page for a logical page including a plurality of tree nodes is provided as an example;
[0015] Figure 4 A schematic diagram of generating index files for incremental pages and base pages provided as an example;
[0016] Figure 5 A schematic diagram showing the relationship between an index file and a data file provided as an example;
[0017] Figure 6 This is a flowchart of a method for garbage collection of state data in a blockchain system provided in an embodiment of this specification;
[0018] Figure 7 This is a schematic diagram of the structure of a blockchain node provided in an embodiment of this specification. DETAILED DESCRIPTION
[0019] In order to enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of this specification, not all of the embodiments. Based on the embodiments in this specification, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of this specification.
[0020] Figure 1 This is an architectural diagram of a blockchain system provided exemplarily in the embodiments of this specification. The blockchain system may include N blockchain nodes, where Figure 1 8 blockchain nodes, i.e., node 1 to node 8, are shown as examples. The lines between the nodes schematically represent the connection between the nodes, and the aforementioned connection is used to support data transmission between different nodes.
[0021] The blockchain system can provide the function of smart contracts. Smart contracts in the blockchain system are contracts that can be triggered and executed by transactions. Smart contracts can be defined in the form of contract code. Calling a smart contract in the blockchain system is to initiate a transaction pointing to the contract address of the smart contract, so that each node in the blockchain system can run the corresponding contract code in a distributed manner.
[0022] In various blockchain systems that introduce smart contracts, accounts can generally be divided into two types:
[0023] Contract account (CA): mainly used to store the contract code of the corresponding smart contract and the values of the state variables defined in the smart contract. It can usually only be activated by calling an external account.
[0024] Externally owned account (EOA): An account registered by an external user in the blockchain system.
[0025] The design of external accounts and contract accounts is actually a mapping from account addresses to account status. The account status of any account usually includes fields such as nonce, balance, storageRoot, and codeHash. Among them, nonce and balance exist in both external accounts and contract accounts, and codeHash and storageRoot attributes are generally only valid on contract accounts.
[0026] More specifically, for external accounts, the value of nonce represents the number of transactions sent from the relevant account address; for contract accounts, the value of nonce can represent the number of smart contracts created by the relevant account address. The value of Balance represents the number of certain digital resources / tokens owned by the relevant account address. The value of Storage root is a tree structure such as the hash value of the root node of an MPT tree, which is used to organize / manage the storage of state variables of the relevant contract account. The value of CodeHash represents the hash value of the contract code of the relevant smart contract. For external accounts, since smart contracts are not included, the values of the storageRoot and CodeHash fields can generally be empty strings / all-0 strings.
[0027] It should be noted that MPT stands for Merkle Patricia Tree, which is a tree structure that combines Merkle Tree and Patricia Tree (a compressed prefix tree, a more space-saving dictionary tree). The Merkle tree algorithm can calculate a hash value for multiple transactions, and then connect them two by two to calculate the hash again until the top Merkle root. Some blockchain systems usually use an improved MPT tree, such as a hexadecimal tree structure, which is usually referred to as the MPT tree.
[0028] The system data that needs to be persistently stored in the blockchain system can be divided into two parts: block data and status data.
[0029] Block data includes one or more blocks that increase in accordance with the block height (or block number). A single block may include a block header and a block body. The block header may include the block hash previous_Hash (or parent hash), timestamp Timestamp, block number BlockNum, state root hash State_Root, transaction root hash Transaction_Root, receipt root hash Receipt_Root, and random number nonce, etc. The block body may include a transaction set and a receipt set.
[0030] A transaction in a blockchain system refers to a task unit executed and recorded in the blockchain system. A single transaction usually includes a send field (From), a receive field (To), and a data field (Data). The From field includes the account that initiated the transaction (i.e., the sender account), and the To field may include another account involved / pointed to by the transaction.
[0031] For any Nth block, multiple transactions in the transaction set belonging to the Nth block can be executed in order according to the state data with block number (or version number) N-1 to obtain the execution results of the multiple transactions. Then, the state data with version number N-1 is updated according to the execution results of the multiple transactions to obtain the state data with version number k.
[0032] In the blockchain system, state data can be managed through a tree structure, and different versions of state data will correspond to different tree structures. A leaf node of the tree structure stores the location information of the variable value of a state variable in the persistent storage medium, and the directed path between the root node and the leaf node of the tree structure stores the key of the state variable; the aforementioned state variable can be the account address of the contract account / external account, or it can be a state variable in a smart contract. The tree structure can include, for example, MPT (Merkle Patricia Tree) or SMT (Sparse Merkle Tree), etc.
[0033] The tree structure used to manage state data may include a state trie, and the hash value of the root node of the state trie is stored in State_Root in the block header. A leaf node of the state trie stores the location information of the account state of an external account / contract account in the persistent storage medium. The directed path from the root node to a leaf node of the state trie stores the account address of an external account / contract account, or part or all of the hash value calculated based on the account address. As mentioned above, the account state of a single account can usually include fields such as Nonce, Balance, Storageroot, CodeHash, etc. Nonce and Balance exist in both external accounts and contract accounts, and CodeHash and Storage root are generally only valid on contract accounts.
[0034] The tree structure for managing state data may also include a storage trie. The hash value of the root node of the storage trie is stored in the storageroot field of the contract account corresponding to the relevant smart contract, thereby locking the contract state of the smart contract to the relevant contract account through the hash value. Similarly, a leaf node of the storage trie stores the location information of the variable value of a state variable defined in the smart contract in the persistent storage medium, and the directed path from the root node of the storage trie to a leaf node stores the state key of a state variable defined in the relevant smart contract. That is, some of the information in the directed path from the root node to the leaf node of the storage trie can be arranged in sequence to form the key of a state variable defined in the relevant smart contract, and the leaf node stores the location information of the variable value of the state variable in the persistent storage medium.
[0035] Based on the aforementioned tree structure, the blockchain system can separate data and indexes, which helps to improve the flexibility and performance of the system. The basic concepts involved here include data files and index files. Among them, the variable values of state variables (including account addresses or state variables in smart contracts) are stored in the data file; the index file stores index entries pointing to the data file. The implementation principle is: write the data content (variable values of state variables or key-value pairs of state variables) into the data file, record the location information of the data content in the data file (such as file identifier, address offset and data length, etc.); then, create an index entry in the index file, and the index entry contains the search key related to the data content and the aforementioned location information. In this way, when it is necessary to query the data content, after obtaining the search key of the data content, you can find the index entry containing the search key in the index file, obtain the location information of the data content from the index entry, and then read the actual data content from the corresponding data file according to the location information.
[0036] For example, refer to Figure 2As shown in the figure, for the tree structure corresponding to the state data with version number N, in the upper-level MPT, i.e., in the state trie, for the leaf node A1, through the shared nibble a7 in the root node A8 (Extension Node)—intermediate node A7 (Branch Node) slot 1—1335 of key-end in leaf node A1, can be combined in sequence to form the key of a state variable: a711335, the value of the state variable is "Nonce=n1, Balance=45.0ETH", the data file used to store "Nonce=n1, Balance=45.0ETH" is file 5, the file identifier of file 5 is "5", the address offset of "Nonce=n1, Balance=45.0ETH" in file 5 is 600, the data length of "Nonce=n1, Balance=45.0ETH" is 100, then the leaf node A1 can, for example, use the locatione field to store the value "Nonce=n1, Balance=45.0ETH" in the persistent storage medium. Location information (5, 600, 100). Similar to the principle of leaf node A1, leaf node A2 can store the location information (3, 805, 150) of the state variable value "Nonce = n2, Balance = 1.00WEI" with key a77d337 in the persistent storage medium through the location field; leaf node A3 can store the location information (2, 100, 180) of the state variable value "Nonce = n3, Balance = 1.1ETH" with key a7f9365 in the persistent storage medium through the location field; leaf node A4 can store the location information P1 of the state variable value "Nonce = n4, Balance = 0.12ETH, CodeHash = c1, Storage root = s1" with key a77d397 in the persistent storage medium through the location field, s1 can be the hash value H (A10) of the tree node A10, that is, the hash value of the root node A10 of the next layer of the tree. It should be noted that Figure 1 In order to illustrate the relationship between the next-level MPT and the previous-level MPT, the leaf node A4 illustrates the variable values "Nonce = n4, Balance = 0.12ETH, CodeHash = c1, Storage root = s1. In fact, the leaf node A4 should store the location information P1 through the location field. However, Figure 2It is not shown in the figure. Among them, leaf nodes A1, A2 and A3 correspond to external accounts, and leaf node A4 corresponds to contract accounts. For contract accounts, it contains the next level MPT, forming a Storage Trie, which is used to store the state of the state variables in the smart contract corresponding to the contract account.
[0037] like Figure 2 As shown in the example, in the next-level MPT (i.e., storage trie), for leaf node A11, slot 3 in root node A10 (Branch Node) - 35b2e4 of key-end in leaf node A11 are sequentially combined to form the key of a state variable: 335b2e4, and the value of the state variable is "Zhang San_A=20". "Zhang San_A=20" means, for example, that the share of type A digital assets belonging to Zhang San defined in the contract is 20, that is, the balance of Zhang San's type A assets is 20, wherein the data file used to store the value "Zhang San_A=20" is file 2, and the file identifier of file 2 is "2". The address offset of "Zhang San_A=20" in file 2 is 7500, and the data length of the value "Zhang San_A=20" is 210. For example, in leaf node A11, the locatione field can be used to store the location information (2, 750, 210) of "Zhang San_A=20" in the persistent storage medium. Similar to the principle of leaf node A11, leaf node A12 can store the location information (3, 350, 210) of the value of the state variable "Li Si_B=50" with the key 7c25988 in the persistent storage medium through the location field. "Li Si_B=50" means, for example, that the share of the B-type digital asset defined in the contract belonging to Li Si is 50, that is, the balance of Li Si's B-type assets is 50; leaf node A15 can store the location information (5, 760, 140) of the value of the state variable "storedData=s" with the key fa6be33 in the persistent storage medium through the location field; leaf node A16 can store the location information (5, 170, 210) of the value of the state variable "Wang Wu_A=35" with the key fa99365 in the persistent storage medium through the location field.
[0038] In the node structure of the aforementioned MPT, the prefix prefix is used to indicate the tree node type. For example, 0 indicates an Extension Node containing an even number of shared nibbles (shared half bytes), 1 indicates an Extension Node containing an odd number of shared nibble(s), 2 indicates a Leaf Node containing an even number of nibbles, and 3 indicates a Leaf Node containing an odd number of nibble(s).
[0039] In the above node structure, the hash value of the entire content of the next tree node is filled in the corresponding position of the previous tree node.
[0040] Based on the tree structure in the above example, the key-value pairs of the tree nodes in the tree structure can be obtained. The key of the tree node can be the result of a hash operation on the entire content of the tree node (i.e., the value of the tree node). In this way, the key-value pairs of the tree node can be used as index entries and stored in the corresponding index file. Figure 2 For example, the index file can be Figure 2 The example tree structure stores the key-value pairs of the tree nodes shown in Table 1 as index entries.
[0041]
[0042]
[0043] Table 1
[0044] In Table 1 above, H() is used to represent hash calculation. In this way, the hash value of the next tree node is anchored in the previous tree node. Through such layers of hashing, the root hash of the entire state trie is obtained, and the root hash is locked in the state root field of the block header. Assume that the kv pairs in Table 1 are saved on disk as index entries in the index file, and the LSM structure is adopted. In this way, after querying the root node of the state trie and matching the key of the state variable to be searched with the shared nibble(s) field (for Extension Node) or slot (for Branch Node) of the root node from the beginning, the hash of the next layer of tree nodes can be read from the matching position, and then the next InternalNode pointed to by the hash value can be found; after unlocking the Internal Node, the fields of the remaining part of the key of the state variable to be read are continued from front to back. If a match is found, the hash value is read from the matching position, and then jump to the next level of tree node pointed to by the hash value. This is repeated continuously, the Internal Node is unlocked level by level, and the remaining fields of the key of the state variable to be read are matched from front to back. The matched hash value is used as the basis for the next search for the intermediate node or leaf node, until the Leafnode is matched, so as to read the location information of the value of the state variable in the persistent storage medium from the Leaf node. Exemplary, for example, the value of the leaf node A11, i.e., the content in "prefix:2, Key-end:35b2e4, location: (2, 750, 210)", can be loaded into the memory in the end, so as to obtain the location information "2, 750, 210" of the value "Zhang San_A=20" of the state variable with the key of 335b2e4 in the persistent storage medium, and then the data content between 750KB and 960KB can be read from file 2 according to the location information "2, 750, 210", and finally the value "Zhang San_A=20" of the state variable with the key of 335b2e4 can be obtained.
[0045] When the key-value pair of a tree node is directly used as an index entry, the key of the tree node is used as a search key. Alternatively, the key-value pair of a tree node may not be directly used as an index entry. For example, for any tree node, the key of the state variable stored in the directed path from the root node to the tree node or a component of the key of the state variable, that is, the key of the lexicographic content distribution from the root node through the intermediate nodes to the tree node (hereinafter referred to as the node ID or nodeID), can be combined with the version number / block number of the state data to form a search key, and the location information of the value of the state variable included in the value of the tree node in the persistent storage medium is combined with the search key to obtain an index entry corresponding to the tree node.
[0046] Based on the tree structure of the above example, the entire tree structure, i.e., the Merkle dictionary tree, can also be divided into multiple logical pages (Logical Pages). Specifically, several upper and lower adjacent tree nodes can be aggregated into one LogicalPage according to the node association relationship of the tree structure, such as aggregating 256 child nodes in the 2-level 16-fork dictionary tree into one LogicalPage. Among them, there are multiple layers of LogicalPages from the root node to the leaf node of the tree structure, and there are parent-child and brother relationships between LogicalPages, thus forming a complete tree structure. In this way, a logical page contains at least one tree node, and different logical pages contain different tree nodes, and the tree nodes corresponding to the current latest version of the status data can be maintained in the logical page.
[0047] like Figure 3As shown, each LogicalPage can maintain a memory page (MemoryPage) to represent the contents of all tree nodes corresponding to the current latest version of the logical page. Furthermore, the contents of the tree nodes maintained by the logical page can be generated into a base page (BasePage) and a delta page (DeltaPage) according to the generation / modification behavior based on the continuous version update. Specifically, the BasePage and DeltaPage are generated according to the generation / modification behavior. For example, for continuous state changes, such as the state variables a=5, b=8, and c=3 included in version N, a=5, b=8, and c=3 can be used as the contents of the base page BasePage. For a=6 in version N+1 and a=8 in version N+2, assuming that b and c have not changed, a=6 in version N+1 and a=8 in version N+2 can be used as DeltaPages based on this BasePage, such as DeltaPage1, and the case where b and c have not changed is not included in DeltaPage1. For a=11 in version N+3 and a=15 in version N+4, a=11 in version N+3 and a=15 in version N+4 can be used as DeltaPages based on this BasePage, such as DeltaPage2. Also, assuming that b and c have not changed, b and c are not included in DeltaPage2, and so on. It should be noted that BasePage and DeltaPage can be used as the smallest unit of the memory structure of the tree structure. In addition, BasePage can be used as the smallest unit of disk persistence of the tree structure. On this basis, DeltaPage can also be used as the smallest unit of disk persistence of the tree structure.
[0048] For BasePage, a BasePage can be generated once for each predetermined number of state modifications (e.g., m times). For example, in the above example, in version N, the state variable a=5 can be used as the content of the base page BasePage1. For version N+1, a=6, and version N+2, a=8, a=6 in version N+1 and a=8 in version N+2 are used as the incremental page DeltaPage1 based on this BasePage. For version N+3, a=11, and version N+4, a=15, a=11 in version N+3 and a=15 in version N+4 are used as the incremental page DeltaPage2 based on this BasePage. For version N+5, a=20, assuming that the modification reaches the predetermined number of times 5, the value a=20 in version N+5, b=8, and c=3 can be used as the content of the new base page BasePage2.
[0049] For DeltaPage, it describes several version modifications of LogicalPage, and aggregates multiple modifications into a set, such as generating a DeltaPage every M modifications, and creating a new DeltaPage to collect the subsequent version modification operations. As mentioned above, for a=6 in version N+1 and a=8 in version N+2, a=6 in version N+1 and a=8 in version N+2 are used as the incremental page DeltaPage1 based on this BasePage. For a=11 in version N+3 and a=15 in version N+4, a=11 in version N+3 and a=15 in version N+4 can be used as the incremental page DeltaPage2 based on this BasePage. It can be seen that DeltaPage is generated every 2 modifications.
[0050] For each modification of BasePage and DeltaPage, corresponding dirty data is generated in the memory. Usually, when the memory occupied by BasePage and DeltaPage reaches a predetermined share, BasePage and DeltaPage can be persisted to disk in batches, so as to avoid persisting all dirty data for each modification and continuously occupying memory and CPU resources.
[0051] The previous article describes DeltaPage and BasePage by the value changes of the state variables in LogicalPage. However, the logical page maintains the contents of the tree nodes. According to the continuous version updates of the state data, DeltaPage and BasePage are generated according to the generation / modification behavior. What is described in DeltaPage and BasePage is still the key-value pairs of the tree nodes associated with the version number.
[0052] Similar to the node ID, the logical page can have a page ID. The page ID can be the lexicographic content from the root node to the top tree node in the logical page (i.e., the memory page), that is, the page ID can be the key of the state variable or a component of the key of the state variable distributed from the root node through the intermediate nodes to the top tree node in the LogicalPage.
[0053] DeltaPage and BasePage can also have versions. BasePage generally corresponds to a memory page, including all tree nodes in the corresponding memory page, which also includes global state variables, so it can have the same version as the corresponding memory page. DeltaPage can correspond to one or more memory pages, including tree nodes in one or more consecutive versions of memory pages that have changed relative to the previous adjacent DeltaPage / BasePage, which also includes changed state variables, but generally not global state variables, so the version of the delta page can be the lowest or highest version in the one or more memory pages corresponding to it.
[0054] For example, a BasePage with version 4 contains state variables a=1, b=2, c=3 and corresponding intermediate nodes and root nodes. A DeltaPage generated after the BasePage may include a=2 with version 5 and corresponding intermediate nodes and root nodes, a=1 with version 6 and corresponding intermediate nodes and root nodes, and b=4 with version 7 and corresponding intermediate nodes and root nodes. In this case, although a=1 with version 6 is the same as the state variable a=1 contained in the BasePage, it is different from the previous adjacent a=2 with version 5, so a=1 with version 6 and corresponding intermediate nodes and root nodes are also included in the DeltaPage. In addition, for the next DeltaPage, that is, the next DeltaPage corresponding to the DeltaPages of versions 5, 6, and 7, for example, the DeltaPage containing tree nodes of versions 7, 8, and 9, is actually in order relative to the previous adjacent DeltaPage and the tree nodes of the version. Therefore, tree nodes with versions 5, 6, and 7 correspond to memory pages with versions 5, 6, and 7 respectively. The DeltaPage containing tree nodes with versions 5, 6, and 7 can have a version of 5 or 7, that is, it can be the lowest or highest version among the corresponding multiple memory pages.
[0055] Moreover, you can set a page type for DeltaPage and BasePage respectively to distinguish DeltaPage from BasePage.
[0056] In combination with the aforementioned determination of index entries based on tree nodes, for example, using the key-value pairs of tree nodes as index entries, when persisting the aforementioned BasePage and DeltaPage pages, the index file corresponding to the BasePage / DeltaPage can actually be specifically stored in the persistent storage medium, and the index file can include the key-value pairs of all tree nodes in the BasePage / DeltaPage. Moreover, the page ID of the logical page corresponding to the BasePage / DeltaPage can be combined with the page type of the BasePage / DeltaPage and the version of the BasePage / DeltaPage to form the file identifier of the index file corresponding to the BasePage / DeltaPage. In addition, the data file usage information of the index file can be determined based on the multiple key-value pairs (i.e., index entries) included in the index file, and the data file usage information can be stored in the corresponding index file.
[0057] The following first describes the process of generating an index file for DeltaPage / BasePage by way of example.
[0058] In some embodiments, when the blockchain system generates the latest version of DeltaPage / BasePage for any logical page of the latest version based on the version update of the status data, the key-value pairs of several tree nodes included in the DeltaPage / BasePage can be directly used as several index entries, and then the data file usage information of the index file corresponding to the DeltaPage / BasePage is determined based on the position information in the several index entries, such as the position information in the value of the tree node.
[0059] Reference Figure 4As shown, it is assumed here that the tree structure corresponding to the state data with version number / block number 100 includes the logical page Leaf Page C1, and Leaf Page C1 includes the extension node Branch Node A, leaf node LeafNode A1, leaf node Leaf Node A2, leaf node Leaf Node A3, leaf node Leaf Node A4 and leaf node Leaf Node A5; the key of a state variable stored in the directed path between the root node Root and Leaf Node A1 is a7r5x1, the key of a state variable stored in the directed path between the root node Root and Leaf Node A2 is a7r5x2, the key of a state variable stored in the directed path between the root node Root and Leaf Node A3 is a7r5x3, the key of a state variable stored in the directed path between the root node Root and Leaf Node A4 is a7r5x4, and the key of a state variable stored in the directed path between the root node Root and Leaf Node A5 is a7r5x5. Moreover, it is assumed here that the value of a7r5x1 in the state data with version number 100 is y1, the value of a7r5x2 in the state data with version number 100 is y2, the value of a7r5x3 in the state data with version number 100 is y3, the value of a7r5x4 in the state data with version number 100 is y4, and the value of a7r5x5 in the state data with version number 100 is y5; in this case, in the tree structure corresponding to the state data with version number 100, more specifically, in the memory page corresponding to Leaf Page C1, the value of Leaf Node A1 will include the location information L(y1) of y1 in the persistent storage medium, the value of Leaf Node A2 will include the location information L(y2) of y2 in the persistent storage medium, the value of Leaf Node A3 will include the location information L(y3) of y3 in the persistent storage medium, and the value of Leaf Node The value of A4 will include the location information L(y4) of y4 in the persistent storage medium, and the value of Leaf Node A5 will include the location information L(y5) of y5 in the persistent storage medium.
[0060] Next, assume that a new DeltaPage is generated for Leaf Page C1 every two modifications in the blockchain system, and assume that in the state data with version number 101, the value of a7r5x1 is updated from y1 to y6, the value of a7r5x2 is updated from y2 to y7, and assume that in the state data with version number 102, the value of a7r5x1 is updated from y6 to y8, and the value of a7r5x3 is updated from y3 to y9. Then, in the tree structure corresponding to the state data with version number 101, the location information L(y1) included in the value of Leaf Node A1 will be updated to L(y6), and the location information L(y2) included in the value of Leaf Node A2 will be updated to L(y7); in the tree structure corresponding to the state data with version number 102, the location information L(y6) included in the value of Leaf Node A1 will be updated to L(y8), and the location information L(y3) included in the value of Leaf Node A3 will be updated to L(y9). Based on the assumptions of the above example, the key-value pairs of the tree nodes shown in the following Table 2 may be recorded in the DeltaPage with version number 101, NodeID a7r5x and type delta.
[0061] The key of the tree node The value of the tree node A1(101) V(y6) A2(101) V(y7) A(101) V(A101) A3(102) V(y9) A1(102) V(y8) A(102) V(A102)
[0062] Table 1
[0063] In the above Table 1, A1(101) represents the key of Leaf Node A1 in the tree structure with version number 101, and V(y6) represents the value of Leaf Node A1 in the tree structure with version number 101; A2(101) represents the key of Leaf Node A2 in the tree structure with version number 101, and V(y7) represents the value of Leaf Node A2 in the tree structure with version number 101; A(101) represents the key of Branch Node A in the tree structure with version number 101, and V(A101) represents the value of Branch Node A in the tree structure with version number 101. Similarly, A3(102) represents the key of LeafNode A3 in the tree structure with version number 102, and V(y9) represents the value of Leaf Node A3 in the tree structure with version number 102; A1(102) represents the key of Leaf Node A1 in the tree structure with version number 102, and V(y8) represents the value of LeafNode A1 in the tree structure with version number 101; A(102) represents the key of Branch Node A in the tree structure with version number 102, and V(A102) represents the value of Branch Node A in the tree structure with version number 102.
[0064] The key-value pairs of several tree nodes included in the DeltaPage can be directly used as several index entries in the index file corresponding to the DeltaPage, or only the key-value pairs of all leaf nodes included in the DeltaPage can be used as index entries in the index file corresponding to the DeltaPage. For example, in the DeltaPage of the aforementioned example with version number 101, NodeID a7r5x and type delta, V(y6) contains the location information L(y6) of the variable value y6 in the persistent storage medium, V(y7) contains the location information L(y7) of the variable value y7 in the persistent storage medium, V(y8) contains the location information L(y8) of the variable value y8 in the persistent storage medium, and V(y9) contains the location information L(y9) of y9 in the persistent storage medium; V(A101) and V(A102) do not contain the location information of the value of any state variable in the persistent storage medium, and there is no need to use the key-value pairs "A(101)-V(A101)" and "A(102)-V(A102)" of the tree nodes as index entries.
[0065] Alternatively, you can use the key-value pairs of tree nodes as index entries instead of directly using them. Figure 4As shown, for the key-value pair of the tree node "A1(101)-V(y6)", the NodeID of the tree node leaf NodeA1 corresponding to the key-value pair is a7r5x1, V(y6) contains the location information L(y6), y6 is the value of a7r5x1 in the state data with version number 101, so the version number such as 101 and the key of the NodeID / state variable such as "a7r5x1" can be used to form a search key "T(101+a7r5x1)" in an index entry, and the location information L(y6) is recorded in the index entry. Furthermore, for the DeltaPage in the above-mentioned Table 2 example, the various index entries in the following Table 3 example can be recorded in the corresponding index file such as Index file D1.
[0066] Retrieve key Location Information T(101+a7r5x1) L(y6) T(101+a7r5x2) L(y7) T(102+a7r5x1) L(y8) T(102+a7r5x3) L(y9)
[0067] Table 3
[0068] As mentioned above, the location information of data content (such as state variable value or key-value pair of state variable) in the persistent storage medium may include the file identifier, address offset and data length of the data file to which the data content belongs. Correspondingly, for the data file usage information included in the index file corresponding to DeltaPage, the data file usage information may include a number of usage records, and the number of usage records are used to indicate a number of first file identifiers and their respective corresponding data occupancy ratios. The aforementioned number of first file identifiers are extracted from the location information of a number of index entries included in the index file corresponding to DeltaPage. The data occupancy ratio is the ratio between the cumulative data length of a number of target variable values corresponding to a number of target index entries and the file size of the data file indicated by the corresponding first file identifier. The aforementioned number of target index entries belong to a number of index entries and include the corresponding first file identifiers.
[0069] Continuing with the example in Table 3 above, refer to Figure 5As shown, it is further assumed here that the variable values y6, y7, y8 and y9 are all stored in the data file Data file F1, that is, in the index file corresponding to the DeltaPage in the above-mentioned Table 2 example, such as Indexfile D1, the position information in all index entries, namely L(y6), L(y7), L(y8), L(y9), all point to the same data file Data file F1, that is, the first file identifiers included in the position information in the several index entries can specifically be the file identifier F1. In addition, assuming that the amount of data or the file size in Data file F1 is M bytes, the data lengths contained in the positions L(y6), L(y7), L(y8), L(y9) containing the same file identifier F1 can be summed to obtain the cumulative data length P1; the ratio between the cumulative data length P1 and the amount of data / file size M is the data occupancy ratio K1 of y6, y7, y8 and y9 corresponding to L(y6), L(y7), L(y8), L(y9) containing the file identifier F1 in Data file F1. That is Figure 4 and Figure 5 In the example Index file D1, Data file useful may include a file identifier F1 and a corresponding data occupancy ratio K1.
[0070] Based on a similar process, please continue to refer to Figure 4 and Figure 5 Here, we continue to assume that for a DeltaPage with version number 103, NodeID a7r5x and type delta, several index entries as shown in the following Table 4 can be determined.
[0071]
[0072]
[0073] Table 4
[0074] Based on the above Table 4, refer to Figure 5As shown, it is further assumed that the variable values y10 and y11 are stored in the data file Data file F1, and the variable values y12 and y13 are stored in the data file Data file F2. L(y10) and L(y11) include the same file identifier F1, and L(y12) and L(y13) include the same file identifier F2. Here, we still assume that the data volume in Data file F1 and Data file F2 is M bytes. Then, we can sum the data lengths contained in the positions L(y10) and L(y11) containing the same file identifier F1 to obtain the cumulative data length P2. The ratio between the cumulative data length P2 and the data volume M is the data occupancy ratio K2 of y10 and y11 corresponding to L(y10) and L(y11) containing the file identifier F1 in Data file F1. Similarly, we can sum the data lengths contained in the positions L(y12) and L(y13) containing the same file identifier F2 to obtain the cumulative data length P3. The ratio between the cumulative data length P3 and the data volume M is the data occupancy ratio K3 of y12 and y13 corresponding to L(y12) and L(y13) containing the file identifier F2 in Data file F2. That is, Figure 4 and Figure 5 For example, for Index file D2, the data file useful information Data file useful may include, for example, the file identifier F1 and the corresponding data occupancy ratio K2, and may also include the file identifier F2 and the corresponding data occupancy ratio K3.
[0075] Continuing from the above Figure 4 , then assume that a new BasePage is generated for Leaf Page C1 every 5 modifications in the blockchain system, and assume that in the state data with version number 105, the value of a7r5x2 is updated from y13 to y14, and the value of a7r5x4 is updated from y11 to y15. Then, in the tree structure corresponding to the state data with version number 105, the location information L(y13) included in the value of LeafNode A2 will be updated to L(y14), and the location information L(y11) included in the value of LeafNode A4 will be updated to L(y15). Based on the assumptions of the above example, the key-value pairs of the tree nodes shown in the example in Table 5 below may be recorded in the BasePage with version number 105, NodeID a7r5x and type base.
[0076] The key of the tree node The value of the tree node A1(105) V(y12) A2(105) V(y14) A3(105) V(y9) A4(105) V(y15) A5(105) V(y5) A(105) V(A105)
[0077] Table 5
[0078] In the above Table 5, A1(105) represents the key of Leaf Node A1 in the tree structure with version number 105, V(y12) represents the value of Leaf Node A1 in the tree structure with version number 105; A2(105) represents the key of Leaf Node A2 in the tree structure with version number 105, V(y14) represents the value of Leaf Node A2 in the tree structure with version number 105; A3(105) represents the key of Leaf Node A3 in the tree structure with version number 105, V(y9) represents the value of Leaf Node A3 in the tree structure with version number 105; A4(105) represents the key of Leaf Node A4 in the tree structure with version number 105, V(y15) represents the value of Leaf Node A5 in the tree structure with version number 105; A(105) represents the key of Branch Node A2 in the tree structure with version number 105. The key of Node A, V(A105), represents the value of BranchNode A in the tree structure with version number 105.
[0079] The key-value pairs of several tree nodes included in the BasePage can be directly used as several index entries in the index file corresponding to the BasePage. Alternatively, only the key-value pairs of all leaf nodes included in the BasePage can be used as index entries in the index file corresponding to the BasePage. For example, in the key-value pair of the tree node in the above-mentioned Table 5 example, V(A105) in "A(105)-V(A105)" does not contain the location information of the value of any state variable in the persistent storage medium, so there is no need to use the key-value pair of the tree node "A(105)-V(A105)" as an index entry.
[0080] Or the key-value pairs of tree nodes are not directly used as index entries. Figure 4As shown, for the key-value pair of the tree node "A1(105)-V(y12)", the NodeID of the tree node leaf Node A1 corresponding to the key-value pair is a7r5x1, V(y12) contains the location information L(y12), y12 is the value of the related state variable in the state data with version number 105, and the version number such as 105 and the key of the NodeID / state variable such as "a7r5x1" can be used to form a search key "T(105+a7r5x1)" in an index entry, and the location information L(y12) is recorded in the index entry. Furthermore, for the BasePage in the above-mentioned Table 5 example, the various index entries in the following Table 6 example can be recorded in the corresponding index file such as Index file D2.
[0081] Retrieve key Location Information T(105+a7r5x1) L(y12) T(105+a7r5x2) L(y14) T(105+a7r5x3) L(y9) T(105+a7r5x4) L(y15) T(105+a7r5x5) L(y5)
[0082] Table 6
[0083] Based on the example in Table 6 above, refer to Figure 6 As shown, it is further assumed here that the variable values y5 and y9 are stored in the data file Data file F1, and the variable values y12, y14, and y15 are stored in the data file Data file F2, that is, in the index file corresponding to the BasePage in the above-mentioned Table 6 example, such as Index file D2, the index entries corresponding to the tree nodes leaf Node A3 and leaf NodeA5 all point to the same data file Data file F1, and the index entries corresponding to the tree nodes leaf Node A1, leafNode A2, and leaf Node A4 all point to the same data file Data file F2. The first file identifiers extracted from the index entries in the above-mentioned Table 6 example may specifically include the file identifiers F1 and F2. Continuing to assume that the data volume in Data file F1 and Data file F2 is M bytes, the data lengths of L(y12), L(y14), and L(y15) containing the same file identifier F2 can be summed to obtain the cumulative data length P4. The ratio between the cumulative data length P4 and the data volume M is the data occupancy ratio K4 of y12, y14, and y15 corresponding to L(y12), L(y14), and L(y15) containing the file identifier F2 in Data file F2. Similarly, the data lengths of L(y9) and L(y5) containing the same file identifier F1 can be summed to obtain the cumulative data length P5. The ratio between the cumulative data length P5 and the data volume M is the data occupancy ratio K5 of y9 and y5 corresponding to L(y9) and L(y5) containing the file identifier F1 in Data file F1. In other words, in the aforementioned Figure 4 and Figure 5 In the example Index file D2, Data file useful may include: a file identifier F2 and a corresponding data occupancy ratio K4; a file identifier F1 and a corresponding data occupancy ratio K5.
[0084] The previous article mainly describes the process of generating corresponding index files for incremental pages and base pages generated based on logical pages in the blockchain when the tree structure is divided into multiple logical pages. However, in actual technical scenarios, the blockchain system may not divide logical pages and generate base pages / incremental pages for logical pages. In this case, the tree structure of each version may be directly divided into logical partitions, and the index entries corresponding to the tree nodes of different logical partitions may be divided into different index files.
[0085] As the number of state data versions increases, the amount of state data stored in the persistent storage medium will increase significantly. It is necessary to continuously prune / manage the state data stored in the data storage system to delete the earlier versions of the state data. In this process, the index file will be deleted according to the specified version number, for example, see the previous article combined with Figure 4 and Figure 5 As described in the relevant description, when the specified version number is 105, the index files corresponding to the incremental pages / base pages with version numbers less than 105, such as Index file B1, Index file D1, and Index file D2, will be deleted. After the index files are deleted, some or all of the data contents (values or key-value pairs of state variables) in some data files become invalid. In this case, some or all of the invalid data contents can be eliminated through the corresponding garbage collection process to save storage space.
[0086] In the embodiments of the present specification, a method for garbage collection of state data in a blockchain system and a blockchain node are provided. The blockchain system stores multiple data files and multiple index files through a persistent storage medium, wherein the variable values of state variables are stored in the data files, and multiple index entries and data file usage information determined based on the multiple index entries are stored in the index files, wherein the index entries include the location information of the variable values of the state variables in the persistent storage medium; if garbage collection of state data is required, the garbage ratio of the multiple data files can be determined according to the data file usage information stored in each of the multiple index files; according to the garbage ratio of the multiple data files, the target data files to be recycled are determined from the multiple data files; valid data are selected from the target data files to generate new data files; the target data files are deleted from the persistent storage medium, the new data files are stored in the persistent storage medium, and the index entries corresponding to the variable values in the new data files are updated in the multiple index files.
[0087] In this way, by directly maintaining the data file usage information of the index file itself in the index file, when some index files are deleted and garbage collection of the status data is required, the garbage ratios of the multiple data files stored in the persistent storage medium at the current moment can be quickly and efficiently predicted based on the data file usage information maintained in the multiple index files that have not been deleted. There is no need to traverse the index entries in the multiple index files to accurately calculate the garbage ratios of the multiple data files each time garbage collection is performed, which is conducive to faster and more efficient implementation of garbage collection of the status data.
[0088] Figure 6 This is a flowchart of a method for garbage collection of state data in a blockchain system provided in an embodiment of this specification.
[0089] As mentioned above, the blockchain system can store multiple data files and multiple index files through a persistent storage medium, the data files store the variable values of the state variables, the index files store multiple index entries and data file usage information determined based on the multiple index entries, and the index entries include the location information of the variable values of the state variables in the persistent storage medium.
[0090] Reference Figure 6 As shown, the method may include but is not limited to part or all of the following steps S601 to S607.
[0091] Step S601, predicting garbage ratios of a plurality of data files according to data file usage information stored in a plurality of index files.
[0092] Referring to the process of generating an index file described in the above exemplary embodiment, the data file usage information stored in the index file may include the first file identifiers of several data files referenced by the index file, and the data occupancy ratios corresponding to the several first file identifiers. More specifically, the data file usage information in the index file may include several usage records, the several usage records are used to indicate several first file identifiers and the data occupancy ratios corresponding to them, the several first file identifiers are extracted from the position information in several index entries, the data occupancy ratio is the ratio between the cumulative data length of the variable values corresponding to several target index entries and the file size of the data file indicated by the corresponding first file identifier, and the several target index entries belong to several index entries and include the corresponding first file identifiers. For example, referring to the example in the previous article, the data file usage information Data file useful stored in the index file Index file D1 may specifically include the file identifier F1 and its corresponding data occupancy ratio K1; the data file usage information Data file useful stored in the index file Index file D2 may specifically include the file identifier F1 and its corresponding data occupancy ratio K2, the file identifier F2 and its corresponding data occupancy ratio K3; the data file usage information Data file useful included in the index file Index file B2 may specifically include the file identifier F2 and its corresponding data occupancy ratio K4, the file identifier F1 and its corresponding data occupancy ratio K5.
[0093] In some embodiments, for any data file among the multiple data files, several target data occupancy ratios corresponding to the file identifier of the data file can be determined from several usage records respectively included in the multiple index files, and the garbage ratio of the data file can be determined based on the several target data occupancy ratios. More specifically, for the multiple data file usage information stored in the multiple index files, all data occupancy ratios corresponding to the same file identifier can be directly summed to obtain a cumulative occupancy ratio, and the garbage ratio of the data file indicated by the file identifier can be obtained by subtracting the cumulative occupancy ratio from 1. For example, assuming that at the current moment, the index files stored in the persistent storage medium specifically include the index files Index file B1, Index file D1, Index file D2 and Index file B2 in the aforementioned example, and the data file usage information Data file useful in Index file B1 specifically includes the file identifier F1 and its corresponding data occupancy ratio k0; continuing the previous exemplary description of the Data file useful included in Index file D1, Index file D2 and Index file B2, it can be calculated that the garbage ratio of Data file F1 is 1-(k0+k1+k2+k5), and the garbage ratio of Data file F2 is 1-(k3+k4).
[0094] For any data file, the predicted garbage ratio may be less than or equal to the actual garbage ratio. This is because the index file corresponding to BasePage, such as Index file B2, may contain index entries corresponding to tree nodes that have not been generated / modified in the previous version, such as the logical page of version 104. Figure 4 In the BasePage with version 105, node ID a7r5x and type base, the key-value pairs of leaf Node A1, leaf Node A3 and leaf NodeA5 have not changed compared with version 104, and the variable values y5 and y9 are stored in Data file F1. The data occupancy ratio corresponding to file identifier F1 is counted in Index file D1 and Index file B2 based on the location information of y5 and y9. However, for the predicted garbage ratios of multiple data files, it still means that the larger the predicted garbage ratio, the smaller the proportion of valid data in the corresponding data file. Therefore, based on the predicted garbage ratio, the target data file with a relatively small proportion of valid data and to be recycled can still be accurately selected from multiple data files.
[0095] The above mainly describes the data file usage information including several first file identifiers and their corresponding data occupancy ratios by way of example. In actual technical scenarios, the data file usage information may not directly include the data occupancy ratio, but may include other information that can be used to calculate the data occupancy ratio. For example, the data occupancy ratio may replace the number of variable values in the data file indicated by the corresponding file identifier and the average data length of the variable values. In this way, in the process of executing the aforementioned step S601, it is necessary to calculate the garbage ratio after calculating the data occupancy ratio corresponding to the first file identifier.
[0096] Step S603: determining a target data file to be recycled from the multiple data files according to the garbage ratios of the multiple data files.
[0097] One or more data files corresponding to garbage ratios may be selected in descending order as target data files to be recycled; or data files with garbage ratios greater than a preset threshold may be selected as target data files to be recycled.
[0098] Step S605: Select valid data from the target data file and generate a new data file using the valid data.
[0099] In the aforementioned steps S603 to S605, valid data may be selected from one target data file to generate a new data file; or valid data may be selected from multiple target data files to generate a new data file.
[0100] The index entries in the multiple index files stored in the persistent storage medium at the current moment can be traversed, and when the location information included in the index entry points to the target data file, the data content corresponding to the location information in the target data file is used as valid data. That is, all index entries in all index files stored in the persistent storage medium at the current moment can be traversed to obtain all target location information containing the file identifier of the target data file, and all valid data (which can be the value of the state variable or the key-value pair of the state variable) can be obtained from the target data file according to the target location information.
[0101] For example, refer to Figure 5As shown, it is further assumed here that the index files Index file B1, Index file D1, and Index file D2 in the persistent storage medium have been deleted at the current moment, the index files stored in the persistent storage medium at the current moment include Index file B2, and the data files Data file F1 and Data file F2 are both determined as target data files to be recovered. When traversing the index entries in Index file B2, valid data y9 and y5 can be selected from the data file Data file F1 according to the position information L(y9) and L(y5), and valid data y12, y14, and y15 can be selected from the data file Data file F2 according to the position information L(y12), L(y14), and L(y15).
[0102] When the key-value pairs of the state variables are stored in the data file, the key-value pairs of the state variables stored in the data file can also be traversed. If there is an index entry pointing to the key-value pair of the state variable in a certain index file, the key-value pair of the state variable can be used as valid data.
[0103] The selected valid data can be stored by generating a new data file. Figure 5 As shown, for the valid data y5, y9, y12, y14 and y15 selected from the target data files Data file F1 and Data file F2, the valid data can be stored in a new data file, such as Data file F3.
[0104] Step S607: deleting the target data file from the persistent storage medium, storing the new data file in the persistent storage medium, and updating the index entries corresponding to the variable values in the new data file in the plurality of index files.
[0105] Continuing with the previous example, refer to Figure 5As shown. For example, the target data files Data file F1 and Data file F2 can be deleted from the persistent storage medium, and a new data file Data file F3 can be stored in the persistent storage medium. In addition, since the location information of valid data such as the aforementioned variable values y5, y9, y12, y14 and y15 in the persistent storage medium has changed, it is also necessary to update the index entries corresponding to the valid data such as y5, y9, y12, y14 and y15 in the corresponding index file, for example, it is necessary to update the location information L(y12) of y12 in the index entry corresponding to the variable value y12 stored in the index file Index file B2, for example, update the file identifier in L(y12) from F2 to F3, and update the address offset of y12 in Data file F2 to the address offset of y12 in Data file F3.
[0106] Correspondingly, when an index entry in a certain index file is updated, referring to the process of determining the data file usage information stored in the index file based on the index entries stored in the index file, it may also be necessary to update the data file usage information stored in the index file accordingly. Continuing with the example in the previous text, for Data file useful in Index file B2, it can be updated to file identifier F3 and its corresponding data occupancy ratio K6, where data occupancy ratio K6 is the ratio between the cumulative data length of y5, y9, y12, y14 and y15 and the data volume / file size M of Data file F3.
[0107] It should be noted that in the exemplary description of the aforementioned steps S601 to S607, some locations describe the situation where an index file, such as Index file B2, is stored in a persistent storage medium. However, this situation is essentially only for the convenience of describing the execution process of a specific step. In fact, the blockchain system supports multiple versions of state data, and it is usually impossible for only one index file to be stored in a persistent storage medium. The relevant description does not constitute a limitation to the present application.
[0108] Based on the same concept as the aforementioned method embodiment, the present specification also provides a blockchain node 700 in a blockchain system. The blockchain system stores multiple data files and multiple index files through a persistent storage medium. The data files store variable values of state variables. The index files store multiple index entries and data file usage information determined based on the multiple index entries. The index entries include the location information of the variable values of the state variables in the persistent storage medium. Figure 7As shown, the blockchain node 700 includes: a prediction processing unit 701, configured to predict the garbage ratio of the multiple data files according to the data file usage information stored in each of the multiple index files; a target determination unit 703, configured to determine the target data file to be recycled from the multiple data files according to the garbage ratio of the multiple data files; a file generation unit 705, configured to select valid data from the target data file and generate a new data file using the valid data; an update processing unit 707, configured to delete the target data file from the persistent storage medium, store the new data file in the persistent storage medium, and update the index entry corresponding to the variable value in the new data file in the multiple index files.
[0109] A computer-readable storage medium is also provided in an embodiment of the present specification, on which a computer program / instruction is stored. When the computer program / instruction is executed in a computer, the computer is caused to execute a method for garbage collection of state data in a blockchain system provided in each of the aforementioned embodiments.
[0110] A computing device is also provided in an embodiment of the present specification, including a memory and a processor, wherein the memory stores a computer program / instruction, and when the processor executes the computer program / instruction, a method for garbage collection of state data in a blockchain system provided in the aforementioned embodiments is implemented.
[0111] In the 1990s, improvements to a technology could be clearly distinguished as hardware improvements (for example, improvements to the circuit structure of diodes, transistors, switches, etc.) or software improvements (improvements to the method flow). However, with the development of technology, many improvements to the method flow today can be regarded as direct improvements to the hardware circuit structure. Designers almost always obtain the corresponding hardware circuit structure by programming the improved method flow into the hardware circuit. Therefore, it cannot be said that an improvement in a method flow cannot be implemented using a hardware entity module. For example, a programmable logic device (PLD) (such as a field programmable gate array (FPGA)) is such an integrated circuit whose logical function is determined by the user's programming of the device. Designers can "integrate" a digital system on a PLD by programming it themselves, without having to ask a chip manufacturer to design and produce a dedicated integrated circuit chip. Moreover, nowadays, instead of manually making integrated circuit chips, this kind of programming is mostly implemented by "logic compiler" software, which is similar to the software compiler used when developing and writing programs, and the original code before compilation must also be written in a specific programming language, which is called hardware description language (HDL). There is not only one HDL, but many kinds, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. The most commonly used ones are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also know that it is only necessary to program the method flow slightly in the above-mentioned hardware description languages and program it into the integrated circuit, and then it is easy to obtain the hardware circuit that implements the logic method flow.
[0112] The controller can be implemented in any appropriate manner, for example, the controller can take the form of a microprocessor or processor and a computer-readable medium storing a computer-readable program code (such as software or firmware) that can be executed by the (micro)processor, a logic gate, a switch, an application-specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art also know that in addition to implementing the controller in a purely computer-readable program code manner, the controller can be implemented in the form of a logic gate, a switch, an application-specific integrated circuit, a programmable logic controller, and an embedded microcontroller by logically programming the method steps. Therefore, this controller can be considered as a hardware component, and the devices included therein for implementing various functions can also be regarded as structures within the hardware component. Or even, the devices for implementing various functions can be regarded as both software modules for implementing the method and structures within the hardware component.
[0113] The systems, devices, modules or units described in the above embodiments may be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a server system. Of course, the present application does not exclude that with the development of computer technology in the future, the computer that implements the functions of the above embodiments may be, for example, a personal computer, a laptop computer, a vehicle-mounted human-computer interaction device, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.
[0114] Although one or more embodiments of the present specification provide method operation steps as described in the embodiments or flow charts, more or less operation steps may be included based on conventional or non-creative means. The order of steps listed in the embodiments is only one way of executing the order of many steps, and does not represent the only execution order. When the device or terminal product in practice is executed, it can be executed in sequence or in parallel according to the method shown in the embodiments or the drawings (for example, a parallel processor or a multi-threaded processing environment, or even a distributed data processing environment). The term "include", "include" or any other variant thereof is intended to cover non-exclusive inclusion, so that the process, method, product or equipment including a series of elements includes not only those elements, but also includes other elements that are not explicitly listed, or also includes elements inherent to such a process, method, product or equipment. In the absence of more restrictions, it is not excluded that there are other identical or equivalent elements in the process, method, product or equipment including the elements. For example, if the words first, second, etc. are used to represent the name, they do not represent any specific order.
[0115] For the convenience of description, the above devices are described in various modules according to their functions. Of course, when implementing one or more of the present specification, the functions of each module can be implemented in the same or more software and / or hardware, or the module implementing the same function can be implemented by a combination of multiple sub-modules or sub-units, etc. The device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0116] The present invention is described with reference to flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0117] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0118] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0119] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0120] The memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.
[0121] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic disk storage, graphene storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.
[0122] It should be understood by those skilled in the art that one or more embodiments of the present specification may be provided as a method, system or computer program product. Therefore, one or more embodiments of the present specification may take the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware. Moreover, one or more embodiments of the present specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0123] One or more embodiments of the present specification may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. One or more embodiments of the present specification may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.
[0124] Each embodiment in this specification is described in a progressive manner, and the same and similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment. In the description of this specification, the description of the reference terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of this specification. In this specification, the schematic representation of the above terms does not necessarily target the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, in the absence of mutual contradiction, a person skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0125] The above description is only an example of one or more embodiments of the present specification and is not intended to limit one or more embodiments of the present specification. For those skilled in the art, one or more embodiments of the present specification may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present specification shall be included in the scope of the claims.
Claims
1. A method for garbage collection of state data in a blockchain system, wherein the blockchain system stores multiple data files and multiple index files through a persistent storage medium, wherein the data files store variable values of state variables, and the index files store multiple index entries and data file usage information determined based on the multiple index entries, wherein the index entries include location information of the variable values of the state variables in the persistent storage medium, and the method comprises: predicting garbage ratios of the plurality of data files according to the data file usage information stored in each of the plurality of index files; Determining a target data file to be recycled from the multiple data files according to the garbage ratios of the multiple data files; Selecting valid data from the target data file, and generating a new data file using the valid data; The target data file is deleted from the persistent storage medium, the new data file is stored in the persistent storage medium, and the index entries corresponding to the variable values in the new data file are updated in the multiple index files.
2. According to the method of claim 1, the location information includes the file identifier of the data file to which the corresponding variable value belongs, the address offset and data length of the corresponding variable value in the data file to which it belongs.
3. According to the method of claim 1, in the blockchain system, state data is managed through a tree structure, a leaf node of the tree structure stores location information of a variable value of a state variable in the persistent storage medium, and a directed path from the root node of the tree structure to the leaf node stores a key of the state variable; in, The index entries include key-value pairs of tree nodes in the tree structure.
4. The method according to claim 3, wherein the tree structure is divided into a plurality of logical pages; in, The method also includes: when the blockchain system generates an incremental page / base page for any logical page of the latest version based on the version update of the status data, the key-value pairs of several tree nodes included in the incremental page / base page are used as several index entries in the index file corresponding to the incremental page / base page; and, based on the position information in the several index entries, the data file usage information in the index file corresponding to the incremental page / base page is determined.
5. According to the method described in claim 4, the data file usage information in the index file includes a number of usage records, and the number of usage records are used to indicate a number of first file identifiers and their respective corresponding data occupancy ratios, the number of first file identifiers are extracted from the location information in the number of index entries, the data occupancy ratio is the ratio of the cumulative data length of each variable value corresponding to a number of target index entries to the file size of the data file indicated by the corresponding first file identifier, and the number of target index entries belong to the number of index entries and include the corresponding first file identifiers.
6. The method according to claim 5, wherein predicting the garbage ratios of the plurality of data files according to the data file usage information in the plurality of index files comprises: For any data file among the multiple data files, several target data occupancy ratios corresponding to the file identifier of the data file are determined from several usage records stored in each of the multiple index files, and the garbage ratio of the data file is determined based on the several target data occupancy ratios.
7. The method according to claim 1, wherein selecting valid data from the target data file comprises: The index entries in the plurality of index files are traversed, and when the position information included in the index entry points to the target data file, the data content in the target data file corresponding to the position information is used as valid data.
8. A blockchain node in a blockchain system, wherein the blockchain system stores multiple data files and multiple index files through a persistent storage medium, wherein the data files store variable values of state variables, and the index files store multiple index entries and data file usage information determined based on the multiple index entries, wherein the index entries include location information of the variable values of the state variables in the persistent storage medium, and the blockchain node includes: A prediction processing unit configured to predict garbage ratios of the plurality of data files according to the data file usage information stored in each of the plurality of index files; a target determination unit configured to determine a target data file to be recycled from the plurality of data files according to the garbage ratios of the plurality of data files; A file generating unit, configured to select valid data from the target data file and generate a new data file using the valid data; The update processing unit is configured to delete the target data file from the persistent storage medium, store the new data file in the persistent storage medium, and update the index entries corresponding to the variable values in the new data file in the multiple index files.
9. A computing device, comprising a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the method according to any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed in a computing device, the computing device executes the method according to any one of claims 1 to 7.