Method for upgrading state data structure in block chain system
By freezing old state data files and creating new state data files, combined with hash entanglement technology to fuse the root hash of the new and old versions, the problem of real-time availability and hash anchor relationship breakage of the blockchain system during upgrade is solved, and smooth upgrades and data traceability is achieved.
Patent Information
- Application Number
- CN202510311474.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-06-24
AI Technical Summary
When the blockchain system is upgrading its state data structure, the existing technology needs to be shut down and upgraded, resulting in damage to real-time availability. The incompatibility of the state data structures of the old and new versions leads to a broken hash anchor relationship, affecting the simple payment verification of light nodes.
By freezing the old state data file, creating a new state data file, and fusing the new and old versions of the state Merck tree root hash through hash entanglement in the upgraded block to generate the state root hash of the current block, thereby realizing the integration of the isolated storage of the new and old state data files and the cross-version root hash.
It realizes smooth upgrade of the state data structure of the blockchain system, ensures data consistency, integrity and traceability, improves the flexibility and stability of the upgrade, and avoids the occurrence of hard forks.
Smart Images

Figure CN120200742A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of this specification belong to the technical field of blockchain, and particularly relate to a method for upgrading the state data structure in a blockchain system. Background Art
[0002] Blockchain is a new application mode of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanism, and encryption algorithms. In a blockchain system, data blocks are combined into a chained data structure in a sequential connection manner according to the time sequence, and it is a distributed ledger that is guaranteed to be tamper-proof and non-forgeable by cryptographic means. Due to the characteristics of blockchain such as decentralization, information immutability, and autonomy, blockchain has received more and more attention and applications.
[0003] In the scenarios where blockchain systems are applied, there is a need to upgrade the state data structure. Usually, it is necessary to shut down the blockchain system for upgrading, and this upgrade method will destroy the real-time availability of the blockchain system and increase the probability of upgrade failure. Moreover, the incompatibility between the old and new version state data structures will lead to the breakage of the hash anchoring relationship between the upgraded state data and the old version state data, and light nodes cannot provide simple payment verification for the state data before the upgrade.
[0004] Therefore, a technical solution is needed that can achieve smooth upgrade of the state data structure in the blockchain system, while maintaining the anchoring of the old state data and supporting simple payment verification, so as to improve the stability and flexibility of the blockchain system in the upgrade scenario. Summary of the Invention
[0005] The purpose of the present invention is to provide a method for upgrading the state data structure in a blockchain system, including:
[0006] A method for upgrading the state data structure in a blockchain system, which is applied to a node in the blockchain system, including:
[0007] In response to the achievement of the upgrade condition, upgrade the first state data software to the second state data software.
[0008] Freeze the first state data file recorded based on the first state data software, and create a second state data file; the first state data file includes a first Merkle tree, and the first Merkle tree includes a first root hash; the second state data file is used to store the state data corresponding to the current block and the blocks after the current block.
[0009] Generate a second Merkle tree based on the current block, and the second Merkle tree includes a second root hash.
[0010] Determine the state root hash of the current block according to the second root hash and the historical root hash, where the historical root hash includes at least the first root hash.
[0011] A computing device includes a memory and a processor. Executable code is stored in the memory, and when the processor executes the executable code, the operations in the above method are implemented.
[0012] According to the method provided in the embodiments of the present specification, when the upgrade condition is met, the old state data file can be encapsulated to create a new state data file for storing the state data generated under the new state data software. At the same time, the Merkle tree root hashes of the old and new versions of the state are hashed and entangled to generate the state root hash of the current block. Thus, the isolated storage of the old and new state data files and the cross-version root hash fusion are realized, which can smoothly upgrade the state data structure in the blockchain system, ensure the consistency, integrity and traceability of the data, and improve the upgrade flexibility of the blockchain system. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] To more clearly illustrate the technical solutions of the embodiments of the present specification, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments recorded in the present specification. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0014] Figure 1 Shows the blockchain architecture diagram in one embodiment;
[0015] Figure 2 Shows the schematic diagram of the block storage structure in one embodiment;
[0016] Figure 3 Shows the schematic diagram of the block storage structure in one embodiment;
[0017] Figure 4 Discloses a method architecture diagram for upgrading the state data structure in the blockchain system;
[0018] Figure 5 Is a flowchart of a method for upgrading the state data structure in the blockchain system provided according to the embodiments of the present specification;
[0019] Figure 6 Is a schematic diagram of the method architecture for testing the state data software provided according to the embodiments of the present specification. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0020] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of this specification. Obviously, the described embodiments are only a part of the embodiments of this specification, rather than all the embodiments. Based on the embodiments in this specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of this specification.
[0021] Figure 1 shows a blockchain architecture diagram in an embodiment. In Figure 1 the shown blockchain architecture diagram, the blockchain includes N nodes, Figure 1 and nodes 1 to 8 are schematically shown. The connections between the nodes schematically represent P2P (Peer to Peer) connections, and such connections can be, for example, TCP connections, etc., for transmitting data between the nodes. All ledger data can be stored on these nodes, that is, the states of all blocks and all accounts are stored. Among them, each node in the blockchain can generate the same state in the blockchain by executing the same transactions, and each node in the blockchain can store the same state database.
[0022] Transactions in the blockchain field can refer to task units that are executed and recorded in the blockchain. A transaction usually includes at least a sending field (From), a receiving field (To), and a data field (Data). For example, in the case of a transfer transaction, the From field represents the account address that sends the transaction (i.e., initiates the transfer task to another account), the To field represents the account address that receives the transaction (i.e., receives the transfer), and the transfer amount can be included in the Data field.
[0023] The blockchain can also provide the function of smart contracts. A smart contract on the blockchain is a piece of code that can be triggered and executed by a transaction on the blockchain system, and the smart contract is defined in the form of code. Invoking a smart contract in the blockchain is actually initiating a transaction pointing to the smart contract address, so that each node in the blockchain runs the smart contract code distributively.
[0024] In various blockchain networks that introduce smart contracts, such as Ethereum, accounts usually can include two types:
[0025] Contract Account: Stores the executed smart contract code and the values of the states in the smart contract code, and usually can only be activated by being called by an external account.
[0026] Externally Owned Account: The account of the user, such as the account of the asset token owner.
[0027] The design of external accounts and contract accounts is actually a mapping from account addresses to account states. The state of an account usually includes fields such as nonce, balance, storageRoot, codeHash, etc. Nonce and balance exist in both external accounts and contract accounts; the codeHash and storageRoot attributes are generally only valid for contract accounts.
[0028] nonce: A counter. For an external account, this number can represent the number of transactions sent from the account address; for a contract account, it can be the number of contracts created by the account.
[0029] balance: The number of asset tokens owned by the account address.
[0030] storageRoot: The hash of the root node of an MPT (Merkle Patricia Trie), which organizes the storage of the state variables of the contract account.
[0031] codeHash: The hash value of the smart contract code. For a contract account, it is the hash value of the smart contract; for an external account, since it does not include a smart contract, the codeHash field can generally be an empty string / a string of all zeros.
[0032] MPT stands for Merkle Patricia Tree, which is a tree structure that combines Merkle Tree and Patricia Tree (a more space-saving Trie tree, also known as a dictionary tree). The Merkle Tree algorithm calculates a Hash value for each transaction, then connects them in pairs and calculates the Hash again until the top-level Merkle root. In Ethereum, an improved MPT tree is adopted, such as a 16-way tree structure, which is usually simply referred to as the MPT tree.
[0033] The data structure of the Ethereum MPT tree includes the State Trie. The State Trie contains key-value pairs (also written as key-value, abbreviated as k-v or kv) of the storage content corresponding to each account in the Ethereum network. The "key" in the State Trie can be a 160-bit identifier (such as the address of an Ethereum account or a part of the hash value of the address, hereinafter collectively referred to as the account address), and this account address is distributed in the storage from the root node to the leaf node of the State Trie. The "value" in the State Trie is generated by encoding the information of the Ethereum account (using the Recursive-Length Prefix encoding (RLP) method). As mentioned above, for an external account, the value includes nonce and balance; for a contract account, the value includes nonce, balance, codeHash, and storageRoot.
[0034] Contract accounts are used to store the states related to smart contracts. After a smart contract is deployed on the blockchain, a corresponding contract account will be generated. This contract account generally has some states, which are defined by the state variables in the smart contract and new values are generated when the smart contract is created and executed. The so-called smart contract usually refers to a contract that can automatically execute terms defined in code form in a blockchain environment. Once an event triggers the terms in the contract (meeting the execution conditions), the code can be automatically executed. In the blockchain, the relevant states of the contract are saved in the Storage Trie, and the Hash value of the root node of the Storage Trie is stored in the above-mentioned storageRoot, so that all the states of the contract are locked under the contract account through Hash. The Storage Trie is also an MPT tree structure, storing the key-value mapping from the state address to the state value. The partial information from the root node to the leaf node of the Storage Trie tree is arranged in order to store the address of a state, and the value of the state is stored in this leaf node.
[0035] Such as Figure 2In some of the blockchain data storages shown, the block header of each block includes several fields, such as the previous block hash previous_Hash (Prev Hash in the figure), nonce (in some blockchain systems, this nonce is not random, or in some blockchain systems, the nonce in the block header is not enabled), timestamp, block number Block Num, state root hash State_Root, transaction root hash Transaction_Root, receipt root hash Receipt_Root, etc. Among them, the Prev Hash in the block header of the next block (such as block N + 1) points to the previous block (such as block N), which is the Hash value of the previous block. In this way, the next block is locked to the previous block through the block header on the blockchain. Among them, State_Root, Transaction_Root, and Receipt_Root lock the state set, transaction set, and receipt set respectively. The state set, transaction set, and receipt set organize states, transactions, and receipts in the form of trees respectively. Generally, they can be the same tree structure or different tree structures. For example, in Ethereum, the same MPT structure is adopted. In the tree structure of the state set including smart contracts in some blockchains such as Ethereum, it includes a two-level MPT structure: the leaf nodes of the upper-level MPT structure include two types: external accounts and contract accounts; each contract account among them includes a lower-level MPT structure, and the leaf nodes of the lower level include the values of the states in the contract account.
[0036] Figure 3 is a schematic diagram of the structure of a blockchain data storage. Still taking Ethereum as an example, it can be combined with Figure 2As shown, state_root is the hash value of the root of the MPT tree composed of the states of all accounts in the current block, that is, the State Trie in the form of an MPT tree points to state_root. The root node of this MPT tree is generally an Extension Node or a Branch Node, and what is stored in state_root is generally the Hash value of this root node. The root node can be connected to one or more layers of Extension Node / Branch Nodes below, and these multi-layer tree nodes can be collectively referred to as Internal Nodes. Part of the values in each node from the root node of this MPT to the leaf node can be concatenated in order to form an account address and used as the key, and the account information stored in the leaf node is the value corresponding to this account address. In this way, a key-value pair is formed. This key can also be a part of sha3(Address), that is, a part of the hash value of the account address (the hash algorithm uses the sha3 algorithm for example), and the value it stores can be RLP(Account), that is, the RLP encoding of the account information. Among them, the account information is
[0037] a quadruple composed of [nonce, balance, storageRoot, codeHash]. As mentioned above, for an external account, generally only the nonce and balance items exist, while the storageRoot and codeHash fields default to storing an empty string / all 0 strings. That is to say, the external account does not store contracts, nor the state variables generated after the contract is executed. A contract account generally includes Nonce, Balance, Storage root, CodeHash. Among them, Nonce is the transaction counter of this contract account; Balance is the account balance; Storage root corresponds to another MPT, and the information related to the contract state can be linked through Storage root; CodeHash is the hash value of the contract code. Whether it is an external account or a contract account, its account information is generally located in a separate Leaf Node. From the Extension Node / Branch Node of the root node to the Leaf Node of each account, there may be several Branch Nodes and Extension Nodes in the middle.
[0038] The State Trie can be a tree in the form of an MPT, generally a 16-way tree, that is, each layer can have at most 16 child nodes. For an Extension Node, it is used to store a common prefix and generally has 1 child node, which can be a Branch Node. For a Branch Node, it can have at most 16 child nodes, which may include Extension Nodes and / or Leaf Nodes.
[0039] In a blockchain system, the efficient storage and verification of state data are core technologies to ensure the scalability and security of the blockchain system. Currently, mainstream platforms adopt diverse data structure designs for the management of state data. For example, Ethereum uses a Merkle Patricia Trie (MPT) to implement the storage and verification of account and contract states; Diem (formerly Libra) selects a Sparse Merkle Tree (SMT) to optimize asset data management; while Hyperledger Fabric implements the storage and isolation of state data before and after transaction execution based on a ReadWriteSet.
[0040] These state data structures with different hash generation rules and data organization logics each have their own advantages in different business scenarios. In addition, business elements such as data fields, business attributes, and calculation logics included in the state data structure also need to be designed specifically according to different business scenarios.
[0041] Therefore, in business scenarios where a blockchain system is applied, there is often a need to upgrade the state data structure. Such upgrade requirements may involve changes to the organization form of state data (e.g., conversion of the state tree structure) or changes to the constituent elements of state data (e.g., adding / deleting fields) to adapt to changing and evolving business requirements and system performance optimization metrics. The embodiments of this specification do not give detailed examples of this.
[0042] In some related technologies, the upgrade of the blockchain system state data structure is completed through an offline migration method. During the blockchain system downtime window, the entire old version state data is exported, converted into the new version data structure, and then re-imported. However, this offline migration solution has various problems. For example, the blockchain service is interrupted due to the downtime window. Especially when the scale of state data is huge, the data migration process may take several hours or even days. During this period, the blockchain system is always in a downtime state, seriously affecting business continuity. In addition, performing data conversion during downtime also increases the risk of errors. Once an unexpected error occurs during the data conversion process, it will cause some nodes to fail to upgrade, resulting in nodes running software with two different versions of state data simultaneously in the blockchain system. The incompatibility of the hash generation rules for the old and new version state data structures will also cause nodes that have not been successfully upgraded to be unable to verify the block state root in the new version format, inevitably leading to the occurrence of a hard fork. In terms of data traceability, the incompatibility of the old and new version state data structures will also cause the hash anchoring relationship between the upgraded state data and the old version state data to break, resulting in light nodes being unable to provide simple payment verification to trace the state data before the upgrade.
[0043] To solve the above technical problems, the inventors designed a method for upgrading the state data structure in a blockchain system. When upgrading the state data structure in the blockchain system, the old version state data file can be encapsulated to create a new version state data file for storing the state data generated under the new version state data software. At the same time, in the upgraded block, the Merkle trees corresponding to the old and new version state data are fused by means of root hash confusion, which can still maintain the traceability of the old version state data under the condition of isolating the old and new version state data files, and provide a complete simple payment verification service. Thus, a smooth upgrade of the state data structure in the blockchain system is achieved, improving the flexibility of the upgrade.
[0044] Figure 4 A method architecture for upgrading the state data structure in a blockchain system is provided. In the accompanying drawings, a block in the blockchain system is exemplarily shown. The blocks form a chain structure through the previous hash (Prev Hash). Each block header contains State Root, and the root node hash of the Merkle tree corresponding to the block is stored in this field. The Merkle tree associated with the block is dynamically constructed by the state data software deployed locally on the node. This software defines the organization form of the state data. Exemplarily, configuration information / parameters such as hash calculation rules and state data file storage directories can also be defined in the state data software.
[0045] Typically, a node can generate a state data file based on state data software, which includes a Merkle tree constructed based on the state data. The hash value of the root node of the Merkle tree is written into the State Root field of the block header to form an association. Take the Figure 4 block N in the appendix as an example. When the first state data software is deployed and activated on the node, according to the protocol specified by the software version, a first state data file can be constructed based on the transaction execution result, which includes a first Merkle tree constructed based on the state data according to a preset hash aggregation algorithm and tree structure. The root hash value of the first Merkle tree will be stored in the State Root of block N to complete the association of the state data.
[0046] It should be noted that the data fields included in the block header shown in the accompanying drawings are only exemplary listings. In actual applications, the data fields in the block header can be added or deleted according to specific business needs. The embodiments of this specification do not make specific limitations on this.
[0047] Continue to refer to the appendix Figure 4 When generating block N+1 from block N, after the node triggers an upgrade to the state data structure (the trigger node for the upgrade is only an example, and a detailed description of the upgrade trigger conditions will be given below), the node upgrades the first state data software to the second state data software, freezes the first state data file while creating a second state data file. It should be noted that the freeze operation can refer to setting the data file to read-only, or it can refer to archiving or snapshotting the data file, or other file operations that retain the current data of the first state data file. Specific limitations are not made here.
[0048] The created second state data file is used to save the state data generated based on the second state data software. It is not difficult to see that in this embodiment, with the upgrade of the state data software, the blockchain system does not need to be shut down, and the old version of the state data file does not need to be converted and imported according to the definition of the new version data structure. Next, according to the state data corresponding to the current block (block N+1), a second Merkle tree is constructed. Then, the root hash of the second Merkle tree and the root hash of the first Merkle tree (i.e., the Merkle tree corresponding to block N) are used to determine a mixed hash through hash entanglement (e.g., hash calculation) as the state root hash of the current block. It should be noted that the "state root hash" and "StateRoot" in the block header shown in the accompanying drawings can be two independent data fields, or only the "state root hash" can be saved in the block header. The accompanying drawings do not represent a limitation on the field settings. In specific applications, relevant fields can be stored in different data structures according to actual needs.
[0049] Through the above-described upgrade method, a smooth upgrade of the state data structure in the blockchain system can be achieved, and in the blocks generated based on the state data software of the new version, the Merkle trees of the old and new versions are merged, thereby ensuring data traceability. At the same time, since the conversion operation on the old version data is avoided in this method, the generation of hard forks can be prevented.
[0050] Next, a method for upgrading the state data structure in the blockchain system will be elaborated in detail in combination with the accompanying drawings and embodiments.
[0051] Following the above technical concept, in Figure 5 shows a method flow for upgrading the state data structure in the blockchain system provided according to an embodiment of this specification. It can be understood that this method can be implemented by any device, equipment, platform, or device cluster with computing and processing capabilities. Referring to Figure 5 the method is applied to a node in the blockchain system and at least includes the following steps: S501: In response to the fulfillment of the upgrade condition, upgrade the first state data software to the second state data software. S503: Freeze the first state data file recorded based on the first state data software and create a second state data file; the first state data file includes a first Merkle tree, and the first Merkle tree includes a first root hash; the second state data file is used to store the state data corresponding to the current block and the blocks after the current block. S505: Generate a second Merkle tree based on the current block, and the second Merkle tree includes a second root hash. S507: Determine the state root hash of the current block according to the second root hash and the historical root hash, and the historical root hash at least includes the first root hash.
[0052] Specifically, in step S501, in response to the fulfillment of the upgrade condition, upgrade the first state data software to the second state data software.
[0053] In this embodiment, the upgrade condition can be a time point when the state data software of the new version takes effect pre-set by the administrator of the blockchain system (for example, 3:00 AM on March 13, 2025), or it can be the height of the block when it takes effect (for example, the 10,000th block). The fulfillment of the upgrade condition can be understood as the system time of the node reaching the pre-set effective time point, or a new block with the effective block height is about to be generated on the node. The upgrade condition can be determined by receiving an upgrade transaction sent by the administrator. That is to say, the administrator of the blockchain system can package the upgrade condition in an upgrade transaction and send this transaction to the blockchain system. The upgrade transaction is broadcast to the network and received and verified by each node. For the upgrade transaction verified through consensus, the upgrade condition therein will be stored in each node. The node can continuously detect the system time or block height until the trigger condition for the upgrade is met.
[0054] When the upgrade condition is met, the node can upgrade the first state data software deployed locally on it to the second state data software. Among them, the first state data software can be understood as the old version software, while the second state data software is the new version software. Here, the old and new versions refer to the software effective order of the state data software relative to this upgrade operation, which does not mean that the version number of the new version must be higher than that of the old version: the first state data software represents the version running before the upgrade, and the second state data software is an iterative version that includes data structure optimization or logic function expansion.
[0055] In some examples, the second state data software can be obtained by the node through network download or management node distribution when triggering the upgrade operation. In other examples, the second state data software can also be pre-locally deployed based on the gray release strategy: in other words, the second state data software has passed the local test of the node and is deployed on the node in an inactive state before the upgrade condition is met. In this case, the upgrade of the software version can also be understood as the dynamic switch between the running states of the state data software. For example, the first state data software working in the production environment (active state) and the second state data software working in the pre-release / test environment (inactive state) are both deployed locally on the node; when the upgrade condition is met, the node can perform a hot switch operation, directly remove the first state data software from the production environment (i.e., deactivate the state), and at the same time set / mount the second state data software to the production environment (i.e., set it to the active state), and update the routing policy to make the second state data software take over the subsequent transaction processing, so as to complete the upgrade operation.
[0056] After the above upgrade operation, the node will process the subsequent transaction execution results based on the second state data software. Therefore, in step S503, freeze the first state data file recorded based on the first state data software, and create a second state data file.
[0057] As mentioned above, the first state data file includes the first Merkle tree, and the first Merkle tree includes the first root hash. The frozen first state data file will be permanently retained as a record of the historical state and cannot be modified.
[0058] The second state data file is used to store the state data corresponding to the current block and the blocks after the current block. The newly created second state data file adopts the data organization form defined by the second state data software (which can be the definition of a tree structure or the definition of a data structure), and stores the corresponding state data based on the transaction execution results of the current block and subsequent blocks. Subsequent blocks refer to each block generated on the node before the next upgrade is successfully executed.
[0059] The second state data file can adopt an incremental storage strategy, that is, only append records to the state change data of the current block and subsequent blocks, rather than performing full-scale redundant storage on historical state data. Taking an account model blockchain as an example, if a transaction is executed after the upgrade is completed for a certain account address, then its latest account state will be recorded in the second state data file. This storage strategy can effectively avoid migrating and converting the old version state data in the first state data file, reducing the resource consumption and downtime risks caused by the upgrade operation, and at the same time avoiding the generation of hard forks.
[0060] According to one implementation, the frozen first state data file and the second state data file can be stored in different data directories respectively to avoid read-write conflicts between them. The data directory information / configuration for storing the state data file can be recorded in a smart contract or other accessible data structures in the blockchain system, which is not specifically limited here.
[0061] Next, in steps S505 and S507, a second Merkle tree is generated based on the current block, and the second Merkle tree includes a second root hash. According to the second root hash and the historical root hash, the state root hash of the current block is determined, and the historical root hash includes at least the first root hash.
[0062] As mentioned above, after the upgrade is completed, the node will process the transaction execution result of the current block (such as block N + 1) based on the second state data software, dynamically construct a Merkle tree (i.e., the second Merkle tree) according to the state data, and generate its root hash (i.e., the second root hash). Among them, the second Merkle tree can be as Figure 3 shown to include a state tree and a storage tree corresponding to the contract. The second Merkle tree only includes the state values of several accounts or contract variables written after executing multiple transactions in the current block. That is to say, the second Merkle tree does not include all the state values in the blockchain system. When reading the state value of any object subsequently, it can be read from the second Merkle tree first. If the state value of the object is not stored in the second Merkle tree, it needs to be read from the first Merkle tree.
[0063] In this case, in order to facilitate simple payment verification (SPV verification) of the read status value, a hash entanglement mechanism can be adopted to generate the state root hash of the current block by combining the first root hash of the first Merkle tree and the second root hash of the second Merkle tree, so that the state root hash not only contains the second root hash information, but also has the first root hash information. In one implementation, when the first Merkle tree also does not store all the status values in the blockchain, for example, the state root hash of block N is generated by combining the first root hash of the first Merkle tree and the third root hash of the third Merkle tree corresponding to block N-1, then the first root hash of the first Merkle tree of block N, the second root hash of the second Merkle tree of block N+1, and the third root hash of the third Merkle tree of block N-1 can be combined to generate the state root hash of the current block N+1. In this case, both the first root hash of the first Merkle tree of block N and the third root hash of the third Merkle tree of block N-1 can be referred to as historical root hashes.
[0064] According to one implementation, the hash entanglement mechanism can be to perform a hash calculation on the second root hash and the historical root hash to generate a mixed hash value. For example, Hash(H1, H2), or Hash(H1 + H2), etc. Where H1 represents the historical root hash and H2 represents the second root hash, and the result of the combined hash calculation on the two is used as the state root hash of the current block. The state root hash obtained in this way not only reflects the latest state data after the upgrade, but also maintains the connection with the historical state data.
[0065] In some embodiments, the blockchain system may experience multiple upgrades of the state data format. At this time, the range of the historical root hash needs to include the root hashes of the Merkle trees of historical versions throughout the entire life cycle, that is, the historical root hash can also include: the root hashes of the Merkle trees corresponding to each data file that records the state data before the first data file. Specifically, when generating the state root hash of the current block, it is necessary to extract, on the blockchain system, the root hashes of the Merkle trees of the old version state data files corresponding to each state data structure upgrade that has occurred in history, and perform a hash calculation with the second root hash to generate a mixed hash value. For example, Hash(H1, H2, ……, H i , H i+1 ), or Hash(H1 + H2 + …… + H i + H i+1 ), etc. Where H i represents the root hashes of the Merkle trees of the old version state data files corresponding to each upgrade, and H i+1 represents the second root hash.
[0066] Typically, the Hash(·) algorithm adopted in the hash entanglement mechanism has the characteristic of algebraic homomorphism. For example, Hash(H1 + H2) = Hash(H1) + Hash(H2). When generating the state root hash of the current block, it is only necessary to perform hash entanglement on the second root hash and the state root hash of the previous block to determine the state root hash of the current block.
[0067] The above is a detailed elaboration of the main process of a method for upgrading the state data structure in a blockchain system provided in this specification. By using the method given in the above embodiments, it is possible to achieve a smooth upgrade of the state data structure in the blockchain system by performing a hot switch between state data software of different versions, avoiding downtime. At the same time, in the current block, the Merkle root hash values of the new and old versions are fused, so as to ensure the traceability of the state data. Below, in combination with the embodiments, a method flow for performing simple payment verification after upgrading based on this method will be elaborated.
[0068] In a specific embodiment, in response to a verification request for the target state corresponding to block N + 1, determine the target state data file to which the target state belongs. Refer to Figure 4 , when the target state is stored in the second state data file, an SPV proof for the target state can be performed based on the Merkle path, state root hash, and historical root hash corresponding to the state root hash in the second Merkle tree; or an SPV proof for the target state can be performed based on the Merkle path in the second Merkle tree and the second root hash (State root) of the second Merkle tree.
[0069] When the target state data file is the pre-upgrade version, trigger simple payment verification for the target state based on the Merkle path, associated root hash, and state root hash in the Merkle tree (such as the first Merkle tree) corresponding to the target state data file; the associated root hash includes at least: the second root hash.
[0070] Specifically, when a node receives a verification request for a target state (for example, account balance verification, contract code verification, etc.), it first parses the verification request to extract the key information of the target state (for example, account address, etc.), and accordingly determines the target state data file storing the target state data.
[0071] Next, if the file belongs to the pre-upgrade version, a Merkle proof can be constructed according to the Merkle tree corresponding to the target state data file. A Merkle proof refers to starting from the leaf node corresponding to the target state, along the Merkle tree, backtracking from this leaf node to the root node's Merkle path, and extracting the hash values of each associated node on the path. These hash values of the associated nodes constitute the Merkle proof of the target state.
[0072] Meanwhile, associated root hashes can also be extracted, which include not only the second root hash of the current block after the upgrade, but also, in addition to the target state data file and the second state data file, the root hashes of the Merkle trees corresponding to each data file that records state data. That is to say, if the target state data file is an intermediate version data file in a series of version upgrades, then other root hashes with a version adjacency relationship with the target state data file can be recursively obtained, so as to obtain various calculation parameters defined in the hash entanglement mechanism, which is defined in the second state data software. These associated root hashes, together with the Merkle tree corresponding to the target state data file, jointly constitute a simple payment proof for the target state.
[0073] In addition, after the state data structure in the blockchain system is successfully upgraded, it may also need to be rolled back for various reasons (for example, the performance metrics after the upgrade do not meet expectations, security vulnerabilities are discovered, etc.). Therefore, in one embodiment, a method for rolling back the upgraded second state data software is provided, including:
[0074] In response to the received rollback instruction, determine the rollback block and the corresponding third state data file according to the rollback height included in the rollback instruction. Roll back the second state data software to the state data software corresponding to the rollback block. Store the state data corresponding to the current block and the blocks after the current block in the third state data file.
[0075] The rollback instruction can be a mandatory rollback instruction sent by the administrator to the blockchain system, or a rollback instruction verified and passed by each node in the blockchain system through the consensus mechanism. In response to this rollback instruction, the node can parse the specified rollback height therein (for example, block No. 10000), and accordingly determine the rollback block and the state data file corresponding to this block (the third state data file). If the third state data file is in a frozen state, it can be unfrozen so that it can be modified. Roll back the currently effective second state data software to the state data software corresponding to the rollback block. After this rollback operation is successfully executed, subsequent blocks will generate state data based on the rolled-back state data software, and these state data will also be stored in the third state data file.
[0076] In other embodiments of this specification, a method for gray-box testing of the first state data software and the second state data software is also provided. Before upgrading the state data software, both the first state data software and the second state data software are deployed in the blockchain system. In response to the completion of the execution of a transaction, the first Merkle tree and the second test Merkle tree are updated respectively based on the first / second state data software, and the state root hash of the current block is determined according to the first Merkle tree. The root hashes of the second test Merkle trees of all nodes in the blockchain system are subjected to consistency verification.
[0077] Figure 6 FIG. shows the method architecture for testing the state data software in this embodiment. Taking node 1 in the blockchain system as an example, before the formal upgrade execution, both the first state data software and the second state data software can be deployed locally on the node. Among them, the first state data software is the data software loaded on the blockchain system in the production environment, which is used to generate state data and the corresponding Merkle tree (the first Merkle tree) according to the transaction execution result, and determine the state root hash of the current block according to the root hash of this Merkle tree. The second state data software can be deployed in an isolated sandbox environment, such as a test environment, a pre-production environment, etc. After the confirmation of the execution of each transaction, both the first state data software and the second state data software can generate their respective Merkle trees according to the transaction execution result. Among them, the first Merkle tree is used as the data in the production environment, and the state root hash of the current block is generated according to its root hash. For the second test Merkle tree generated by the second state data software, its root hash can be temporarily stored locally on the node.
[0078] Next, a preset verification mechanism can be adopted to perform consistency verification on the root hashes of the second test Merkle trees of all nodes in the blockchain system. If the verification passes and the number of passes reaches the preset test pass threshold, it can be determined that the second state data software has sufficient robustness and stability and can be upgraded; if the verification fails or the number of failures reaches the preset test failure threshold, it can be determined that there are logical or compatibility problems with the second state data software and it cannot be upgraded, and further debugging and optimization are required.
[0079] After determining that there is a problem with the second state data software, if it is necessary to remove this version of the software from the node, the test data and the second state data software can be removed by polling the node one by one. If necessary, the node can be restarted based on the environment where a single state data software (the first state data software) is deployed.
[0080] According to one implementation, the consistency verification includes at least one of the following: verification based on a consensus mechanism, offline comparison verification, and third-party online comparison. Each will be introduced below.
[0081] Verification based on consensus mechanism is a method that relies on consensus algorithm to verify the second test Merkle tree root hash. The current mainstream consensus mechanisms include: Proof of Work (POW), Proof of Stake (POS), Delegated Proof of Stake (DPOS), Practical Byzantine Fault Tolerance (PBFT) algorithm, Honey Badger Byzantine Fault Tolerance (HoneyBadgerBFT) algorithm, etc. Through the consensus mechanism, the consistency of the root hash calculated on each node can be guaranteed. However, even if the second test Merkle tree root hash of a node does not conform to the consensus mechanism, it will not affect the public results in the production environment, because the first Merkle tree root hash generated by the first state data software is still used in the production environment. For nodes that have not passed the consensus mechanism verification, error information can be output.
[0082] Offline comparison verification, in some application scenarios, it may be necessary to export the second test Merkle tree root hash of each node to a secure offline environment (for example, a local file or a centralized file system) for comparison. This method is usually used in test scenarios that require manual intervention. By directly comparing the root hash values exported by each node, its consistency can be verified.
[0083] Third-party online comparison: deploy a verification node in the blockchain system, call the API of each node to obtain the root hash value of the second test Merkle tree under the specified version, and compare it on the verification node. This method takes advantage of the convenience of the API interface and can automatically obtain and compare the root hash values of each node to ensure their consistency.
[0084] The above verification methods have their own advantages and can meet different testing requirements. In combination with specific practical scenarios, selecting a suitable consistency verification method can effectively test and verify the second state data software.
[0085] The above is a detailed description of a method for upgrading the state data structure in a blockchain system provided in an embodiment of this specification. It should be noted that although the above embodiment mainly uses state data as an example to illustrate the method flow, the technical concept embodied therein can also be applied to the software upgrade execution process of other data in the blockchain system.
[0086] According to an embodiment of another aspect, the present specification also provides a computing device, including a memory and a processor, characterized in that executable code is stored in the memory, and when the processor executes the executable code, the steps of the foregoing method combined with Figure 5 are implemented.
[0087] In this specification, the "first" in terms such as the first data file and the first root hash, and the corresponding "second", "third" (if any) in the text are only for the convenience of distinction and description and do not have any limiting meaning.
[0088] In the 1990s, improvements to a technology could be clearly distinguished as either hardware improvements (e.g., improvements to circuit structures such as diodes, transistors, switches, etc.) or software improvements (improvements to method flows). However, with the development of technology, many method flow improvements today can be regarded as direct improvements to hardware circuit structures. Designers almost always obtain the corresponding hardware circuit structure by programming the improved method flow into the hardware circuit. Therefore, it cannot be said that an improvement to a method flow cannot be implemented using a hardware entity module. For example, a Programmable Logic Device (PLD) (such as a Field Programmable Gate Array (FPGA)) is an integrated circuit whose logic function is determined by the user programming the device. Designers can program themselves to "integrate" a digital system onto a single PLD, without having to ask a chip manufacturer to design and fabricate a dedicated integrated circuit chip. Moreover, nowadays, instead of manually fabricating integrated circuit chips, this programming is mostly implemented using "logic compiler" software, which is similar to the software compilers used in program development and writing. The original code before compilation also has to be written in a specific programming language, which is called a Hardware Description Language (HDL), and there is not just one type of HDL, but many types, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. The most commonly used ones currently are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also be aware that by simply performing a little logical programming on the method flow using the above-mentioned several hardware description languages and programming it into the integrated circuit, it is easy to obtain the hardware circuit that implements the logical method flow.
[0089] The controller can be implemented in any suitable manner. For example, the controller can take the form of, for example, a microprocessor or a processor and a computer-readable medium storing computer-readable program code (such as software or firmware) executable by the (micro)processor, logic gates, switches, an application specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller. Examples of the controller include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art also know that in addition to implementing the controller in the form of pure computer-readable program code, it is entirely possible to logically program the method steps to enable the controller to be implemented in the form of logic gates, switches, application specific integrated circuits, programmable logic controllers, embedded microcontrollers, etc. to achieve the same function. Therefore, such a controller can be considered a hardware component, and the devices included therein for implementing various functions can also be regarded as the structures within the hardware component. Or even, the devices for implementing various functions can be regarded as either software modules for implementing the method or structures within the hardware component.
[0090] The systems, devices, modules, or units illustrated in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a server system. Of course, this application does not exclude that with the development of future computer technologies, the computers implementing the functions of the above embodiments can be, for example, personal computers, laptop computers, in-vehicle human-machine interaction devices, cellular phones, camera phones, smart phones, personal digital assistants, media players, navigation devices, email devices, game consoles, tablet computers, wearable devices, or any combination of these devices.
[0091] Although one or more embodiments of this specification provide method operation steps as described in the embodiments or flowcharts, more or fewer operation steps may be included based on conventional or non-creative means. The order of steps listed in the embodiments is only one way among many orders of step execution and does not represent the only execution order. When the actual device or terminal product is executed, it may be executed in the order of the method shown in the embodiments or the drawings or in parallel (for example, in a parallel processor or multi-threaded processing environment, or even in a distributed data processing environment). The term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, product or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or also includes elements inherent to such process, method, product or device. Without further limitation, it does not exclude the existence of additional identical or equivalent elements in the process, method, product or device comprising the said elements. For example, if terms such as first and second are used to denote names, they do not denote any particular order.
[0092] For convenience of description, the above device is described by dividing it into various modules according to functions. Of course, when implementing one or more of this specification, the functions of each module may be implemented in the same or multiple software and / or hardware, or the modules implementing the same function may be implemented by a combination of multiple sub-modules or sub-units, etc. The device embodiments described above are only illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other may be through some interfaces, and the indirect coupling or communication connection of the device or unit may be in electrical, mechanical or other forms.
[0093] The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, and the combination of processes and / or blocks in the flowchart and / or block diagram can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0094] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction device that implements the functions specified in one or more processes and / or blocks Figure 1 in the process or processes and / or blocks Figure 1 in the block or blocks.
[0095] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operational steps are performed on the computer or other programmable device to produce a computer-implemented process, thereby providing steps for implementing the functions specified in one or more processes and / or blocks Figure 1 in the process or processes and / or blocks Figure 1 in the block or blocks.
[0096] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.
[0097] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM), and / or non-volatile memory such as read-only memory (ROM) or flash memory (flash RAM). Memory is an example of computer-readable media.
[0098] Computer-readable media includes both permanent and non-permanent, removable and non-removable media implemented by any method or technology for storing information. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape disk storage, graphene storage or other magnetic storage devices, or any other non-transmission media that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transitory media such as modulated data signals and carrier waves.
[0099] Those skilled in the art should understand that one or more embodiments of this specification can be provided as a method, a system, or a computer program product. Therefore, one or more embodiments of this specification can take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Moreover, one or more embodiments of this specification can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0100] One or more embodiments of this specification can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. One or more embodiments of this specification can also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.
[0101] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for system embodiments, since they are basically similar to method embodiments, the description is relatively simple, and reference can be made to the relevant parts of the method embodiments for the related content. In the description of this specification, the description of reference terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of this specification. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0102] The above description is only for the embodiments of one or more embodiments of this specification and is not intended to limit one or more embodiments of this specification. For those skilled in the art, one or more embodiments of this specification can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of this specification shall be included within the scope of the claims.
Claims
1. A method for upgrading a state data structure in a blockchain system, the method being applied to a node in the blockchain system, comprising: In response to the upgrade condition being met, upgrading the first state data software to the second state data software; Freeze the first state data file recorded by the first state data software, and create a second state data file; The first state data file includes a first Merkle tree, and the first Merkle tree includes a first root hash; the second state data file is used to store state data corresponding to the current block and the blocks after the current block; Generate a second Merkle tree based on the current block, wherein the second Merkle tree includes a second root hash; Determine a state root hash of the current block according to the second root hash and the historical root hash, wherein the historical root hash includes at least the first root hash.
2. The method according to claim 1, wherein: The historical root hash also includes: Before the first data file, the root hash of the Merkle tree corresponding to each data file of the status data is recorded.
3. The method according to claim 1, wherein: Determining the state root hash of the current block includes: The second root hash is hashed with the historical root hash to generate a mixed hash value as the state root hash.
4. The method according to claim 1, wherein: The upgrade condition is any one of the following: effective time point, effective block height; The upgrade condition is determined by the received upgrade transaction.
5. The method according to claim 1, wherein: The method further comprises: In response to the received rollback instruction, determining a rollback block and a third state data file corresponding to the block according to a rollback height included in the rollback instruction; Rolling back the second state data software to the state data software corresponding to the rollback block; In the third status data file, status data corresponding to the current block and the blocks after the current block are stored.
6. The method according to claim 1, wherein: The method further comprises: In response to a received verification request for a target state corresponding to a current block, determining a target state data file to which the target state belongs; In the case where the target state data file is a version before the upgrade, a simple payment verification for the target state is triggered based on the Merkle tree, the associated root hash and the state root hash corresponding to the target state data file; the associated root hash includes at least: the second root hash.
7. The method according to claim 6, wherein: The associated root hash also includes: In addition to the target state data file and the second state data file, the root hash of the Merkle tree corresponding to each data file of the state data is recorded.
8. The method according to claim 1, wherein: The first state data software and the second state data software are simultaneously deployed in the blockchain system; Before upgrading the first state data software to the second state data software in response to the achievement of the upgrade condition, the method further includes: In response to the completion of the execution of the transaction, the first Merkle tree and the second test Merkle tree are updated based on the first / second state data software respectively, and the state root hash of the current block is determined according to the first Merkle tree; The second test Merkle tree root hash of each node in the blockchain system is verified for consistency.
9. The method according to claim 8, wherein: The consistency verification includes at least one of the following: verification based on a consensus mechanism, offline comparison verification, and third-party online comparison.
10. A computing device, comprising a memory and a processor, wherein the memory stores executable code, and when the processor executes the executable code, the method according to any one of claims 1 to 9 is implemented.
Citation Information
Cited By
Sparse Merkel tree storage and verification method and device, medium and equipment
CN120950508A