Method for persisting transaction data in block chain
By splitting transaction data into sub-data areas in the blockchain system and using composite key storage, the read amplification problem is solved, and the performance and query efficiency of the blockchain system are improved.
Patent Information
- Application Number
- CN202510349307.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-07-11
AI Technical Summary
Existing blockchain systems have read amplification problems in high-frequency fine-grained query scenarios, resulting in performance bottlenecks, especially when retrieving specific transaction data, which requires full reading of block data, resulting in a significant decline in storage system throughput.
The transaction data in the block is split into sub-data areas for independent storage, and the key-value pair format is adopted to build composite keys through block numbers and area indexes to achieve efficient query and reduce the read amplification effect.
By reducing the storage capacity of each sub-data area, the amount of data during reading is reduced, and the overall performance and transaction query efficiency of the blockchain system are improved.
Smart Images

Figure CN120295718A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of this specification belong to the technical field of blockchain, and particularly relate to a method for persisting transaction data in a blockchain. Background Art
[0002] Blockchain is a new application mode of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanism, and cryptographic algorithms. In a blockchain system, data blocks are combined into a chained data structure in a sequential connection manner according to the time sequence, and a distributed ledger that is tamper-proof and non-forgeable is guaranteed by cryptographic means. Due to the characteristics of decentralization, information immutability, and autonomy of the blockchain, the blockchain has received more and more attention and applications.
[0003] In a blockchain system, the transaction data packed into a block and its corresponding receipt data (collectively referred to as transaction data) need to be permanently stored (i.e., persisted) in the underlying database to support historical verification and state tracing. The current mainstream storage solutions usually persist the entire block as an atomic storage unit directly into the database. Although this persistence solution can reduce the read amplification effect by large continuous reads in the scenario of batch pulling of transaction data in units of the whole block, it causes serious performance bottlenecks in the face of high-frequency fine-grained query scenarios. Since a single block contains a large amount of transaction data, when a user only needs to retrieve a specific piece of transaction data, the blockchain system stored by the above solution has to read the entire block data from the database / disk into the memory and then parse it. This "full-scale reading - partial use" mode will cause a significant decrease in the effective throughput of the storage system. With the improvement of the throughput of the blockchain system and the continuous expansion of the block size, the read amplification problem caused by this coarse-grained storage strategy will be exponentially amplified, becoming a bottleneck restricting the scalability of the blockchain system.
[0004] Therefore, a solution is needed to efficiently persist the transaction data in a block with the goal of reducing read amplification, thereby improving the overall performance of the blockchain system. Summary of the Invention
[0005] The purpose of the present invention is to provide a method for persisting transaction data in a blockchain, aiming to comprehensively consider the disk reading mechanism and design the storage of the transaction data of the block, so as to reduce the read amplification problem generated during transaction query. By optimizing the storage structure and query mechanism, the data reading performance of the blockchain system can be improved on the premise of ensuring the integrity and consistency of the transaction data, making it more suitable for processing high-frequency transaction data queries and large-scale transaction data storage requirements.
[0006] In a first aspect, a method for persisting transaction data in a blockchain is provided. The method is applied to a node in a blockchain system and includes:
[0007] Receiving a persistence request for a transaction in a target block.
[0008] Persisting first transaction data in a first sub-data area of a database in a key-value pair format. The first transaction data corresponds to some transactions in the target block. The key of the first sub-data area includes the block number of the target block and the area index of the first sub-data area, and the value of the first sub-data area includes the first transaction data.
[0009] In a second aspect, a method for querying transaction data in a blockchain is provided. The method is applied to a node in a blockchain system and includes:
[0010] Receiving a query request for target transaction data, where the target transaction data is persisted in a database by the method described in the first aspect.
[0011] Determining a target sub-data area in the database to which the target transaction data belongs.
[0012] Querying the target transaction data in the target sub-data area.
[0013] In a third aspect, a computing device is provided, including a memory and a processor. An executable code is stored in the memory. When the processor executes the executable code, the method described in any one of the first aspect and the second aspect is implemented.
[0014] In the technical solution provided in the embodiments of this specification, the transaction data packed in a block can be split into several parts, independently stored in corresponding sub-data areas respectively, and key-value pairs are constructed based on the area index of the sub-data area and the transaction data therein, so that the persistence of the block transaction data can be "broken into pieces", reducing the scale of the data stored per unit, thereby alleviating the read amplification problem and improving the overall performance of the blockchain system. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] To more clearly illustrate the technical solutions in the embodiments of this specification, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings described below are only some embodiments recorded in this specification. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0016] Figure 1 Shows the architecture diagram of a blockchain system in an embodiment;
[0017] Figure 2 It is a schematic diagram showing the relationship between modules, CPU, memory, and disk involved in the transaction processing process in one embodiment;
[0018] Figure 3A It shows a schematic architecture diagram for persisting transaction data in a blockchain disclosed in this specification;
[0019] Figure 3B It shows a schematic architecture diagram for persisting transaction data in a blockchain disclosed in this specification;
[0020] Figure 4 It is a flowchart of a method for persisting transaction data in a blockchain according to an embodiment of this specification;
[0021] Figure 5 It is a flowchart of a method for querying transaction data in a blockchain according to an embodiment of this specification. Detailed implementation manners
[0022] To enable those skilled in the art to better understand the technical solutions in this specification, the following will clearly and completely describe the technical solutions in the embodiments of this specification with reference to the accompanying drawings in the embodiments of this specification. Obviously, the described embodiments are only a part of the embodiments of this specification, rather than all the embodiments. Based on the embodiments in this specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of this specification.
[0023] Figure 1 It is an architecture diagram of a blockchain system exemplarily provided in an embodiment of this specification. The blockchain system may include N blockchain nodes, where Figure 1 8 blockchain nodes such as Node 1 - Node 8 are exemplarily shown. The connections between the nodes schematically represent the connections between the nodes, and the aforementioned connections are used to support data transmission between different nodes.
[0024] The blockchain system can provide the function of smart contracts. The smart contracts in the blockchain system are contracts that can be triggered and executed by transactions. The smart contracts can be defined in the form of contract codes. Invoking a smart contract in the blockchain system is to initiate a transaction pointing to the contract address of the smart contract, enabling each node in the blockchain system to distributively run the corresponding contract code.
[0025] In various blockchain systems introducing smart contracts, accounts can generally be divided into two types:
[0026] Contract account (CA): mainly used to store the contract code of the corresponding smart contract and the values of the state variables defined in the smart contract, and usually can only be activated by being called by an external account;
[0027] Externally owned account (EOA): an account registered by an external user in the blockchain system.
[0028] The design of external accounts and contract accounts is actually a mapping from account addresses to account states. The account state of any account usually includes fields such as nonce, balance, storageRoot, and codeHash. Among them, nonce and balance exist in both external accounts and contract accounts, and the codeHash and storageRoot attributes are generally only valid on contract accounts.
[0029] More specifically, for an external account, the value of nonce represents the number of transactions sent from the relevant account address; for a contract account, the value of nonce can represent the number of smart contracts created by the relevant account address. The value of balance represents the amount of a certain digital resource / token owned by the relevant account address. The value of storage root is the hash value of the root node of a tree structure, such as a Merkle Patricia Tree (MPT) tree, which is used to organize / manage the storage of the state variables of the relevant contract account. The value of codeHash represents the hash value of the contract code of the relevant smart contract. For an external account, since it does not include a smart contract, the values of the storageRoot and CodeHash fields can generally be an empty string / all 0 strings.
[0030] It should be noted that MPT stands for Merkle Patricia Tree, which is a tree structure that combines Merkle Tree and Patricia Tree (a more space-saving trie). Among them, the Merkle tree algorithm can calculate a Hash value for each of multiple transactions, and then connect them in pairs and calculate the Hash again until the top-level Merkle root. In some blockchain systems, an improved MPT tree is usually adopted, such as a 16-ary tree structure, which is usually also simply referred to as an MPT tree.
[0031] The system data that needs to be persistently stored in the blockchain system can be divided into two parts: block data and state data.
[0032] The block data includes one or more blocks in increasing order of block height (or block number). A single block may include a block header and a block body. The block header may include the block hash previous_Hash of the previous block (or called the parent hash), Timestamp, BlockNum, State_Root, Transaction_Root, Receipt_Root, and nonce, etc. The block body may include a transaction set and a receipt set.
[0033] A transaction in the blockchain system refers to a task unit that is executed and recorded in the blockchain system. A single transaction usually includes a sending field (From), a receiving field (To), and a data field (Data). The From field includes the account that initiates the transaction (i.e., the sender account), and the To field may include another account involved / pointed to by the transaction.
[0034] For any Nth block, multiple transactions included in the transaction set of the Nth block can be executed in order based on the state data with block number (or version number) N - 1, and the execution results of the multiple transactions can be obtained. Then, the state data with version number N - 1 can be updated according to the execution results of the multiple transactions to obtain the state data with version number N.
[0035] The state data of the blockchain system can usually be organized in a tree structure. The aforementioned tree structure may include but is not limited to MPT (Merkle Patricia Tree) and SMT (Sparse Merkle Tree), etc. A leaf node in the tree structure stores the value of a state variable, and the key of a state variable is stored in the directed path from the root node to a leaf node. The data storage system can store the key-value pairs of the tree nodes in the tree structure; for the state variables in the smart contract, the key of the corresponding tree node can be determined based on the contract address of the smart contract and the tree node itself.
[0036] The tree structure for organizing state data may include a state trie. The hash value of the root node of the state trie is stored in State_Root in the block header. The location information of the account state of an external account / contract account in the persistent storage medium is stored in a leaf node of the state trie. In the directed path from the root node to a leaf node of the state trie, the account address of an external account / contract account, or part or all of the hash value calculated based on the account address is stored. As mentioned above, the account state of a single account usually includes fields such as Nonce, Balance, Storageroot, CodeHash, etc. Nonce and Balance exist in both external accounts and contract accounts, and CodeHash and Storage root are generally only valid for contract accounts.
[0037] In some blockchain systems, the code of the blockchain platform may include a P2P (Peer to Peer) module, a consensus module, an execution module, and a storage module. P2P is a way of forming a computer network. Different from the common web network, P2P is decentralized. The P2P module can complete the distributed dissemination of data. For blockchain nodes, receipts can be disseminated and received in a peer-to-peer manner through the P2P module. Different participants can establish a distributed blockchain network by deploying nodes. The ledger constructed using the chained block structure is stored on each node (or most nodes, such as consensus nodes) in the distributed blockchain network, which is also called a decentralized (or multi-centered) distributed ledger. Such a blockchain system needs to solve the problems of consistency and correctness of the ledger data on each of the decentralized (or multi-centered) multiple nodes. The same blockchain platform program runs on each node. Under the design with certain fault tolerance requirements, through the consensus module, it can be ensured that all honest nodes have the same transactions, so as to ensure that all honest nodes have the same execution results for the same transactions, and the transactions and execution results are packaged to generate blocks. The current mainstream consensus mechanisms include: Proof of Work (POW), Proof of Stake (POS), Delegated Proof of Stake (DPOS), Practical Byzantine Fault Tolerance (PBFT) algorithm, HoneyBadgerBFT algorithm, etc. During the consensus process, the consensus module generally can also generate the timestamp of the block corresponding to the current transaction set, etc. The execution module can execute transactions, including ordinary transfer transactions and transactions involving contracts, either before or after the consensus module completes the consensus. For transactions involving contracts, the execution module can introduce a virtual machine to execute the code of the smart contract, such as the Ethereum Virtual Machine (EVM), so as to shield the differences in the hardware configurations and software environments of each node through the EVM, ensure that the processes and results of executing the smart contract on each node are the same, and avoid the impact of the execution of the smart contract on the blockchain platform code, other programs, or the operating system on the host through the sandbox environment. For a case of a consortium chain, the nodes can determine the transaction content and transaction order in a transaction set through the consensus module, and then output a deterministic transaction set of the consensus result to the execution module. The execution module generates execution results by executing ordinary transfer transactions / transactions involving contracts and sends them to the storage module. The storage module can be responsible for storing the execution results in the persistent storage medium local to the node.
[0038] In Figure 2 a blockchain node as shown, physically including a CPU, memory, disk, etc. In the blockchain platform code executed by this blockchain node, it may include a P2P module, a consensus module, an execution module, and a storage module. The realization of the functions of the P2P module, the consensus module, and the execution module generally requires the participation of the CPU and memory. The storage module may include a tree construction module, a block header generation module, a WAL (Write Ahead Log) module, and a state database module. Among them, the tree construction module is used to construct a tree (such as an MPT tree) based on the state k-v passed in by the execution module, such as the aforementioned state trie and storage trie, so as to obtain the k-v of the tree nodes, which generally requires the participation of the CPU and memory. The block header generation module is used to generate a block header according to the root node of the tree constructed by the tree construction module and some other data (such as the previous block hash, timestamp, block number, etc.), which generally requires the participation of the CPU and memory. The WAL module is used to persistently store the k-v of the leaf nodes of the tree generated by the tree construction module before writing the k-v of the leaf nodes of the tree generated by the tree construction module into the state database module, to prevent data loss caused by power failure or other situations during the process of writing the k-v of the leaf nodes of the tree generated by the tree construction module into the state database module, and to recover data in case of such a situation, which generally requires the participation of the CPU, memory, and disk. The state database module is used to store the k-v of the tree nodes constructed by the tree construction module on a persistent storage device; since the tree node data will eventually be written into a persistent storage medium (such as the disk in the figure), the state database module generally requires the participation of the disk in addition to the CPU and memory.
[0039] In terms of the storage structure, the above-mentioned Merkle tree structure, such as Ethereum's MPT and Libra's SMT (Sparse Merkle Tree, similar to MPT), can be located in the tree construction module and stored in memory. Among them, the upper-layer Merkle tree is a prefix tree (trie), which can organize data and obtain a unique Merkle root for the organized data. The leaf nodes can store the state Value, and the root node to the intermediate nodes to the leaf nodes implement a lexicographical index for the state key. These tree nodes are encoded as Key according to a certain rule, and their content is encoded as Value, and finally stored in the lower-layer database. Most databases use a NoSQL Key-Value DB (DataBase, also simply referred to as KVDB) with an LSM (Log-Structured Merge-Tree) - like structure, located in the state database module and finally saved on disk. Specifically, the database is, for example, Ethereum's levelDB and Libra's RocksDB. Both of these KVDBs are based on the LSM storage engine.
[0040] The LSM storage engine is a hierarchical, ordered, disk-oriented storage engine. It draws on the characteristic of continuously appending (rather than modifying) logs. The core idea is to make full use of the fact that the sequential write of a disk batch is much more efficient than random writes, and to sacrifice some read efficiency in exchange for maximizing write operation efficiency. Generally speaking, the way to maximize the use of disk characteristics is to read or write a fixed-size block of data at one time and minimize random addressing operations as much as possible. The design idea of LSM is based on this disk characteristic and assumes that the memory is sufficient. Instead of writing the updated data to the disk every time, the latest data is first resident in the memory. After the data volume accumulates enough, the data in the memory is merged with the data on the disk using merge sort and then batch appended to the disk.
[0041] In some related technologies, the entire block is usually directly stored in the database as the smallest storage unit. The choice of using the entire block as the basic unit for Key-Value Pair storage mainly stems from two engineering considerations: Firstly, in databases based on the LSM mechanism such as LevelDB and RocksDB, the storage mode with the block as the storage granularity can reduce the read amplification effect during the block pulling process. For example, when a node needs to pull historical data one by one in the order of block heights, consecutive sequential reads in units of blocks can effectively avoid the multiple disk addressing overheads caused by fragmented data storage and reduce disk I / O. Secondly, since the block data encapsulates strongly correlated data of transactions and receipts, the need to store the correspondence between transaction hashes and receipt hashes additionally can be avoided. In the scenario of querying the execution results of transactions, the complete transaction context can be directly obtained by deserializing the block data, reducing the random read overheads generated by associated queries.
[0042] However, this solution of storing the entire block as a key-value pair will also have performance problems due to its insufficient flexibility in storage and reading. In the scenario of retrieving a specific transaction or verifying a certain receipt, since the key-value pair data is stored in units of the entire block, the entire block data has to be loaded into memory for parsing. Even if the target data to be queried only accounts for one ten-thousandth of the block data, obviously, this will cause a serious read amplification problem. Especially when transaction queries become a high-frequency operation in the scenario, the impact of this read amplification will be more significant, thus reducing the performance and efficiency of the blockchain system.
[0043] Typically, taking the average data size of transactions and receipts as 1KB as an example, when a single block packages 10,000 transactions and their receipts, if the entire block is used as the storage unit for persistence, the data volume of this block in the database will be 10MB. At this time, if a single transaction and receipt (1KB of data) need to be read, the entire 10MB of block data needs to be read out and then searched. In this example, reading transaction data that accounts for 0.01% of the block data size requires consuming 10MB of disk bandwidth, and the read amplification ratio is as high as 10,000 times. Taking this as an example, if the upper limit of the disk read bandwidth is 300MB, then theoretically, the performance of such read operations will be limited within 30 TPS, which obviously cannot meet the requirements of high-performance transaction processing.
[0044] In addition to the above storage strategy of persisting transaction data with blocks as a whole, the inventors found that there is also a common storage strategy in practice: a decentralized strategy of persisting each single transaction data as a storage unit. This strategy encapsulates each transaction and its corresponding receipt into an independent key-value pair (for example, the key is the transaction hash, and the value is the serialized binary tuple of the transaction receipt), and persists it in the database. This storage strategy adopts a completely opposite design idea to the previous strategy. Obviously, in the design of this storage strategy, more attention is paid to the performance of random reading in the scenario of single transaction query. For single transaction query, theoretically, random reading without read amplification effect can be achieved. The target transaction data can be directly located through an exact key-value match without loading redundant information.
[0045] However, the disadvantages of this storage strategy are also very obvious. Firstly, each transaction data is stored as a separate data entry in the database. In order to support basic functions such as block pulling operation and block reconstruction, a complex reverse index mechanism has to be introduced, that is, the association relationship between each block and transaction data is additionally maintained in the database. Secondly, in some scenarios that require batch sequential reading of block data, especially in the scenario of traversing the complete transaction data of a block, the blockchain system applying this storage strategy will have to generate a large number of discrete read requests to read out the transaction data one by one. Such high-frequency random requests and responses will bring at least two negative impacts. On the one hand, it increases the probability of network-level anomalies (such as TCP retransmission, query thread deadlock, etc.). On the other hand, the disk jumping addressing caused by discrete persistence cannot utilize the advantage of sequential reading of the disk, resulting in physical-level disk throughput I / O congestion. These drawbacks will lead to the actual performance of this storage strategy being far lower than expected in some core application scenarios of the blockchain (such as block synchronization).
[0046] Figure 3AThis is a schematic diagram of an architecture for persisting transaction data in a blockchain, corresponding to the storage strategy of persisting block data as a whole described previously. In the embodiment shown in this figure, 3 blocks in the blockchain are listed, as well as the database corresponding to storing the transaction data (transactions and their corresponding receipt data) in the block. In the database, the data is persisted in the form of key-value pairs, shown in the figure as "key => value". Generally, the key of the key-value pair is an integer value, such as uint64 (unsigned 64-bit integer value) in the example, represented in hexadecimal (HEX) conversion as a 16-bit unsigned HEX. In the database, the key of the key-value pair is used to represent the block number, usually with consecutive, dense, and increasing numbers. For example, 0x0d in the figure represents block N-1, 0x0e represents block N, and 0x0f represents block N+1. In practical applications, other numbering rules can also be set for the key, and no specific limitation is made here.
[0047] Continue to refer to the attached Figure 3A Figure. Taking the example where both block N-1 and block N have packed 5000 transactions, and block N+1 has packed 10000 transactions, in the database, the value of the key-value pair corresponding to these blocks is the 5000 / 10000 transactions and their corresponding receipt data of their respective blocks. The transaction data is represented with the prefix TX, and the receipt data is represented with the prefix RCPT. The TX and RCPT data with the same numerical number attached to the prefix are the corresponding transaction and receipt data. It should be noted that the value representation of the key-value pair given in the figure (the one-dimensional array composed of TX and RCPT) is only for example and does not represent a limitation on the data organization form. In specific applications, a suitable data structure can be set according to needs to carry the persistence of transaction data.
[0048] As mentioned previously, in the database that persists transaction data based on the storage strategy shown in Figure 3A Figure, if you need to query a certain transaction or receipt data, such as TX1, you need to first read out the key-value pair of block N-1 from the database (which can be read according to its corresponding key 0x0d). Calculated at 1KB per single transaction data, this time it is necessary to read 5000 transaction data, that is, 5MB of block data. Then, among the values of the read key-value pair, retrieve the 1KB TX1 data, and the remaining transaction data in the 5MB block data are all redundant data for this query, and the read amplification is 5000 times. In the same query scenario, if it is placed in the data of block N+1, the read amplification is 10000 times. That is to say, under this storage strategy, as the number of transactions packed in the block increases, the read amplification effect generated by random queries increases proportionally.
[0049] To solve the above technical problems, the inventor designed a method for persisting transaction data in a blockchain system, which can achieve efficient query of transaction data while maintaining the read and traversal performance of block data. The following will briefly elaborate on the ideological framework of this technical discovery.
[0050] Figure 3B The following is an architecture diagram of a method for persisting transaction data in a blockchain disclosed in an embodiment of this specification. Compared with Figure 3A the data storage structure shown, in Figure 3B the embodiment shown, the inventor split the key key of the key-value pair bit by bit, reserved 2 bits of HEX as the area index of the sub-data area, and the remaining bits are still used to represent the block number. The sub-data area can be understood as a "partition" obtained by dividing the blocks according to certain rules. Each sub-data area independently persists transaction data within a specific range (for example, as shown in the attached figure, each sub-data area stores 1250 transaction data), and anchors the belonging block with the block number, and uses the area index as its unique identifier. As shown in this example, when the width of the area index bit is 8bit (2 bits of HEX), a single block can be divided into at most 256 sub-data areas, and each sub-data area is uniquely identified by the area index in the key key. Taking the attached figure as an example, when block N-1 is divided into 4 sub-data areas, calculating with a single transaction data size of 1KB, each sub-data area can persist 1250KB of transaction data. To retrieve a single transaction data among them, only need to determine the sub-data area through the transaction index in the transaction data, read the 1250KB content of this sub-data area, and retrieve 1KB of transaction data from it. The read amplification factor is reduced from 5000 times to 1250 times. It can be understood that as the number of sub-data areas divided for a single block increases, the size of the transaction data persisted therein decreases, thereby effectively reducing the read amplification effect in the scenario of random transaction queries. There are various rules for dividing the sub-data area, and it can also be selected and applied in combination with specific system characteristics, which will be specifically explained below and will not be introduced in detail here.
[0051] Through the process of the foregoing example, it can be realized in the blockchain to store the transaction data of the block in small batches, and each sub-data area only contains part of the transaction data of the block, rather than the entire block data. In this way, in the scenario of querying a single transaction data, only the sub-data area containing the target query transaction data needs to be read, reducing the data loading amount, thereby achieving the effect of reducing the read amplification effect and improving the performance and response speed of the entire blockchain.
[0052] The following will elaborate in detail on a method for persisting transaction data in a blockchain in combination with the attached figure and the embodiment.
[0053] Following the above technical concept, in Figure 4The figure shows a method flow for persisting transaction data in a blockchain according to an embodiment of this specification. It can be understood that this method can be implemented by any device, equipment, platform, or device cluster with computing and processing capabilities. Refer to Figure 4 The method is applied to a node in a blockchain system and at least includes the following steps: S401: Receive a persistence request for a transaction in a target block. S403: Persist the first transaction data in a key-value pair format in a first sub-data area of the database, where the first transaction data corresponds to some transactions in the target block; the key of the first sub-data area includes the block number of the target block and the area index of the first sub-data area, and the value of the first sub-data area includes the first transaction data.
[0054] Specifically, in step S401, receive a persistence request for a transaction in a target block.
[0055] As mentioned above, nodes in the blockchain can determine the transaction content and transaction order in a transaction set through a consensus module, and then output a deterministic transaction set of the consensus result to an execution module. The execution module generates an execution result by executing ordinary transfer transactions / transactions involving contracts and sends it to a storage module. The storage module can be responsible for storing the execution result in a persistent storage medium local to the node.
[0056] The persistence request is generated for several transactions packaged in a target block that has completed consensus verification and is about to be chained, aiming to permanently store the transaction data in the database. The database in the blockchain system can be a key-value pair database or a streaming database, and no specific limitation is made here.
[0057] Next, in step S403, persist the first transaction data in a key-value pair format in a first sub-data area of the database, where the first transaction data corresponds to some transactions in the target block.
[0058] Generally speaking, along with the execution of a transaction, a receipt corresponding to the transaction is generated. In a blockchain system, when each transaction is executed, in addition to updating the state data and generating a new block, corresponding receipt data will also be generated. The receipt data contains key information about the transaction execution, such as the execution status, consumed Gas, etc. The receipt data corresponds one-to-one with the transaction data, ensuring the transparency and traceability of transactions in the blockchain system. In some implementation manners, the transaction data may include transaction data and its corresponding receipt data.
[0059] As mentioned when elaborating on the technical discovery above, the embodiments of the present invention can reduce the storage capacity of each sub-data area and reduce the read amplification effect by dispersedly storing the transaction data corresponding to transactions.
[0060] In this step, the first transaction data corresponding to some transactions in the target block is stored in the first sub-data area of the database. Taking the key-value pair database as an example, when storing key-value pairs in the first sub-data area, the key "key" can be set as the block number of the target block and the area index of the first sub-data area, and the value "value" can be set as the first transaction data. It should be noted that the setting of the key and value data given here is only an example of its necessary data elements. In specific applications, other data elements can be added and set according to specific needs, and this embodiment does not make any limitations in this regard.
[0061] The sub-data area is a "partition" obtained by dividing the database according to certain rules. That is to say, in the blockchain database, there can be several sub-data areas. Each sub-data area independently persists transaction data within a specific range, anchors to the block it belongs to through the block number, and uses the area index as its unique identifier. In order to embed both the area index and the block number information in the key "key" of the key-value pair, a composite key strategy can be adopted to construct the key: encode the block number and the area index through bit splicing as the key. For example, the block number occupies the high-order part, and the area index occupies the low-order part.
[0062] In different databases, there are different range constraints for the value of the key. Taking a uint64 type number (the value range is from 0 to 2 64 -1) as an example, it contains 64 binary bits. Even if 20 blocks are generated per second, 50 binary bits (the value range is from 0 to 2 50 -1) are used to represent the block number, it is already sufficient for more than 100 years. Therefore, it can be considered to take a specific number of binary bits (for example, the low 8 bits) from uint64 to store the area index of the sub-data area, and the remaining binary bits (for example, the high 56 bits) are used to store the block number. Taking the area index represented by 8 binary bits (i.e., 2 bits HEX) as an example, each block can be divided into at most 256 sub-data areas.
[0063] The key of the key-value pair constructed based on the above composite key strategy (block number + area index) not only reasonably utilizes the bit width of the key but also provides a flexible identification space for the sub-data area. By combining the block number and the area index, the sub-data area can be efficiently managed and retrieved in the database. For example, when the area index in the key "key" is given an associated mapping relationship with the transaction index of the transaction data, the discrete transaction data originally scattered in the block can be arranged in a standardized manner, laying a foundation for subsequent efficient queries.
[0064] According to one implementation, the area indexes of the several sub-data areas and the transaction indexes of the transaction data satisfy the monotonicity rule. Specifically, there is an increasing or decreasing correspondence between the area index and the transaction index, that is, as the value of the area index increases, the transaction index also shows an increasing trend, and vice versa. This monotonicity rule can ensure the ordered distribution of transaction data in the sub-data areas, facilitating the rapid positioning and retrieval of specific transaction data. For example, the sub-data areas can be divided according to the value range of the transaction index. Each sub-data area is responsible for storing transaction data within a certain range, and the transaction index value of the latter sub-data area is greater than that of the previous sub-data area. In this way, when querying transaction data, the sub-data area where it is located can be directly determined according to the transaction index, and then the corresponding data can be quickly read, effectively improving the query efficiency.
[0065] Typically, in a specific example, the area indexes of the several sub-data areas can be continuously numbered based on the monotonicity rule. In this way, a strict linear mapping relationship can be formed between the transaction indexes of the block transaction data and the area indexes of the sub-data areas. For example, if the capacity of a single sub-data area is set to 1250 transaction data, then the sub-data area with area index 1 stores the transaction data with transaction indexes 1 - 1250; the sub-data area with area index 2 stores the transaction data with transaction indexes 1251 - 2500, and so on; the sub-data area with area index n stores the transaction data with transaction indexes from (n - 1)*1250 + 1 to min(n*1250, total number of transactions in the block). This linear mapping relationship can be abstracted as: area index = ceiling(transaction index / sub-data area capacity), where ceiling(·) is the ceiling function. Based on this linear mapping relationship, in the scenario of querying transaction data, the sub-data area where it is located can be directly determined according to the transaction index through simple arithmetic operations, achieving a faster data query.
[0066] The sub-data areas designed through the above implementation methods can avoid a serious read amplification effect caused by loading too much transaction data at one time during the reading process of transaction data by reducing the storage capacity of each sub-data area. Further, since the keys of all sub-data areas have an increasing property (the block number increases, and the area indexes within the same block number increase), the sub-data areas can be iterated through the iteration interface of the database (such as RocksDB, LevelDB, etc.) to achieve the purpose of efficiently obtaining all transaction data within a certain range.
[0067] The above is a detailed description of the area index definition and transaction data storage of the sub-data area. Through the previous introduction, it can be understood that the division of the sub-data area is essentially to find a balance between the efficiency of sequential batch loading and random query of transaction data: a larger sub-data area capacity can reduce storage fragmentation and improve the sequential I / O throughput when reading all block data, but it will cause a certain amount of data overload (read amplification) during single transaction retrieval; while an overly divided sub-data area can effectively reduce the read amplification effect during random query, but it will increase the metadata overhead of key-value storage (for example, occupying more key digits to express the area index), resulting in more fragmented data storage.
[0068] Therefore, it is necessary to divide the sub-data area according to appropriate rules according to the specific application scenario. The sub-data area division rules in multiple embodiments will be further exemplified below.
[0069] According to one implementation, the capacity of the sub-data area can be a divisor of the number of transactions in the target block. In other words, the capacity of the sub-data area can be determined according to the N-equal division of the number of transactions in the target block. By designing the sub-data area capacity as an integer divisor of the total number of transactions in the target block, uniform cutting of the transaction data within the block can be achieved, and at the same time, it has a certain degree of flexibility, and the value of N can be adjusted according to actual needs to adapt to different transaction volumes and performance requirements. This uniform division strategy can make each sub-data area contain an equal amount of transaction data, and fast and accurate area index positioning can be achieved when calculating the area index through the transaction index. It should be noted that the N-equal division given in this embodiment is only used to illustrate the equality of the upper limit of the sub-data area capacity, and does not represent the divisor of the equal division nature of the number of transactions in the target block. When the number of block transactions cannot be divided evenly, the nearest divisor strategy can be adopted to select the equal division divisor closest to the ideal partition capacity. For example, if the target block packages 5001 transactions and the ideal partition capacity is 1250 transaction data, 5 sub-data areas can be divided, with the first 4 each storing 1250, and the 5th sub-data area storing 1 transaction data.
[0070] According to another embodiment, the capacity of the sub-data area can be a multiple of the physical sector size of the database storage medium to comply with the reading principle of the disk sector, further reduce read amplification, and reduce system load. For example, if the physical sector size of the storage medium used by the database is 4KB, the capacity of the sub-data area can be set to an integer multiple of 4KB (e.g., 8KB, 16KB, etc.). In this way, when reading and writing data, the data in each sub-data area can be aligned with the physical sector to avoid storage fragmentation caused by spanning multiple sectors. At the same time, this alignment method helps to improve the cache hit rate and reduce data loading time when reading data, making full use of the physical characteristics of the storage medium, reducing the seek and loading delays in I / O operations, reducing the read amplification effect, and improving data reading and writing efficiency.
[0071] According to another embodiment, when a streaming database is used as the persistence medium of the blockchain, the system will pay more attention to the window processing of data (that is, when receiving a continuous data stream, the data in a certain range is divided by windowing for persistence). In this scenario, the capacity of the sub-data area can be a preset fixed capacity to match the underlying characteristics of streaming data window processing. Based on the fixed-capacity sub-data area design, when receiving a continuous transaction data stream, batch persistence of transaction data can be performed in units of windows: when the accumulated transaction data in the memory buffer reaches a preset fixed capacity (for example, 64KB), the entire block of data is immediately triggered to be dropped to disk (persistent) and the area index is incremented to receive the transaction data of the next window.
[0072] In the above-mentioned embodiment of the sub-data area division, when the sub-data area capacity is set according to the data size, there will no longer be a unique mapping relationship between the transaction data stored in the sub-data area and the area index. This is because the size of the transaction data is uncertain, resulting in different numbers of transaction data stored in each sub-data area. For example, a sub-data area with the same capacity may store 20 small-volume transaction data, or it may only store 3 large-volume transaction data. Therefore, in the query scenario, in order to be able to quickly determine the sub-data area to which the transaction data belongs based on the transaction data, an additional index table can be stored in the database, which contains the transaction data and the area index information of the sub-data area to which it belongs, for fast data retrieval. In other words, after the first transaction data is persisted in the first sub-data area in the format of a key-value pair, the correspondence between the first transaction data and the area index of the first sub-data area can be recorded in the database.
[0073] The above is a detailed description of a method for persisting transaction data in a blockchain. Based on this method, the persistence of block transaction data can be "broken down into parts", reducing the scale of data stored per unit, thereby alleviating the read amplification problem, improving the overall performance of the blockchain system, and enhancing the query efficiency of transactions. Based on this, in some other embodiments of this specification, a method for querying transaction data in a blockchain is also provided. Figure 5 The flowchart of the method is shown, and this method is applied to nodes in a blockchain system, specifically including:
[0074] After receiving a query request for target transaction data (step S501), first parse the query request to extract key information of the target transaction data, such as transaction hash, transaction index, etc.
[0075] Determine the target sub-data area to which the target transaction data belongs according to the extracted key information (step S503). As mentioned before, when storing transaction data, either the transaction data is stored in the corresponding sub-data area according to a certain linear mapping relationship, or the index relationship information between the transaction data and the sub-data area is written into the index table. Therefore, after extracting the key information of the target transaction data, the target sub-data area to which the target transaction data belongs can be quickly determined based on the preset mapping relationship or by querying the index table, avoiding blind search in the full block data.
[0076] After that, query the target transaction data in the target sub-data area (step S505). Since only the transaction data corresponding to some block transactions is stored in the sub-data area, the I / O throughput required to read the target sub-data area from the database will be lower than that of reading the entire block data. Thus, not only can unnecessary data loading be reduced, but also the data query range can be narrowed, and the target transaction data can be quickly queried at a smaller time cost and space cost. Through this query method, the query efficiency can be improved while ensuring data integrity.
[0077] According to an embodiment of another aspect, this specification also provides a computing device, including a memory and a processor, characterized in that executable code is stored in the memory, and when the processor executes the executable code, the steps of the foregoing method are implemented in combination with Figure 4 、 Figure 5 the method.
[0078] In this specification, the "first" in terms such as the first sub-data area and the first transaction data, and the corresponding "second", "third" (if any) in the text are only for the convenience of distinction and description, and do not have any limiting meaning.
[0079] In the 1990s, improvements to a technology could be clearly distinguished as either hardware improvements (e.g., improvements to circuit structures such as diodes, transistors, switches, etc.) or software improvements (improvements to method flows). However, with the development of technology, many method flow improvements today can be regarded as direct improvements to hardware circuit structures. Designers almost always obtain the corresponding hardware circuit structure by programming the improved method flow into the hardware circuit. Therefore, it cannot be said that an improvement to a method flow cannot be implemented using a hardware entity module. For example, a Programmable Logic Device (PLD) (e.g., a Field Programmable Gate Array (FPGA)) is such an integrated circuit whose logical function is determined by the user programming the device. Designers can program themselves to "integrate" a digital system onto a single PLD without having to ask a chip manufacturer to design and fabricate a dedicated integrated circuit chip. Moreover, nowadays, instead of manually fabricating integrated circuit chips, this programming is mostly implemented using "logic compiler" software, which is similar to the software compilers used in program development and writing. The original code before compilation also has to be written in a specific programming language, which is called a Hardware Description Language (HDL), and there is not just one type of HDL, but many, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. The most commonly used ones currently are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also be aware that by simply performing a little logical programming on the method flow using the above-mentioned several hardware description languages and programming it into an integrated circuit, it is easy to obtain the hardware circuit that implements the logical method flow.
[0080] The controller can be implemented in any suitable manner. For example, the controller can take the form of, for example, a microprocessor or a processor and a computer-readable medium storing computer-readable program code (such as software or firmware) executable by the (micro)processor, logic gates, switches, an application specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller. Examples of the controller include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art also know that in addition to implementing the controller in the form of pure computer-readable program code, it is entirely possible to logically program the method steps to enable the controller to implement the same functions in the form of logic gates, switches, application specific integrated circuits, programmable logic controllers, and embedded microcontrollers. Therefore, such a controller can be considered a hardware component, and the devices included therein for implementing various functions can also be regarded as the structures within the hardware component. Or even, the devices for implementing various functions can be regarded as either software modules for implementing the method or the structures within the hardware component.
[0081] The systems, devices, modules, or units illustrated in the above embodiments can be specifically implemented by a computer chip or an entity, or by a product with certain functions. A typical implementation device is a server system. Of course, this application does not exclude that with the development of future computer technologies, the computers implementing the functions of the above embodiments can be, for example, personal computers, laptop computers, in-vehicle human-machine interaction devices, cellular phones, camera phones, smart phones, personal digital assistants, media players, navigation devices, email devices, game consoles, tablet computers, wearable devices, or any combination of these devices.
[0082] Although one or more embodiments of this specification provide method operation steps as described in the embodiments or flowcharts, more or fewer operation steps may be included based on conventional or non-creative means. The order of steps listed in the embodiments is only one way among the execution orders of numerous steps and does not represent the only execution order. When the actual device or terminal product is executed, it may be executed in the order of the method shown in the embodiments or the drawings or in parallel (for example, in an environment of parallel processors or multi-threaded processing, or even in a distributed data processing environment). The term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, product or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or also includes elements inherent to such process, method, product or device. Without further limitation, there is no exclusion of additional identical or equivalent elements in the process, method, product or device comprising the said elements. For example, if terms such as first and second are used to denote names, they do not denote any particular order.
[0083] For convenience of description, when describing the above device, it is divided into various modules according to functions for separate description. Of course, when implementing one or more of this specification, the functions of each module may be implemented in the same or multiple software and / or hardware, or the modules implementing the same function may be realized by a combination of multiple sub-modules or sub-units, etc. The device embodiments described above are only illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other may be through some interfaces, and the indirect coupling or communication connection of the device or unit may be in electrical, mechanical or other forms.
[0084] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, and the combination of processes and / or blocks in the flowcharts and / or block diagrams, can be realized by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate a device for realizing the function specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0085] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction device that implements the functions specified in one process Figure 1 one process or multiple processes and / or boxes Figure 1 specified in one box or multiple boxes.
[0086] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one process Figure 1 one process or multiple processes and / or boxes Figure 1 specified in one box or multiple boxes.
[0087] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.
[0088] The memory may include non-permanent memory in the computer-readable medium, in the form of random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash memory (flash RAM). Memory is an example of a computer-readable medium.
[0089] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology for information storage. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic disk storage, graphene storage or other magnetic storage devices, or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media, such as modulated data signals and carrier waves.
[0090] Those skilled in the art should understand that one or more embodiments of this specification can be provided as a method, a system, or a computer program product. Therefore, one or more embodiments of this specification can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, one or more embodiments of this specification can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0091] One or more embodiments of this specification can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. One or more embodiments of this specification can also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.
[0092] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for system embodiments, since they are basically similar to method embodiments, the description is relatively simple, and reference can be made to the relevant parts of the method embodiments for the related content. In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of this specification. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0093] The above is only the embodiment of one or more embodiments of this specification and is not used to limit one or more embodiments of this specification. For those skilled in the art, one or more embodiments of this specification can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of this specification shall be included within the scope of the claims.
Claims
1. A method for persisting transaction data in a blockchain, the method being applied to a node in a blockchain system, comprising: Receiving a persistence request for a transaction in a target block; Persisting first transaction data in a first sub-data area of a database in a key-value pair format, the first transaction data corresponding to some transactions in the target block; The key of the first sub-data area includes the block number of the target block and the area index of the first sub-data area, and the value of the first sub-data area includes the first transaction data.
2. The method according to claim 1, wherein, The transaction data includes: transaction data and its corresponding receipt data.
3. The method according to claim 1, wherein The database includes a plurality of sub-data areas, and the area indexes of the plurality of sub-data areas and the transaction indexes of the transaction data satisfy a monotonicity rule.
4. The method according to claim 3, wherein, The area indexes of the plurality of sub-data areas are continuously numbered based on the monotonicity rule.
5. The method according to claim 1, wherein The capacity of the sub-data area is a divisor of the number of transactions in the target block.
6. The method according to claim 1, wherein The capacity of the sub-data area is a multiple of the physical sector size of the database storage medium.
7. The method according to claim 1, wherein, The capacity of the sub-data area is a preset fixed capacity.
8. The method according to claim 6 or 7, wherein, After persisting the first transaction data in the first sub-data area of the database in a key-value pair format, the method further includes: Recording the correspondence between the first transaction data and the area index of the first sub-data area in the database.
9. A method for querying transaction data in a blockchain, the method being applied to a node in a blockchain system, comprising: Receiving a query request for target transaction data, the target transaction data being persisted in a database by the method according to any one of claims 1-8; Determining the target sub-data area to which the target transaction data belongs in the database; Querying the target transaction data in the target sub-data area.
10. A computing device, comprising a memory and a processor, wherein the memory stores executable code, and when the processor executes the executable code, the method according to any one of claims 1-9 is implemented.