Method and device for accessing data in block chain system
By designing the decoupling of logical paths and physical storage media addresses in the blockchain system, a unified data access interface is provided, and the problem of inconsistency in storage media management and file access mechanisms in the blockchain storage system is solved, and efficient cross-storage level data access is achieved.
Patent Information
- Application Number
- CN202510349553.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-07-11
AI Technical Summary
The differences in interface protocols and file access methods of different storage media in blockchain storage systems lead to inefficient data access and high development and maintenance complexity, making it difficult to achieve a unified file access mechanism.
Design a method to access data in a blockchain system, and provide a unified data access interface through the decoupling of the logical path and the address of the physical storage medium, automatically converting the logical path into a physical path, and achieving efficient access across the storage level.
It reduces the complexity of system development and maintenance, improves the efficiency and convenience of data access, and realizes unsensed cross-level data access to different storage media and levels.
Smart Images

Figure CN120295991A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of this specification belong to the technical field of blockchain, and particularly relate to a method and device for accessing data in a blockchain system. Background Art
[0002] Blockchain is a new application mode of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanism, and encryption algorithms. In a blockchain system, data blocks are combined into a chained data structure in a sequential connection manner according to the time sequence, and a distributed ledger that is tamper-proof and forgery-proof is guaranteed by cryptographic means. Due to the characteristics of blockchain such as decentralization, information immutability, and autonomy, blockchain has received more and more attention and applications.
[0003] All kinds of data generated by the blockchain system are stored in the underlying storage devices, and these storage devices can be storage media such as disks and solid-state drives. With the continuous development of blockchain technology and the continuous enrichment of application scenarios, the amount of data generated by the blockchain system is also increasing continuously, and the storage devices participating in data storage are also increasing accordingly. These storage devices are usually designed with a storage framework with multiple levels (for example, hot data caching, cold data archiving), multiple types (for example, a mix of NVME storage media and SATA storage media), and multiple disks (for example, disk arrays, distributed storage clusters). This storage design can make full use of the advantages of different storage media, for example, the fast random read and write capabilities of caches and solid-state drives, and the large-capacity storage characteristics of mechanical hard disks, so as to improve the performance and reliability of the blockchain storage system.
[0004] However, although a complex blockchain storage system can provide convenience in hierarchical storage of data and save storage costs, it brings certain troubles to data access. Different levels of storage media in the storage system have their own different interface protocols and file access methods. When the blockchain accesses data, it has to set corresponding file access logics for different storage media, resulting in the need for the blockchain application layer (data access party) to frequently switch access interfaces, which not only increases the complexity of development and maintenance, but also affects the data access efficiency to a certain extent.
[0005] Therefore, a technical solution is needed that can uniformly manage various storage media, form a file access mechanism with a unified interface specification, and improve the access convenience of the storage system. Summary of the Invention
[0006] The purpose of the present invention is to provide a method and device for accessing data in a blockchain system, aiming to solve the problems of storage medium management and file access mechanism specification in the blockchain hot and cold hierarchical storage system.
[0007] In a first aspect, a method for accessing data in a blockchain system is provided. The blockchain system stores data through a number of data files, and the number of data files are divided into a number of storage levels according to their respective storage media. The method includes:
[0008] Receiving an access request for target data.
[0009] According to a pre-stored file directory, based on the logical path included in the access request, determining the physical path corresponding to the target data; the physical path includes the target data file corresponding to the target data and the storage level of the target data file.
[0010] In a second aspect, an apparatus for accessing data in a blockchain system is provided. The blockchain system stores data through a number of data files, and the number of data files are divided into a number of storage levels according to their respective storage media. The apparatus includes:
[0011] A receiving unit configured to receive an access request for target data.
[0012] A determining unit configured to, according to a pre-stored file directory, based on the logical path included in the access request, determine the physical path corresponding to the target data; the physical path includes the target data file corresponding to the target data and the storage level of the target data file.
[0013] In a third aspect, a computing device is provided, including a memory and a processor. An executable code is stored in the memory, and when the processor executes the executable code, the method described in the first aspect is implemented.
[0014] In the technical solution provided in the embodiments of this specification, a logical path with a unified file access specification is designed, realizing the decoupling between the logical path of the data file and the physical storage medium address. In the scenario of accessing blockchain data, the logical path can be automatically converted into a physical file path, enabling the data access party to ignore the different data access methods between different storage media and different storage levels, and realizing efficient cross-storage-level data access without perceiving the underlying storage architecture, thereby reducing the complexity of system development and maintenance and improving the efficiency and convenience of data access. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] In order to more clearly illustrate the technical solutions of the embodiments of this specification, the drawings required for describing the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this specification. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0016] Figure 1 Shows the architecture diagram of a blockchain system in an embodiment;
[0017] Figure 2 Is a schematic diagram of the relationship between modules, CPU, memory, and disk involved in the transaction processing process in an embodiment;
[0018] Figure 3 Shows the schematic diagram of the method architecture for accessing data in a blockchain system disclosed in this specification;
[0019] Figure 4 Is a flowchart of a method for accessing data in a blockchain system provided according to an embodiment of this specification;
[0020] Figure 5 Shows the schematic diagram of a storage system architecture disclosed in this specification;
[0021] Figure 6 Is a rule for writing a logical path provided according to an embodiment of this specification;
[0022] Figure 7 Is a device for accessing data in a blockchain system provided according to an embodiment of this specification. Detailed implementation manners
[0023] In order to enable those skilled in the art to better understand the technical solutions in this specification, the following will clearly and completely describe the technical solutions in the embodiments of this specification with reference to the accompanying drawings in the embodiments of this specification. Obviously, the described embodiments are only a part of the embodiments of this specification, rather than all the embodiments. Based on the embodiments in this specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of this specification.
[0024] Figure 1 Is an architecture diagram of a blockchain system provided exemplarily in an embodiment of this specification. The blockchain system may include N blockchain nodes, where Figure 1 Exemplarily shows 8 blockchain nodes such as Node 1 - Node 8. The connections between the nodes schematically represent the connections between the nodes, and the foregoing connections are used to support data transmission between different nodes.
[0025] The blockchain system can provide the function of smart contracts. The smart contracts in the blockchain system are contracts that can be triggered and executed by transactions. The smart contracts can be defined in the form of contract codes. Invoking a smart contract in the blockchain system is to initiate a transaction pointing to the contract address of the smart contract, so that each node in the blockchain system runs the corresponding contract code distributively.
[0026] In various blockchain systems introducing smart contracts, accounts can generally be divided into two types:
[0027] Contract account (CA): mainly used to store the contract code of the corresponding smart contract and the values of the state variables defined in the smart contract, and usually can only be activated by an external account call;
[0028] Externally owned account (EOA): an account registered by an external user in the blockchain system.
[0029] The design of external accounts and contract accounts is actually a mapping from account addresses to account states. The account state of any account usually includes fields such as nonce, balance, storageRoot, and codeHash. Among them, nonce and balance exist in both external accounts and contract accounts, and the codeHash and storageRoot attributes are generally only valid for contract accounts.
[0030] More specifically, for an external account, the value of nonce represents the number of transactions sent from the relevant account address; for a contract account, the value of nonce can represent the number of contracts of the smart contract created by the relevant account address. The value of balance represents the number of a certain digital resource / token owned by the relevant account address. The value of storage root is the hash value of the root node of a tree structure such as a Merkle Patricia Trie (MPT) tree, which is used to organize / manage the storage of the state variables of the relevant contract account. The value of codeHash represents the hash value of the contract code of the relevant smart contract. For an external account, since it does not include a smart contract, the values of the storageRoot and CodeHash fields can generally be an empty string / all-zero string.
[0031] It should be noted that MPT stands for Merkle Patricia Tree, which is a tree structure combining Merkle Tree and Patricia Tree (a more space-saving trie). Among them, the Merkle tree algorithm can calculate a Hash value for multiple transactions respectively, and then connect them in pairs and calculate the Hash again until the top-level Merkle root. In some blockchain systems, an improved MPT tree is usually adopted, such as a 16-ary tree structure, which is usually also simply referred to as the MPT tree.
[0032] The system data that needs to be persistently stored in the blockchain system can be divided into two parts: block data and state data.
[0033] The block data includes one or more blocks in increasing order of block height (or block number). A single block may include a block header and a block body. The block header may include the block hash previous_Hash of the previous block (or called parent hash), timestamp Timestamp, block number BlockNum, state root hash State_Root, transaction root hash Transaction_Root, receipt root hash Receipt_Root, and nonce, etc. The block body may include a transaction set and a receipt set.
[0034] A transaction in the blockchain system refers to a task unit that is executed and recorded in the blockchain system. A single transaction usually includes a sending field (From), a receiving field (To), and a data field (Data). The From field includes the account that initiates the transaction (i.e., the sender account), and the To field may include another account involved / pointed to by the transaction.
[0035] For any Nth block, multiple transactions included in the transaction set belonging to the Nth block can be executed in sequence according to the state data with block number (or version number) N - 1, and the execution results of the multiple transactions can be obtained. Then, the state data with version number N - 1 can be updated according to the execution results of the multiple transactions to obtain the state data with version number N.
[0036] Generally, a tree structure can be used to organize the state data of the blockchain system. The aforementioned tree structure may include, but is not limited to, MPT (Merkle Patricia Tree) and SMT (Sparse Merkle Tree), etc. A leaf node in the tree structure stores the value of a state variable, and the key of a state variable is stored in the directed path from the root node to a leaf node. The data storage system can store the key-value pairs of the tree nodes in the tree structure; for the state variables in the smart contract, the key of the corresponding tree node can be determined based on the contract address of the smart contract and the tree node itself.
[0037] The tree structure for organizing state data may include a state trie. The hash value of the root node of the state trie is stored in State_Root in the block header. The location information of the account status of an external account / contract account in the persistent storage medium is stored in a leaf node of the state trie. In the directed path from the root node to a leaf node of the state trie, the account address of an external account / contract account, or part or all of the hash value calculated based on the account address is stored. As mentioned above, the account status of a single account usually includes fields such as Nonce, Balance, Storageroot, and CodeHash. Nonce and Balance exist in both external accounts and contract accounts, while CodeHash and Storage root are generally only valid for contract accounts.
[0038] In some blockchain systems, the code of the blockchain platform may include a P2P (Peer to Peer) module, a consensus module, an execution module, and a storage module. P2P is a composition method of computer networks. Different from common web networks, P2P is decentralized. The P2P module can complete the distributed dissemination of data. For blockchain nodes, through the P2P module, receipts can be disseminated and received in a peer-to-peer manner. Different participating parties can establish a distributed blockchain network by deploying nodes. The ledger constructed using the chained block structure is stored on each node (or most nodes, such as consensus nodes) in the distributed blockchain network, which is also called a decentralized (or multi-centered) distributed ledger. Such a blockchain system needs to solve the problems of the consistency and correctness of the ledger data on each of the decentralized (or multi-centered) multiple nodes. The same blockchain platform program runs on each node. Under the design with certain fault tolerance requirements, through the consensus module, it can be ensured that all honest nodes have the same transactions, so as to ensure that all honest nodes have the same execution results for the same transactions, and package the transactions and execution results into blocks. Currently, the mainstream consensus mechanisms include: Proof of Work (POW), Proof of Stake (POS), Delegated Proof of Stake (DPOS), Practical Byzantine Fault Tolerance (PBFT) algorithm, HoneyBadgerBFT algorithm, etc. During the consensus process, the consensus module generally can also generate the timestamp of the block corresponding to the current transaction set, etc. The execution module can execute transactions, including ordinary transfer transactions and transactions involving contracts, which can be before or after the consensus module completes consensus. For transactions involving contracts, the execution module can introduce a virtual machine to execute the code of the smart contract, such as the Ethereum Virtual Machine (EVM), so as to shield the differences in the hardware configurations and software environments of each node through the EVM, ensure that the processes and results of executing the smart contract on each node are the same, and avoid the impact of the execution of the smart contract on the blockchain platform code, other programs, or the operating system on the host through the sandbox environment. For a case of a consortium blockchain, through the consensus module, nodes can determine the transaction content and transaction order in a transaction set, and then output a deterministic transaction set of the consensus result to the execution module. The execution module generates execution results by executing ordinary transfer transactions / transactions involving contracts and sends them to the storage module. The storage module can be responsible for storing the execution results in the local persistent storage medium of the node.
[0039] In Figure 2 a blockchain node as shown, physically includes a CPU, memory, disk, etc. In the blockchain platform code executed by this blockchain node, it may include a P2P module, a consensus module, an execution module, and a storage module. The implementation of the functions of the P2P module, the consensus module, and the execution module generally requires the participation of the CPU and memory. The storage module may include a tree construction module, a block header generation module, a WAL (Write Ahead Log) module, and a state database module. Among them, the tree construction module is used to construct a tree (such as an MPT tree) based on the state k-v passed in by the execution module, such as the state trie and storage trie mentioned above, so as to obtain the k-v of the tree nodes, generally requiring the participation of the CPU and memory. The block header generation module is used to generate a block header based on the root node of the tree constructed by the tree construction module and some other data (such as the previous block hash, timestamp, block number, etc.), generally requiring the participation of the CPU and memory. The WAL module is used to persistently store the k-v of the leaf nodes of the tree generated by the tree construction module before writing the k-v of the leaf nodes of the tree generated by the tree construction module into the state database module, to prevent data loss caused by power failure and other situations during the process of writing the k-v of the leaf nodes of the tree generated by the tree construction module into the state database module, and to recover data in case of such a situation, generally requiring the participation of the CPU, memory, and disk. The state database module is used to store the k-v of the tree nodes constructed by the tree construction module on a persistent storage device; since the tree node data will eventually be written into a persistent storage medium (such as the disk in the figure), the state database module generally requires the participation of the disk in addition to the CPU and memory.
[0040] In terms of the storage structure, the above-mentioned Merkle tree structure, such as the MPT of Ethereum and the SMT (Sparse Merkle Tree, similar to MPT) of Libra, can be located in the tree construction module and stored in memory. Among them, the upper-layer Merkle tree is a prefix tree (trie), which can organize data and obtain a unique Merkle root for the organized data. The leaf nodes can save the state Value, and the root node to the intermediate node to the leaf node realizes the lexicographical index of the state key. These tree nodes are encoded as Key according to a certain rule, and their content is encoded as Value, and finally stored in the lower-layer database.
[0041] Most databases adopt NoSQL Key-Value DB (DataBase, also simply referred to as KVDB) with an LSM (Log-Structured Merge-Tree) - like structure, which is located in the state database module and is ultimately saved on disk. Specifically, the database is, for example, levelDB of Ethereum and RocksDB of Libra. Both of these KVDBs are based on the LSM storage engine.
[0042] The LSM storage engine is a hierarchical, ordered, disk - oriented storage engine. It draws on the feature that Logs are continuously appended (instead of modified). Its core idea is to make full use of the fact that disk - based batch sequential writes are far more efficient than random writes, sacrificing some read efficiency in exchange for maximizing write operation efficiency. Generally speaking, the way to maximize the use of disk characteristics is to read or write a fixed - size block of data at one time and minimize random addressing operations as much as possible. The design idea of LSM is based on this disk characteristic and assumes that the memory is sufficient. Instead of writing data to disk every time there is an update, the latest data is first resident in memory. After the data volume accumulates enough, the data in memory is merged with the data on disk using merge sort and then batch - appended to disk.
[0043] With the continuous development of blockchain and the gradual enrichment of application scenarios, the huge amounts of data generated are stored in multiple storage media, and the storage cost and data access cost will also increase rapidly as the data volume grows.
[0044] The hot - cold hierarchical storage of blockchain data designed to balance the blockchain data storage cost and access cost usually divides the storage media into multiple storage levels and evaluates the storage level to which the data should be migrated according to the access characteristics of the data (such as the most recent access timestamp, access frequency per unit time, access interval, etc.). With the continuous influx of new data in the blockchain system and the continuous evolution of business application scenarios, the hot - cold characteristics of the data are constantly changing, thus forming a mobile migration of hot and cold data in the storage system and achieving a dynamic balance between storage cost and access cost.
[0045] However, this adaptive data migration process often has a black box effect on the data access party. That is, the blockchain system / storage system autonomously determines the cold and hot nature of the data according to pre-set rules and completes the data migration. The data access party does not know which storage level and which storage medium the target data to be accessed will be moved to by the storage system. Therefore, when accessing the data, the data access party cannot accurately know the physical storage location / path of the target data and has to traverse and access each storage medium in the storage system, which not only increases the interface call frequency, generates additional I / O overhead and results in low access efficiency, but may even result in the situation where the target data cannot be obtained.
[0046] It should be noted that the data access party mentioned above generally refers to the entity that retrieves and accesses data. According to different object abstraction levels, the data access party can refer to a user, a node of the blockchain system, or an application layer program module, etc., and no specific limitation is made here.
[0047] To solve the above technical problems, the inventor designed a method for accessing data in a blockchain system, which can provide a unified data access interface on a storage system with multiple storage media and improve the convenience of data access. The architecture of this technical solution will be briefly described below.
[0048] Figure 3 As an architecture diagram of a method for accessing data in a blockchain system disclosed in the embodiments of this specification, storage modules / devices with multiple different storage media are deployed in the storage system, and the blockchain data is scattered and stored in these storage media according to certain storage rules. The composition of the storage system and the storage rules will be elaborated in detail below and will not be introduced too much here. In specific practice, a storage system containing multiple storage media can be implemented as a separate system module and named for easy management and access. For example, the storage system can be implemented as the TieredPoolManager module.
[0049] As mentioned above, the data is the system data that needs to be persistently stored in the blockchain system and can be divided into two parts: block data and state data, where the block data can include transaction data and receipt data.
[0050] In a storage system, a file directory records the correspondence between data, data directories and data files, as well as between data files and storage media (hereinafter referred to as data storage information), that is, it records the logical identifiers of each data unit and data file and their corresponding storage location information, including storage levels, identifiers of specific storage media, etc. The data storage information recorded in the file directory can be generated and recorded when data is written, or it can be data that has been persisted in the storage system and is dynamically updated when migrated to other storage media or storage levels according to pre-set rules. For example, when operations such as addition, modification, migration, and deletion occur to data units and data files in the storage system, corresponding updates are made. Regarding the update timing of the data storage information in the file directory, the embodiments of this specification will not list them one by one. It should be noted that the data unit is the basic component unit of data in the storage system. It can be an instantiated data body, a data value, or a data directory, etc., and can be flexibly set according to specific application scenario requirements. The data unit also reflects the smallest data access granularity supported by the storage system.
[0051] During the data access process, the access request needs to carry the logical path corresponding to the target data. The logical path is a virtual addressing identifier designed by the inventor to facilitate data access parties to locate and access data, and is decoupled from the physical storage architecture and the physical address of the storage media. The writing of the logical path needs to follow a unified writing rule, so that even when the data access party does not know the specific storage location of the data, it can accurately obtain the required target data through the logical path.
[0052] In a blockchain system, the physical path corresponding to the actual storage location of the target data can be determined according to the logical path in the access request, so as to achieve accurate location and access to the target data. This process depends on the file directory introduced above. As a bridge between the logical path and the physical path, it realizes the conversion and determination from the logical path to the physical path by retrieving the data storage information recorded therein. And during this process, the data access party only needs to give the logical path of the target data based on the unified writing rule, and the process of resolving and mapping the corresponding physical path and matching the storage medium access protocol is imperceptible to the data access party, thus realizing the decoupling between data access and the physical architecture of the storage system, unifying the data access mechanism, and improving the convenience and flexibility of data access.
[0053] Next, a method for accessing data in a blockchain system will be elaborated in detail in combination with the accompanying drawings and embodiments.
[0054] Following the above technical concept, in Figure 4A method flow for persisting transaction data in a blockchain provided according to an embodiment of this specification is shown. It can be understood that this method can be implemented by any device, equipment, platform, or device cluster with computing and processing capabilities. Refer to Figure 4 , the method at least includes the following steps: S401: Receive an access request for target data. S403: Based on the pre-stored file directory and the logical path included in the access request, determine the physical path corresponding to the target data; the physical path includes the target data file corresponding to the target data and the storage level of the target data file.
[0055] As mentioned above, according to the access characteristics of data, the blockchain system can store data through several data files and divide the several data files into several storage levels according to their respective storage media. For the purpose of ensuring the global performance of the storage system, the design of the storage level needs to consider both the data access cost and the data storage cost at the same time. In other words, frequently accessed data can be stored in data files on storage media with high cost and high random read / write performance, while infrequently accessed data can be stored in data files on storage media with large capacity and low cost characteristics. In some scenarios, historical snapshot data with extremely low access frequency and expired intermediate data can also be archived and stored in data files on cold storage media such as tapes.
[0056] Exemplarily, Figure 5 A schematic diagram of a storage system architecture is given. According to this implementation, there is a negative correlation between the access frequency of data by the blockchain system and the storage level (Tier) of the data file corresponding to the data, that is, the higher the data access frequency, the lower the storage level where it is located; at the same time, there is a negative correlation between the access efficiency of the storage medium and the storage level, that is, the lower the storage level, the higher the access efficiency (read / write performance) of the storage medium used.
[0057] Refer to the appendix Figure 5, according to the access characteristics of the data (e.g., the frequency of access), the data files storing the data can be sequentially divided into multiple data layers. As the access characteristics of the data decrease (e.g., the frequency of access decreases), the storage level of the data file corresponding to the data gradually deepens. For example, the hot data layer (Layer 0), warm data layer (Layer 1), cold data layer (Layer 2), and archive data layer (Layer 3) shown in the figure. To balance the data storage cost, different storage media with different read and write performances can be used for different data layers. As the storage level deepens, the access efficiency of the storage media used in the storage level gradually decreases. As shown in the figure, in Layer 0, a high-performance solid-state drive SSD that supports the NVME specification can be used to store data files; in Layer 1, a storage medium with relatively high performance, such as an SSD using the SATA protocol, can be used; in Layer 2, a traditional mechanical hard disk HDD can be used. Although the read and write performance is not as good as that of the SSD, the HDD has a large capacity and low cost; in Layer 3, low-cost persistent storage media such as tapes, stacked magnetic disks SMR, or archive storage service systems can be used. The read and write efficiency is extremely low, but still supports on-demand recovery for normal data access.
[0058] The data access process based on this storage system will be described in detail below.
[0059] Back to the appendix Figure 4 , in step S401, an access request for the target data is received.
[0060] In this step, according to the specific blockchain data access scenario, the target data can be a data directory. For example, an access request for all data under a certain directory; the target data can also be a data file, such as an access request for a certain transaction data file. Typically, the access request can include a read access request and a write access request; the read access request can be a query of the meta-information of the target data or a read of the content of the target data; the write access request can be a file replacement or addition for the data directory, or an incremental write for the data file.
[0061] Next, in step S403, according to the pre-stored file directory, based on the logical path included in the access request, the physical path corresponding to the target data is determined; the physical path includes the target data file corresponding to the target data and the storage level of the target data file.
[0062] In this step, based on the pre-constructed file directory, the logical path in the access request is converted into the physical path corresponding to the target data. The file directory serves as a mapping table between the logical path and the physical path, and the data storage information recorded therein may include the storage hierarchy configuration of the storage system; the correspondence between data, data directories, and data files, as well as between data files and storage media. By searching the file directory, the data file storing the target data and the storage hierarchy of the data file can be accurately located, thereby completing the conversion of the logical path to the physical path. In this process, the data access party does not need to care about the underlying storage architecture and data storage relationship of the storage system, and only needs to give the logical path of the target data according to the unified writing rules to obtain the physical path that can access the target data. Next, the writing rules of the logical path will be elaborated with examples.
[0063] Figure 6 An exemplary logical path writing rule is shown. According to this embodiment, the logical path can adopt a writing rule in a format similar to URI (Uniform Resource Identifier). Among them, curly braces {·} indicate that specific parameters need to be filled in, and parentheses (·) indicate that the content therein is an optional item.
[0064] According to one implementation, the logical path constructed according to the above writing rules may include a first keyword, and the first keyword is used to define a first storage protocol. For example: storage01: / / ……, where storage01 is the first keyword, indicating that a storage engine supporting the storage01 storage protocol is used to convert the logical path to the physical path to determine the physical path corresponding to the target data. The storage protocol defines the message format for communication between the data access party and the storage engine (different storage protocols may correspond to different writing rules for logical paths). The storage engine is a middleware responsible for specific data storage and access, implements the functions specified by the storage protocol, receives the logical paths it can support, interacts with the storage system, and converts the logical paths into corresponding physical paths. In some application scenarios, multiple storage systems can be instantiated, each storage system has its corresponding storage engine and supported storage protocol, and the blockchain system can flexibly match its corresponding storage engine and storage system according to the first keyword given in the logical path.
[0065] According to one implementation, the logical path constructed according to the above writing rules may include an optional second keyword, which is used to define the access storage level. When the data access party knows the storage level where the target data to be accessed is located, it can indicate to the storage engine through the second keyword to determine the physical address of the target data in the specified storage level. For example: storage01: / / tier:1 / ……, where tier: is the second keyword, which specifies that the access storage level is 1, that is to say, the storage engine that supports the storage01 storage protocol will determine the physical path of the target data in all storage media of the first layer in its corresponding storage system. In other words, the storage level in the physical path corresponding to the target data determined according to this logical path is the access storage level.
[0066] In some scenarios, when the data access party is not sure about the storage level where the target data to be accessed is located, in the logical path, the second keyword can be ignored to indicate that the storage engine determines the physical address of the target data in all levels of its corresponding storage system. For example: storage01: / / / ……, the storage engine will search for the target data according to the file directory in all storage levels of its corresponding storage system to determine its physical path. In other words, in the several storage levels into which the storage medium is divided, retrieve and determine the storage level in the physical path corresponding to the target data.
[0067] Continue to refer to Figure 6 , the relative path in the logical path writing rules is used to specify the relative position of the target data in the storage system, which is the position information relative to the storage level, rather than the absolute path starting from the root directory of the storage system, such as: / data / dg0.log. In this way, without caring about the storage level to which the target data belongs, the data access party can give the relative path of the target data to indicate the storage engine to locate the target data and complete the conversion and determination of the physical path.
[0068] In summary, in practical applications, the data access party can flexibly construct the logical path according to specific requirements and in accordance with the writing rules. Some logical path - physical path examples in application scenarios will be given below to assist in explaining the above method process.
[0069] In an application scenario, the data access party needs to access the data directory / active / stream_1001 (i.e., the relative path of the target data), but is not clear which storage levels of the storage system this data directory exists in, or the data access party needs to access all storage levels that have this data directory. In this application scenario, the data access party does not need to give the second keyword, that is, does not need to specify the storage level, and leaves it to the storage engine to retrieve and determine the physical path corresponding to this data directory among all storage levels by itself; this scenario is usually applicable to scenarios such as parameter setting for batch access to target data. For example, if a default root directory is set on a certain storage system, then the logical path paradigm shown in this scenario can be used to instruct the storage engine to set the target data (i.e., the data directory given in the logical path) as the default root directory on each storage level. The example is as follows:
[0070] Logical path: storage01: / / / active / stream_1001 (storage01 is the first keyword, indicating the use of a storage engine that supports the storage01 storage protocol).
[0071] Physical path: file: / / medium0 / active / stream_1001 (find this data directory in the 0th layer of storage medium), file: / / medium1 / active / stream_1001 (find this data directory in the 1st layer of storage medium) (the file in the physical path indicates that the storage medium corresponding to the target data uses the local file transfer protocol for access).
[0072] In another application scenario, the data access party needs to access the data file
[0073] / active / stream_1001 / data / ss1.log (i.e., the relative path of the target data), and also clearly knows the storage level in the storage system where this data file is stored. In this application scenario, the data access party can directly give the second keyword to specify the access storage level, and leave it to the storage engine to retrieve and determine the physical path corresponding to the target data in this access storage level; this scenario is usually applicable to scenarios such as reading and modifying target data. For example, if a log file is read on a certain storage system, then the logical path paradigm shown in this scenario can be used to instruct the storage engine to read and write the target data file on a specific access storage level. The example is as follows:
[0074] Logical path: storage03: / / tier:2 / active / stream_1001 / data / ss1.log (storage03 is the first keyword, indicating the use of a storage engine that supports the storage protocol of storage03) (tier2 specifies accessing the data file in the second-layer storage medium)
[0075] Physical path: cluster: / / medium2 / active / stream_1001 / data / ss1.log (In the second-layer storage medium, find this data file) (cluster in the physical path indicates that the storage medium corresponding to the target data uses the data cluster protocol for access)
[0076] The storage protocols given in the above scenarios are only for illustrative purposes and do not represent a limitation. In actual use, depending on different storage systems and storage engines, the supported storage protocols vary, and need to be flexibly set according to specific usage scenarios. In addition, the storage hierarchy design given above is also only exemplary. In actual use, more / fewer storage hierarchies can be divided according to data storage needs and the access performance of the storage medium.
[0077] The above is a detailed elaboration of a method flow for accessing data in a blockchain system provided by an embodiment of this specification. Based on this method, decoupling between the data logical path and the physical storage medium address can be achieved. In the scenario of accessing blockchain data, the logical path is automatically converted into a physical file path, enabling the data access party to ignore the different data file access methods between different storage media and different storage hierarchies, and realizing efficient cross-storage-hierarchy data access without having to perceive the underlying storage architecture, thereby reducing the complexity of system development and maintenance and improving the efficiency and convenience of data access.
[0078] Based on the same concept as the foregoing method embodiment, an embodiment of this specification also provides an apparatus for accessing data in a blockchain system. The blockchain system stores data through a number of data files, and the number of data files are divided into a number of storage hierarchies according to their respective storage media. Referring to Figure 7 As shown, the apparatus 700 includes:
[0079] A receiving unit 701, configured to receive an access request for target data.
[0080] A determining unit 702, configured to determine, according to a pre-stored file directory and based on the logical path included in the access request, the physical path corresponding to the target data; the physical path includes the target data file corresponding to the target data and the storage hierarchy of the target data file.
[0081] According to an embodiment of another aspect, the present specification further provides a computing device, including a memory and a processor, characterized in that executable code is stored in the memory, and when the processor executes the executable code, the methods in the foregoing various embodiments are implemented.
[0082] In this specification, the "first" in terms such as the first storage protocol and the first keyword, and the corresponding "second", "third" (if any) in the text are only for the convenience of distinction and description, and do not have any limiting meaning.
[0083] In the 1990s, improvements to a technology could be clearly distinguished as either hardware improvements (e.g., improvements to circuit structures such as diodes, transistors, switches, etc.) or software improvements (improvements to method flows). However, with the development of technology, many method flow improvements today can be regarded as direct improvements to hardware circuit structures. Designers almost always obtain the corresponding hardware circuit structure by programming the improved method flow into the hardware circuit. Therefore, it cannot be said that an improvement to a method flow cannot be implemented using a hardware entity module. For example, a Programmable Logic Device (PLD) (such as a Field Programmable Gate Array (FPGA)) is an integrated circuit whose logic function is determined by the user programming the device. Designers can program themselves to "integrate" a digital system onto a single PLD, without having to ask a chip manufacturer to design and fabricate a dedicated integrated circuit chip. Moreover, nowadays, instead of manually fabricating integrated circuit chips, this programming is mostly implemented using "logic compiler" software, which is similar to the software compilers used in program development and writing. The original code before compilation also has to be written in a specific programming language, which is called a Hardware Description Language (HDL). There is not just one type of HDL, but many types, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. The most commonly used ones currently are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also be aware that by simply performing a little logical programming on the method flow using the above-mentioned several hardware description languages and programming it into an integrated circuit, it is easy to obtain the hardware circuit that implements the logical method flow.
[0084] The controller can be implemented in any suitable manner. For example, the controller can take the form of, for example, a microprocessor or a processor and a computer-readable medium storing computer-readable program code (such as software or firmware) executable by the (micro)processor, logic gates, switches, an application specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller. Examples of the controller include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art also know that in addition to implementing the controller in the form of pure computer-readable program code, it is entirely possible to logically program the method steps to enable the controller to be implemented in the form of logic gates, switches, application specific integrated circuits, programmable logic controllers, embedded microcontrollers, etc. to achieve the same function. Therefore, such a controller can be considered a hardware component, and the devices included therein for implementing various functions can also be regarded as the structures within the hardware component. Or even, the devices for implementing various functions can be regarded as either software modules for implementing the method or the structures within the hardware component.
[0085] The systems, devices, modules, or units illustrated in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a server system. Of course, this application does not exclude that with the development of future computer technologies, the computers implementing the functions of the above embodiments can be, for example, personal computers, laptop computers, in-vehicle human-machine interaction devices, cellular phones, camera phones, smart phones, personal digital assistants, media players, navigation devices, email devices, game consoles, tablet computers, wearable devices, or any combination of these devices.
[0086] Although one or more embodiments of this specification provide method operation steps as described in the embodiments or flowcharts, more or fewer operation steps may be included based on conventional or non-creative means. The order of steps listed in the embodiments is only one way among the execution orders of numerous steps and does not represent the only execution order. When the actual device or terminal product is executing, it may be executed in the order of the method shown in the embodiments or the drawings or in parallel (such as in an environment of parallel processors or multi-threaded processing, or even in a distributed data processing environment). The terms "comprising", "including" or any other variant thereof are intended to cover non-exclusive inclusion, such that a process, method, product or device comprising a series of elements not only includes those elements but also other elements not explicitly listed, or also includes elements inherent to such process, method, product or device. Without further limitation, there is no exclusion of additional identical or equivalent elements in the process, method, product or device comprising the said elements. For example, if terms such as first and second are used to denote names, they do not indicate any specific order.
[0087] For convenience of description, when describing the above device, it is described by dividing it into various modules according to functions. Of course, when implementing one or more of this specification, the functions of each module can be implemented in the same or multiple software and / or hardware, or the modules implementing the same function can be realized by a combination of multiple sub-modules or sub-units, etc. The device embodiments described above are only illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces, and the indirect coupling or communication connection of the device or unit can be in electrical, mechanical or other forms.
[0088] The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be realized by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for realizing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0089] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction device that implements the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 specified in one block or multiple blocks.
[0090] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 specified in one block or multiple blocks.
[0091] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.
[0092] The memory may include non-permanent memory in the computer-readable medium, in the form of random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash memory (flash RAM). Memory is an example of a computer-readable medium.
[0093] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology for information storage. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic disk storage, graphene storage or other magnetic storage devices, or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media, such as modulated data signals and carrier waves.
[0094] Those skilled in the art should understand that one or more embodiments of this specification can be provided as a method, a system, or a computer program product. Therefore, one or more embodiments of this specification can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, one or more embodiments of this specification can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0095] One or more embodiments of this specification can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. One or more embodiments of this specification can also be practiced in a distributed computing environment, where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.
[0096] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the description of the method embodiment. In the description of this specification, the description referring to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of this specification. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0097] The above is only the embodiments of one or more embodiments of this specification and is not used to limit one or more embodiments of this specification. For those skilled in the art, one or more embodiments of this specification can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of this specification shall be included within the scope of the claims.
Claims
1. A method for accessing data in a blockchain system, wherein the blockchain system stores data through a plurality of data files, and the plurality of data files are divided into a plurality of storage levels according to their respective storage media; the method includes: Receiving an access request for target data; Based on a pre-stored file directory and a logical path included in the access request, determining a physical path corresponding to the target data; The physical path includes a target data file corresponding to the target data and a storage level of the target data file.
2. The method according to claim 1, wherein The logical path includes a first keyword for defining a first storage protocol; the determining the physical path corresponding to the target data includes: Using a storage engine that supports the first storage protocol to determine the physical path corresponding to the target data.
3. The method according to claim 1, wherein The logical path includes an optional second keyword for defining an access storage level; The determining the physical path corresponding to the target data includes: If the second keyword is included in the logical path, using the access storage level as the storage level in the physical path; If the second keyword is not included in the logical path, retrieving and determining the storage level in the physical path from the plurality of storage levels.
4. The method according to claim 1, wherein There is a negative correlation between the access frequency of data by the blockchain system and the storage level of the data file corresponding to the data; there is a negative correlation between the access efficiency of the storage medium and the storage level.
5. The method according to claim 1, wherein, The target data includes one of the following: a data directory, a data file.
6. The method according to claim 1, wherein, The storage medium includes one or more of the following: a solid state drive SSD that supports the NVME specification, a solid state drive SSD, a hard disk drive HDD, a shingled magnetic recording SMR, a tape.
7. The method according to claim 1, wherein The access request includes one or more of the following: a read access request, a write access request.
8. The method according to claim 1, wherein The data includes one or more of the following: status data, transaction data, receipt data.
9. An apparatus for accessing data in a blockchain system, wherein the blockchain system stores data through a plurality of data files, and the plurality of data files are divided into a plurality of storage levels according to their respective storage media; the apparatus includes: A receiving unit configured to receive an access request for target data; A determining unit configured to determine a physical path corresponding to the target data based on a pre-stored file directory and a logical path included in the access request; The physical path includes a target data file corresponding to the target data and a storage level of the target data file.
10. A computing device including a memory and a processor, wherein the memory stores executable code, and when the processor executes the executable code, the method according to any one of claims 1-8 is implemented.