Verifiable ledger database storage extension method and system

By introducing shard expansion protocol and synchronization control protocol, the locality and coding strategy of the storage engine are optimized, and the scalability bottlenecks and insufficient storage performance optimization of verified ledger database systems under data growth are solved, and efficient storage and synchronization performance and reliability of cross-shash transactions are achieved.

WO2025111750A1PCT designated stage expired Publication Date: 2025-06-05SHENZHEN INST OF ADVANCED TECH CHINESE ACAD OF SCI

Patent Information

Application Number
PCT/CN2023/134393
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-11-27
Publication Date
2025-06-05

AI Technical Summary

Technical Problem

Existing verifiable ledger database systems face scalability bottlenecks in the case of data growth, insufficient storage performance optimization, and it is difficult to balance the storage and computing overhead caused by security requirements.

Method used

The shard expansion protocol and synchronization control protocol are introduced, transparent services are provided through the abstract database layer, the storage modules of the blockchain system are adapted to the storage engine, the index data localization and coding strategy are optimized, and the cross-chip transaction snapshot format is used to define the cross-chip transaction snapshot format, realizing read-write separation and concurrent lock-free operations.

Benefits of technology

It improves the storage and synchronization performance of the verifiable ledger database, ensures the reliability of cross-shash transaction services, realizes collaborative optimization of the data layer and the storage layer, adapts to different distributed consensus algorithms, and improves the scalability and performance of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2023134393_05062025_PF_FP_ABST
    Figure CN2023134393_05062025_PF_FP_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of computers, and in particular to a verifiable ledger database storage extension method and system, capable of solving, to a certain extent, the problems of how to improve authenticated data structures to improve query efficiency, and balancing storage and calculation overheads caused by security requirements. The method comprises: introducing a sharding extension protocol, wherein the sharding extension protocol abstracts a database layer interface to provide transparent service for a lower-layer application; introducing a synchronization control protocol, wherein the synchronization control protocol improves a storage engine interface to implement two-tier storage with read-write separation, and defines a cross-shard transaction snapshot format on the basis of a Merkle tree and a derivative verifiable index structure; and designing a storage engine to optimize the locality of index data.
Need to check novelty before this filing date? Find Prior Art

Description

A method and system for storage expansion of verifiable ledger database Technical Field

[0001] The present application relates to the field of computer technology, and more specifically, to a method and system for expanding storage of a verifiable ledger database. Background Art

[0002] Verifiable ledger databases are a direct result of the innovative applications and engineering practices of blockchain systems in data management and storage. Blockchain systems leverage distributed ledger technology, bringing new features to traditional data services and opening up new application scenarios for trusted ledger data services, targeting scenarios requiring data traceability and mutual authentication, such as finance, healthcare, and the Internet of Things. However, industrial applications require more robust technical solutions tailored to specific scenarios. Limited by performance and scalability, blockchain systems are not yet sufficient to completely revolutionize traditional data storage systems. Traditional databases, distributed computing, and cloud service architectures remain the primary technological drivers, giving rise to hybrid blockchain-database data service solutions and independent verifiable ledger database systems.

[0003] Blockchain systems enable decentralized data management from networking to application. Multiple nodes in the network autonomously and synchronously participate in data management and verification, enabling independent verification of data correctness and consistency even in the face of malicious attacks and single points of failure. Verifiable ledger databases are abstracted from the blockchain storage layer. The monotonically increasing ledger data storage requirements further promote the ledger model as a trusted storage proof for interactive applications in scenarios where data computation and storage are separated. These models provide data verification and auditing features, and as independent storage components, they can be customized to provide more flexible and efficient services for different application scenarios. Both of these systems face scalability bottlenecks. As ledger data grows, storage architecture design becomes a major challenge in system expansion, and storage performance is a key optimization point.

[0004] Blockchain defines a basic block-chain data structure, including cryptographic techniques and distributed consensus mechanisms. As the foundation of distributed ledger technology, it enables data synchronization and consistency verification across distributed networks. The ledger data services provided by the hybrid blockchain database architecture must implement: 1) a verifiable query interface, enabling clients to verify the validity of query results from untrusted nodes. This requires storage nodes to traverse the entire blockchain ledger to execute queries; and 2) a distributed consensus interface, ensuring consistency in data synchronization mechanisms across logical units in the distributed network. Therefore, a verifiable ledger database typically consists of two components: an abstract database layer and a ledger storage engine. These components provide transparent ledger data storage and query interfaces to various upper-layer applications, requiring the processing and persistence of large amounts of data.

[0005] Currently, the main blockchain database hybrid architecture systems are difficult to compare and evaluate horizontally due to the lack of uniformity in the abstract database layer. This also means that they are difficult to apply to heterogeneous computing across systems. The verifiable ledger database will combine the abstract database layer and the storage engine layer in the expansion plan to achieve linear expansion and availability assurance. In general, it includes off-chain storage expansion and sharded storage expansion.

[0006] However, there is less attention paid to improving the authenticatable data structure to improve query efficiency and balance the storage and computing overhead brought by security requirements, while there is less attention paid to the scalability of the underlying storage protocol and the performance optimization of the storage engine.

[0007] Summary of the Invention

[0008] In order to solve the problem of how to improve the authenticatable data structure to improve query efficiency and balance the storage and computing overhead brought by security requirements, while paying less attention to the scalability of the underlying storage protocol and storage engine performance optimization, this application provides a verifiable ledger database storage expansion method and system.

[0009] The embodiment of the present application is implemented as follows:

[0010] In a first aspect, the present application provides a method for expanding storage of a verifiable ledger database, comprising:

[0011] Introducing the sharding extension protocol, which abstracts the database layer interface and provides transparent services for lower-layer applications. It does not need to consider the specific sharding scheme and form, and can be adapted to the storage module of any current blockchain system.

[0012] A synchronization control protocol is introduced. This protocol improves the storage engine interface to implement two-tier storage with read-write separation. Based on the Merkle tree and its derived verifiable index structure, it defines a cross-shard transaction snapshot format. Through snapshot isolation, it supports concurrent lock-free implementation and ensures the integrity of distributed transactions.

[0013] Design the storage engine to optimize the locality of index data to match the sharding strategy and ensure that the data within each shard can be compactly stored in physical storage to reduce latency.

[0014] One possible implementation also includes improving the storage engine of the LSM structure, optimizing the index structure encoding and data page disk placement strategy.

[0015] In one possible implementation, the synchronization control protocol also introduces a concept based on a Merkle tree and a verifiable index structure to define a snapshot format for cross-shard transactions.

[0016] In one possible implementation, sharding divides storage network nodes into small groups, each of which can process transactions in parallel and reduce the storage burden on each node.

[0017] In a possible implementation, the sharding includes:

[0018] Sharding based on data replicas: This method uses network grouping to achieve data consensus, which generally reduces data redundancy in storage engine processing. This requires the implementation of sharding and aggregation algorithms.

[0019] Sharding based on data indexing: Optimizes the ability of parallel computing by separating state data, improving data service performance overall. It requires dynamic load balancing and transactional control. Sharding is isolated, and each operation can be provided by a single shard.

[0020] In one possible implementation, the sharding extension protocol includes a sharding routing protocol and a synchronization control protocol, which can simultaneously adapt to the sharding extension of the data replica-based sharding and the data index-based sharding, and customize a verifiable query interface and a storage engine component that supports transactions and synchronization.

[0021] In one possible implementation, for cross-shard transactions, the following steps are included:

[0022] Start a transaction: After execution, a cross-shard transaction number is generated based on the block and location of the transaction. This number is updated to the current active transaction list, and a snapshot corresponding to the transaction is generated and stored in the transaction context.

[0023] Continue transaction: Based on the transaction ID of the previous chain executed on this chain, inherit the number and snapshot of the cross-shard transaction and continue the transaction operation;

[0024] Commit transaction: Based on the id of the previous on-chain transaction executed on this chain, inherit the number and snapshot of the cross-shard transaction and commit all states related to the cross-shard transaction. The specific operation is to remove the transaction from the system's active transaction list, so that the state changed by the transaction can be read by other transactions;

[0025] Rollback transaction: According to the last on-chain transaction ID executed on this chain, inherit the number and snapshot of the cross-chain transaction, and roll back all states related to the cross-shard transaction.

[0026] In a second aspect, the present application provides a verifiable ledger database storage expansion system, comprising an abstract database layer, multiple storage engine layers, and multiple data pages;

[0027] The abstract database layer is connected to the plurality of storage engine layers, and each data page is connected to a different storage engine page.

[0028] In one possible implementation, the abstract database layer includes multiple tasks, a shard router and a synchronization controller, the shard router is connected to the synchronization controller, and the multiple tasks are connected to the shard router and the synchronization controller.

[0029] In a possible implementation, the storage engine layer includes a disk buffer, concurrency control, a codec, a verifiable index, and a block chain log, and the verifiable index is connected to the block chain log.

[0030] The technical solution provided by this application can achieve at least the following beneficial effects:

[0031] The storage expansion method of the verifiable ledger database provided in the present application provides an implementation scheme for expanding a general distributed storage architecture for a verifiable ledger database, which is applicable to blockchain systems and distributed ledger technology architectures, and is oriented to application scenarios such as data traceability and auditing, meeting both sharding expansion and off-chain storage expansion modes. In the context of sharding technology, it ensures the reliability of cross-shard transaction services, and proposes an LSM storage and synchronization optimization scheme for verifiable data structures. In the sharded storage scheme of the verifiable ledger database, the performance of storage and synchronization is improved, the design of the LSM storage engine is analyzed from the perspective of data, and a semantic bridge is built between the semantic layer of the verifiable ledger database and the underlying storage layer based on the verifiable data structure and data access mode, thereby combining the data layer and the storage layer to achieve collaborative optimization, unifying the storage protocol of the sharded scheme of the verifiable ledger database, and having good scalability in the data interface, which can adapt to different distributed consensus algorithms; at the same time, the present invention integrates the semantic characteristics of the verifiable data layer and the storage layer to optimize the storage and synchronization performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, a brief introduction will be given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0033] FIG1 is a flow chart of a method for expanding storage of a verifiable ledger database according to an exemplary embodiment of the present application;

[0034] FIG2 is a schematic diagram of the system architecture of a verifiable ledger database storage expansion system according to an exemplary embodiment of the present application;

[0035] FIG3 is a schematic diagram of an algorithm of a verifiable query interface shown in an exemplary embodiment of the present application;

[0036] FIG4 is a schematic diagram of data isolation control based on a version chain-block structure, shown in an exemplary embodiment of the present application;

[0037] FIG5 is a schematic diagram of a cross-shard transaction processing process according to an exemplary embodiment of the present application;

[0038] FIG6 is a schematic diagram of a transactional control process according to an exemplary embodiment of the present application;

[0039] FIG7 is a schematic diagram of the architecture of a ledger storage engine shown in an exemplary embodiment of the present application. DETAILED DESCRIPTION

[0040] In order to make the purpose, implementation methods and advantages of the present application clearer, the exemplary implementation methods of the present application will be clearly and completely described below in conjunction with the drawings in the exemplary embodiments of the present application. Obviously, the described exemplary embodiments are only part of the embodiments of the present application, not all of the embodiments. It should be understood that the specific embodiments described here are only used to explain the present application and are not used to limit the present application.

[0041] It should be noted that the brief descriptions of terms in this application are only for the purpose of facilitating the understanding of the embodiments described below, and are not intended to limit the embodiments of this application. Unless otherwise specified, these terms should be understood according to their ordinary and usual meanings.

[0042] In the specification and claims of this application and the accompanying drawings, the terms "first," "second," "third," etc. are used to distinguish similar or similar objects or entities, and are not necessarily intended to limit a particular order or sequence, unless otherwise noted. It should be understood that the terms used in this manner are interchangeable under appropriate circumstances.

[0043] The terms "comprise," "include," and "have," and any variations thereof, are intended to cover but not exclude inclusion; for example, a product or device comprising a list of components is not necessarily limited to all the components expressly listed but may include other components not expressly listed or inherent to such product or device.

[0044] To facilitate the technical solution of the application, some concepts involved in this application are first explained below.

[0045] Verifiable Ledger Database: A Verifiable Ledger Database is a data management and storage system for data objects and their historical transaction records. By maintaining a verifiable data structure and utilizing cryptographic interactive authentication technology, it constructs proofs of existence and integrity for data query results, enabling effective data traceability and auditing.

[0046] Distributed Ledger Technology (DLT): Distributed ledger technology (DLT) is the foundational framework for building shared ledgers across distributed networks. It includes data transmission protocols, ledger synchronization mechanisms, and distributed consensus algorithms, with blockchain technology being a representative form. This distributed architecture is the foundation for scaling verifiable ledger database storage.

[0047] Ledger Storage Engine: A ledger storage engine is an implementation based on the ledger storage protocol. It defines the organizational structure, storage method, and access interface of ledger data, maintains data integrity, reliability, and security, and provides efficient data access and query capabilities. The storage engine is a universal basic unit of distributed systems.

[0048] Authenticated Data Structure (ADS): The Authenticated Data Structure is the core of ledger data storage, interactive authentication, and distributed synchronization, and is a key design for storage protocols and distributed extensions.

[0049] Before explaining the verifiable ledger database storage expansion method provided in the embodiment of the present application, the application scenario and implementation environment of the embodiment of the present application are first introduced.

[0050] Verifiable ledger databases are a direct result of the innovative applications and engineering practices of blockchain systems in data management and storage. Blockchain systems leverage distributed ledger technology, bringing new features to traditional data services and opening up new application scenarios for trusted ledger data services, targeting scenarios requiring data traceability and mutual authentication, such as finance, healthcare, and the Internet of Things. However, industrial applications require more robust technical solutions tailored to specific scenarios. Limited by performance and scalability, blockchain systems are not yet sufficient to completely revolutionize traditional data storage systems. Traditional databases, distributed computing, and cloud service architectures remain the primary technological drivers, giving rise to hybrid blockchain-database data service solutions and independent verifiable ledger database systems.

[0051] Blockchain systems enable decentralized data management from networking to application. Multiple nodes in the network autonomously and synchronously participate in data management and verification, enabling independent verification of data correctness and consistency even in the face of malicious attacks and single points of failure. Verifiable ledger databases are abstracted from the blockchain storage layer. The monotonically increasing ledger data storage requirements further promote the ledger model as a trusted storage proof for interactive applications in scenarios where data computation and storage are separated. These models provide data verification and auditing features, and as independent storage components, they can be customized to provide more flexible and efficient services for different application scenarios. Both of these systems face scalability bottlenecks. As ledger data grows, storage architecture design becomes a major challenge in system expansion, and storage performance is a key optimization point.

[0052] Blockchain defines a basic block-chain data structure, including cryptographic techniques and distributed consensus mechanisms. As the foundation of distributed ledger technology, it enables data synchronization and consistency verification across distributed networks. The ledger data services provided by the hybrid blockchain database architecture must implement: 1) a verifiable query interface, enabling clients to verify the validity of query results from untrusted nodes. This requires storage nodes to traverse the entire blockchain ledger to execute queries; and 2) a distributed consensus interface, ensuring consistency in data synchronization mechanisms across logical units in the distributed network. Therefore, a verifiable ledger database typically consists of two components: an abstract database layer and a ledger storage engine. These components provide transparent ledger data storage and query interfaces to various upper-layer applications, requiring the processing and persistence of large amounts of data.

[0053] The current hybrid architecture of major blockchain database systems is difficult to compare and evaluate due to the lack of uniformity in the abstract database layer. This also means that it is difficult to apply to heterogeneous computing across systems. Verifiable ledger databases combine the abstract database layer and the storage engine layer in their expansion solutions to achieve linear expansion and availability guarantees. Generally speaking, they can be divided into two categories:

[0054] 1. Off-chain storage expansion: This refers to separating computing tasks through offline storage protocols, avoiding processing each transaction through the consensus mechanism, and instead only using consensus to perform key tasks (for example, settlement and dispute resolution).

[0055] 2. Sharded storage expansion: This refers to the horizontal separation of the read and write loads of computing and storage, avoiding system throughput bottlenecks caused by congestion of computing tasks due to a large number of transactions, and achieving better concurrency performance through data isolation.

[0056] The above technical solutions all have great limitations.

[0057] 1. Network Sharding Mechanism: Data sharding is the primary means of blockchain and database expansion. In a distributed environment, the ledger storage engine maintains the consistency of ledger data, requiring the use of complex coordination and synchronization mechanisms. The ledger storage engine is required to provide concurrency control and distributed service interfaces.

[0058] 2. Concurrent execution mechanism: When multiple users or multiple nodes access the ledger storage engine simultaneously, it is necessary to ensure the performance and efficiency of concurrent access. Concurrent read and write operations may cause contention and conflicts, and appropriate concurrency control mechanisms are required to ensure data consistency and concurrent performance.

[0059] 3. Storage index optimization: Ledger storage engines typically need to support indexing and querying of ledger data, as well as meet users' needs for batch access to ledger data. As data volume increases and data storage expands across shards, the efficiency of indexing and querying may decrease, leading to increased query latency.

[0060] Based on this, this application provides a method for expanding storage of a verifiable ledger database, defines a general ledger storage protocol, decouples and abstracts the authentication data service module and the ledger storage engine interface, and supports data sharding expansion and distributed synchronous storage. On the basis of storage expansion, the application further optimizes the performance of the storage engine design, balances storage efficiency and performance, and maintains a certain degree of load balancing flexibility.

[0061] Next, the technical solution of the present application and how the technical solution of the present application solves the above-mentioned technical problems will be described in detail through embodiments and in conjunction with the accompanying drawings. The various embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. Obviously, the described embodiments are only part of the embodiments of the present application, not all of them.

[0062] Its application scenarios are: The fundamental use case for verifiable ledger databases is to ensure data integrity, auditability, and security within distributed systems. Key applications include storing transactions and data in blockchain systems, ensuring that each node can verify data consistency, thereby achieving a decentralized, trusted database. Furthermore, its technological breakthroughs will help advance applications in centralized financial transactions, medical records, supply chain management, IoT data storage, digital identity verification, and vertical industry private chain value interaction.

[0063] FIG1 is a flow chart of a method for expanding storage of a verifiable ledger database, shown in an exemplary embodiment of the present application.

[0064] In an exemplary embodiment, as shown in FIG1 , a method for expanding storage of a verifiable ledger database is provided. In this embodiment, the method may include:

[0065] Step 100: Introduce the sharding extension protocol. The sharding extension protocol abstracts the database layer interface and provides transparent services for lower-layer applications. There is no need to consider the specific sharding scheme and form, and it can be adapted to the storage module of any current blockchain system.

[0066] Step 200: Introduce a synchronization control protocol. The synchronization control protocol improves the storage engine interface to implement two-tier storage with read-write separation, and defines a cross-shard transaction snapshot format based on the Merkle tree and derived verifiable index structure. Through snapshot isolation, it supports concurrent lock-free implementation and ensures the integrity of distributed transactions.

[0067] Step 300: Design the storage engine to optimize the locality of index data to match the sharding strategy and ensure that the data in each shard can be compactly stored in the physical storage to reduce latency.

[0068] One possible implementation involves the introduction of a sharding extension protocol, whose core goal is to improve the scalability of verifiable ledger databases. A data storage layer extension protocol based on sharding is proposed, providing a transparent service interface to upper-layer applications. This partitioning can be adapted to different blockchain systems, such as Ethereum, without modifying the underlying blockchain system's storage module. Sharding allows for parallel processing of transactions and data across different shards, thereby improving the throughput and performance of the entire system. The design is not limited to a specific sharding scheme, but can adapt to a variety of different sharding strategies and forms.

[0069] One possible implementation involves a synchronization control protocol. To support collaboration across multiple shards, the solution introduces a distributed synchronization protocol. This protocol improves the storage engine interface to implement a two-tier storage structure with read-write separation. Based on the concepts of Merkle trees and verifiable index structures, it defines the snapshot format for cross-shard transactions. This makes cross-shard transaction isolation more efficient and supports concurrent operations without the need for traditional locking mechanisms. Through these improvements, the solution can guarantee the integrity of distributed transactions and ensure data consistency across different shards.

[0070] In one possible implementation, a storage engine design is introduced, which manages data storage and retrieval. This solution optimizes the locality of index data to ensure compatibility with the sharding strategy, ensuring that data within each shard can be compactly stored on the physical storage medium to reduce I / O latency. Some embodiments of this application improve the storage engine with an LSM (Log-Structured Merge) structure. By improving the encoding of the index structure and the data page placement strategy, this further enhances the performance and efficiency of the storage engine, helping to reduce the time overhead of data reading and writing.

[0071] In some embodiments of the present application, by deconstructing the verifiable ledger database into an abstract database layer and a storage engine layer, designing a sharding protocol and a synchronization protocol as scalable basic components, and customizing a reusable ledger storage engine for key data models, corresponding performance optimization is achieved.

[0072] Figure 3 is an algorithm diagram of a verifiable query interface shown in an exemplary embodiment of the present application, Figure 4 is a data isolation control diagram based on a version chain-block structure shown in an exemplary embodiment of the present application, Figure 5 is a flow diagram of cross-shard transaction processing shown in an exemplary embodiment of the present application, Figure 6 is a flow diagram of transactional control shown in an exemplary embodiment of the present application, and Figure 7 is an architectural diagram of an account storage engine shown in an exemplary embodiment of the present application.

[0073] In one possible implementation, some embodiments of the present application include abstract storage protocol design to ledger storage engine design and optimization, with the core logic unit comprising the following three parts:

[0074] 1. Storage Extension Protocol

[0075] Sharding is a key technical approach for scaling verifiable ledger databases: Sharding divides storage network nodes into groups, called shards. Each shard can process transactions in parallel and reduce the storage burden on each node. This sharding for database storage and workload introduces a new requirement: cross-shard database services, namely data aggregation for queries and workload balancing for management. Sharding provides a transparent storage expansion protocol for ledger data services, providing a unified application interface for the following common sharding expansion solutions:

[0076] 1) Sharding based on data replicas: This refers to achieving data consensus through network grouping, which generally reduces data redundancy in storage engine processing. This requires the implementation of sharding and aggregation algorithms, including account-based sharding, side chains, offline transactions, and off-chain storage solutions in Ethereum 2.0.

[0077] 2) Data index-based sharding: This refers to optimizing the ability to parallelize computing by separating state data, improving overall data service performance. It requires dynamic load balancing and transactional control. It is common in many blockchain database hybrid architectures. Shards are isolated, and each operation can be provided by a single shard.

[0078] This solution decouples the storage expansion protocol into a shard routing protocol and a synchronization control protocol. It can adapt to both shard expansion methods and customizes the verifiable query interface shown in Figure 3, as well as a storage engine component that supports transactions and synchronization.

[0079] 2. Synchronous Control Protocol

[0080] In account / balance-based state sharding, distributing a large number of transactions raises a natural problem: how to achieve load balance across all shards. Workload imbalance in hot shards may lead to longer transaction confirmation delays, threatening the eventual atomicity of cross-shard transactions. This problem becomes even more challenging when most transactions are cross-shard transactions. Ensuring the eventual atomicity of cross-shard transactions becomes a challenge, hindering the widespread application of state sharding mechanisms in practice.

[0081] MVCC is often used in blockchain systems in combination with block structure to provide data version control and support conventional concurrent operations, but it does not provide data isolation guarantees, which may lead to transaction integrity issues in a sharded environment.

[0082] In some embodiments of the present application, as shown in FIG4 , the underlying storage of the blockchain does not fully utilize the characteristics of the block chain structure. By adjusting the MVCC concurrency control protocol in the ledger storage engine, feature support for online and transaction mode evolution is provided.

[0083] The blockchain structure often uses a Merkle tree-related index structure, such as MPT, in which each block contains a Merkle tree that stores the historical state data of the blockchain. Each leaf node has a pointer to the previous state of the state, thereby effectively supporting snapshot generation and state rollback in multi-version concurrency control. The leaf node in the Merkle tree is L and the middle node is M. Each leaf node contains the specific value of the state it manages and a pointer to the previous state of the state. The leaf node can be represented as Among them, k i is the state key value corresponding to the leaf node, v j is the version number of the state, t n The transaction number that updates the status to the current value is The value of the middle node of the Merkle tree is k M =hash(L i ,f(L i )=M), the function obtains L i The parent node of the node. The hash is the hash of the sequence after the values ​​of all nodes in the set are concatenated.

[0084] For a cross-shard transaction, as shown in Figure 5, the specific operations of the transaction method at each stage are as follows:

[0085] 1. Begin transaction BeginTx: After execution, a cross-shard transaction number is generated based on the block and location of the transaction. This number is updated to the current active transaction list, and a snapshot corresponding to the transaction is generated and stored in the transaction context.

[0086] 2. Continue transaction ContinueTx: Based on the last on-chain transaction ID executed on this chain, it inherits the number and snapshot of the cross-shard transaction and continues the transaction operation;

[0087] 3. Commit transaction CommitTx: Based on the id of the previous on-chain transaction executed on this chain, inherit the number and snapshot of the cross-shard transaction and commit all states related to the cross-shard transaction. The specific operation is to remove the transaction from the system's active transaction list, so that the state changed by the transaction can be read by other transactions;

[0088] 4. Rollback Transaction RollbackTx: Based on the ID of the previous on-chain transaction executed on this chain, it inherits the number and snapshot of the cross-chain transaction and rolls back all states related to the cross-shard transaction. The specific operation is to read the write set of the transaction from the snapshot, obtain the version of all states in the write set before being modified by the transaction, use this version to update the state, and then remove it from the system's active transaction list.

[0089] In the basic read method, as shown in FIG6 , the state and its corresponding version number are first read.

[0090] Then determine whether the state version number is greater than the maximum version number of the transaction in the active list. If it is, it means that this state has been updated by other transactions after the transaction snapshot was generated. The state at this time should be invisible to this transaction. Therefore, the state is continuously rolled back to the previous version until the state version is less than the maximum number in the active transaction list.

[0091] Next, the state version is checked to see if it is less than the minimum version in the active transaction list. If so, the transaction that updated the state has completed, and this version of the state is visible to all transactions, including this one. The state value of this version is returned. If not, the state version number is between the minimum and maximum transaction version numbers in the active list, and the process continues to determine whether the transaction that updated the state is the current transaction.

[0092] If so, the update is visible to the current transaction and the status value of the version is returned; if not, continue to determine whether the transaction that updated the status is in the active list.

[0093] If it is in the active list, it means that the status update has not been committed. The current status is not visible to other transactions, including this transaction. The status needs to be rolled back to the previous version until the version update is completed and committed. If it is not in the active list, it means that the transaction that updated the status has been committed. The version is visible to all transactions, including this transaction. The status value of this version is returned.

[0094] 3. Storage engine optimization

[0095] Some embodiments of this application observe that the ledger storage engine may experience service interruptions and performance bottlenecks during data synchronization, as shown in Figure 7. This led to the design of an asynchronous, non-blocking migration method. This method employs a data aggregation technology based on the multi-level storage structure of the LSM storage engine and employs a dual mode of cross-shard synchronization and concurrency control to minimize the impact of the migration process on the tables being migrated.

[0096] This solution deeply studies the process of node serialization storage in the secure Merkle Patricia Trie (MPT) structure implemented in Ethereum and its impact on the system query performance. The serialized storage implemented by the hash function causes the key values ​​to be randomly mapped to different key ranges, resulting in unplanned data layout and loss of the original position semantics. This introduces the reading of multiple ordered string (Sorted String Table, SST) files in the state data request, which seriously reduces the query performance of the system. To solve this problem, we propose a strategy based on semantic mapping, which aggregates the key-value pairs in the same query path into adjacent SST files by prefixing, so as to reduce the additional disk IO during the query process. Combined with the data construction supported by transactions in the previous article, the stored procedure encoding scheme is as follows: k M = prefix + hash (L i ,f(L i )=M)

[0097] Each access to the MPT involves a complete branch from the root node to the leaf nodes. Due to the tree-structured index design, adjacent nodes must be parent and child nodes of each other, resulting in a certain degree of temporal locality.

[0098] Therefore, some embodiments of the present application cluster nodes that are parent-child relationships in adjacent locations on the storage device to ensure sequential reading of data and improve storage performance. In this regard, the LSM-Tree-based storage component has shown great potential. Its ordered SSTtable and multi-level data cache structure provide adjacent storage locations for keys with similar lexicographical order, and exert spatial locality through block cache, further improving storage performance.

[0099] It can be seen that transparent expansion of the verifiable ledger database does have certain requirements for the storage engine to support an efficient, scalable, and reliable database sharding mechanism. Some embodiments of this application focus on the following optimization points:

[0100] 1. Concurrent transaction control: The storage engine needs to combine the index structure to handle high-concurrency random reads and writes, and even coordinate and manage cross-shard transactions.

[0101] 2. Efficient read and write capabilities: The storage engine should optimize data locality to match the sharding strategy and ensure that the data within each shard can be compactly stored in the physical storage to reduce IO latency.

[0102] 3. Partial synchronization mechanism: The storage engine should support efficient migration of data. As data access patterns change, the sharding strategy may need to be readjusted without causing large amounts of data migration or downtime.

[0103] It should be understood that although the steps in the flowcharts of the above embodiments are shown in sequence as indicated, these steps are not necessarily executed in the order indicated. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps can be executed in other orders. Moreover, at least a portion of the steps in the flowcharts of the above embodiments may include multiple steps or multiple stages, and these steps or stages are not necessarily executed at the same time, but can be executed at different times. The order of execution of these steps or stages is not necessarily sequential, but can be executed in rotation or alternation with other steps or at least a portion of steps or stages in other steps.

[0104] Corresponding to the aforementioned embodiment of the verifiable ledger database storage expansion method, using the same technical concept, this application also provides an embodiment of a verifiable ledger database storage expansion system.

[0105] FIG2 is a schematic diagram of the system architecture of a verifiable ledger database storage expansion system shown in an exemplary embodiment of the present application.

[0106] In an exemplary embodiment, as shown in FIG2 , the verifiable ledger database storage expansion system includes an abstract database layer, multiple storage engine layers, and multiple data pages;

[0107] The abstract database layer is connected to the plurality of storage engine layers, and each data page is connected to a different storage engine page.

[0108] In one possible implementation, the abstract database layer includes multiple tasks, a shard router and a synchronization controller, the shard router is connected to the synchronization controller, and the multiple tasks are connected to the shard router and the synchronization controller.

[0109] In a possible implementation, the storage engine layer includes a disk buffer, concurrency control, a codec, a verifiable index, and a block chain log, and the verifiable index is connected to the block chain log.

[0110] The specific limitations of the verifiable ledger database storage expansion system can be found in the limitations of the verifiable ledger database storage expansion method described above and will not be further elaborated here. Each module in the aforementioned verifiable ledger database storage expansion system can be implemented in whole or in part through software, hardware, or a combination thereof. Each of these modules can be embedded in or independent of a processor within a computer device in hardware form, or stored in a computer device memory in software form, allowing the processor to call and execute the corresponding operations of each module.

[0111] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0112] The embodiments described above merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that a person skilled in the art could make various modifications and improvements without departing from the spirit of the present application, all of which fall within the scope of protection of the present application. Therefore, the scope of protection of the present patent application shall be determined by the appended claims.

Claims

1. A method for expanding the storage of a verifiable ledger database, characterized in that, it includes: Introduce a sharding expansion protocol. The sharding expansion protocol abstracts the database layer interface to provide transparent services for lower-layer applications without considering specific sharding schemes and forms, and can be adapted to the storage module of any current blockchain system; Introduce a synchronization control protocol. The synchronization control protocol improves the storage engine interface to implement a two-layer storage with read-write separation, and based on the Merkle tree and the derived verifiable index structure, defines the snapshot format for cross-shard transactions. Through snapshot isolation, it supports concurrent lock-free implementation and ensures the integrity of distributed transactions; Design the storage engine to optimize the locality of index data to match the sharding strategy, and ensure that the data within each shard can be compactly stored in physical storage to reduce latency.

2. The method for expanding the storage of a verifiable ledger database according to claim 1, characterized in that, it also includes optimizing the index structure encoding and data page disk dropping strategy through a storage engine that improves the LSM structure.

3. The method for expanding the storage of a verifiable ledger database according to claim 1, characterized in that, the synchronization control protocol also introduces the concept based on the Merkle tree and the verifiable index structure for defining the snapshot format of cross-shard transactions.

4. The method for expanding the storage of a verifiable ledger database according to claim 1, characterized in that, sharding divides the storage network nodes into groups, and each shard can process transactions in parallel and reduce the storage burden on each node.

5. The method for expanding the storage of a verifiable ledger database according to claim 4, characterized in that, the sharding includes: Sharding based on data replicas: Through network grouping for data consensus, generally reducing the data redundancy processed by the storage engine, and it is necessary to implement a sharding algorithm and an aggregation algorithm; Sharding based on data indexes: Optimize the parallel computing ability by separating state data, generally improving the data service performance, and it is necessary to implement dynamic load balancing and transactional control. The shards are isolated, and each operation can be provided by a single shard.

6. The method for expanding the storage of a verifiable ledger database according to claim 5, characterized in that, the sharding expansion protocol includes a sharding routing protocol and a synchronization control protocol, which can simultaneously adapt to the sharding expansion of the sharding based on data replicas and the sharding based on data indexes, and customize a verifiable query interface, as well as a storage engine component that supports transactions and synchronization.

7. The method for expanding the storage of a verifiable ledger database according to claim 3, characterized in that, for cross-shard transactions, it includes the following steps: Start a transaction: After execution, generate the number of this cross-shard transaction according to the block and position where this transaction is located, update this number to the current active transaction list, and generate the snapshot corresponding to this transaction and store it in the transaction context; Continue the transaction: According to the in-chain transaction id executed last on this chain, inherit the number and snapshot of this cross-shard transaction, and continue to perform this transaction operation; Commit transaction: Inherit the number and snapshot of the cross-shard transaction according to the in-chain transaction ID of the last in-chain transaction executed on this chain, and commit all the states related to this cross-shard transaction. The specific operation is to remove this transaction from the system active transaction list, so that the states changed by this transaction can be read by other transactions; Rollback transaction: Inherit the number and snapshot of the cross-chain transaction according to the in-chain transaction ID of the last in-chain transaction executed on this chain, and roll back all the states related to this cross-shard transaction.

8. A verifiable ledger database storage extension system, characterized in that, it includes an abstract database layer, multiple storage engine layers and multiple data pages; The abstract database layer is connected to multiple storage engine layers, and each data page is respectively connected to a different storage engine page.

9. The verifiable ledger database storage extension system according to claim 8, characterized in that, the abstract database layer includes multiple tasks, a shard router and a synchronization controller, the shard router is connected to the synchronization controller, and multiple tasks are connected to the shard router and the synchronization controller.

10. The verifiable ledger database storage extension system according to claim 8, characterized in that, the storage engine layer includes a disk buffer, concurrency control, a codec, a verifiable index and a block chain log, and the verifiable index is connected to the block chain log.

Citation Information

Patent Citations

  • Method for realizing transverse expansion of distributed account book based on fragmentation mechanism

    CN110310115A

  • Block chain-based under-chain capacity expansion technology

    CN116388957A

  • Alliance chain account book extension storage method based on state data collaboration

    CN116662443A

  • Distributed verifiable ledger database

    WO2023177358A1

Cited By

  • Data storage method for artificial intelligence learning mode

    CN120723945A

  • Efficient database reading method based on multi-level cache optimization and dynamic index fragmentation

    CN121092593A

  • Asynchronous verification and batch storage system and method for high-concurrency voucher data

    CN121681152A

  • Semantic identification block chain retrieval method and system for large-scale distributed data

    CN122285778A