Persistent memory data storage method and system for distributed database
By designing a persistent in-memory data storage method for distributed databases in distributed databases, the problem of performance bottlenecks and high failure recovery overhead in distributed databases and persistent data storage is solved, and efficient and reliable data storage and concurrent access are achieved.
Patent Information
- Application Number
- CN202310196821.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-27
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2043-02-27
AI Technical Summary
When the prior art faces efficient and reliable data storage of distributed databases and persistent memory, there are problems such as low write performance, poor fragmentation management efficiency, unsuitable index structure for range search and large logging overhead, resulting in high performance bottlenecks and failure recovery overhead.
A persistent in-memory data storage method for distributed databases is designed by sequentially storing persistent data in PMems in block-aligned manner in the storage engine, and recording meta information and indexes in DRAM. The method includes partitioning of thread access areas, spatial management of implicit linked lists, key-to-address mapping, and non-log validity determination, and supports efficient read and write operations and concurrent access.
Improves the write operation and concurrent access performance of persistent memory, reduces the transactional overhead of write operations, improves the read efficiency of key-value pairs, and supports range reads, reducing the overhead of failure recovery.
Smart Images

Figure CN116226232B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data storage, and in particular to a method and system for persistent memory data storage for distributed databases. Background Art
[0002] Database systems are important carriers of contemporary network services and are widely used in online user service scenarios such as e-commerce, finance, and browser search. Under the current mainstream memory-disk storage architecture, database systems can benefit from the high read and write performance of memory and the persistence of disks, but are also subject to the volatility of memory and the low performance of disks, facing a trade-off between performance and persistence. As the demand for online processing of database systems increases, the number of requests processed per second by a single server node in a cluster can reach hundreds of thousands or even more. Each request is accompanied by persistent device I / O ranging from tens of bytes to tens of KB. In addition, the system regularly performs compression and copy operations to maintain persistent storage content, resulting in the existing persistent device performance having a significant limit on system throughput, bringing severe challenges.
[0003] Persistent Memory (PMem) is a new type of storage medium that has low read and write latency and byte addressing characteristics, as well as the ability to provide persistent storage. Its performance is between that of memory and disk, making it an ideal storage device to make up for the shortcomings of memory and disk. Currently, the most prominent application scenario of persistent memory is to provide faster and more stable persistence support than disk for traditional memory applications (such as memory data storage engines), greatly reducing the cost of system recovery. Industry practice shows that persistent memory is suitable for data storage scenarios with high data persistence requirements, high access frequency, and low tolerance for performance fluctuations.
[0004] Tests show that compared with DRAM, persistent memory still has a large performance gap in some aspects, such as low write bandwidth, fine-grained write amplification, and poor random and parallel access performance. How to solve the above performance problems and better apply persistent memory to scenarios such as distributed databases is a hot topic in the field of persistent memory.
[0005] The data storage system for persistent memory is a typical persistent memory application case. Take a typical PMem-based key-value storage system as an example. Figure 1 As shown in the figure, the data to be persisted from the upper-layer application is stored in the PMem device by calling the write interface in the form of key-value pairs; the hash index in DRAM supports fast access for read and delete operations. In order to avoid PMem's write performance issues as much as possible, the data stored in PMem is usually organized in a sequential, block-aligned form. In the data access scenario where large-scale, multiple read and write operations are interleaved, there are the following three performance challenges:
[0006] (1) In order to ensure efficient use of storage space, the released fragments need to be properly managed. In the block-aligned storage mode, the current mainstream fragment management mechanism for PMem is that when the total number of valid entries in a block is too small, the entries in the block are merged to other locations in order to release a whole block of space for reallocation; this mechanism will lead to additional write-time copying, which not only occupies read and write bandwidth, but also cannot be multi-threaded and has poor performance. A new fragment management mechanism needs to be designed;
[0007] (2) In order to ensure efficient access to stored data, the key-value storage system needs to maintain a suitable index structure. In systems such as distributed databases, there are a large number of range search requirements, and unordered hash indexes are difficult to effectively support range operations; ordered search tree indexes exist based on DRAM, PMem, and hybrid solutions, facing a trade-off between performance and reliability. It is urgent to design a reasonable index mechanism to meet the data persistence requirements in distributed databases;
[0008] (3) In order to ensure the transactional nature of persistent operations and prevent persistent data from becoming invalid due to failures, additional records are required for data modification operations. Traditional logging solutions record every modification operation, resulting in significant write amplification. In some scenarios where it is uncertain whether the modification will be committed (such as database concurrency control), a large number of modification operations and their log records do not need to be generated, and the persistence overhead is significantly higher. There is an urgent need to design a PMem-based persistence mechanism to reduce fault recovery overhead.
[0009] Based on the above analysis, it is urgent to design an efficient and reliable data storage method based on the characteristics of distributed databases and persistent memory, improve the performance of database systems when dealing with large numbers of read and write operations, and reduce fault recovery overhead. Summary of the invention
[0010] The technical task of the present invention is to address the above shortcomings and provide a persistent memory data storage method and system for distributed databases to solve the technical problems of how to achieve efficient and reliable data storage for distributed databases and persistent memory, improve the performance of the database system when dealing with a large number of read and write operations, and reduce fault recovery overhead.
[0011] In a first aspect, the present invention provides a persistent memory data storage method for a distributed database, which is applied to a storage engine and multiple access devices, wherein the storage engine includes a DRAM and a PMem, wherein the DRAM is used to map the content in the PMem and record the saved data and the number of blocks in the meta information, wherein the storage space of the PMem is divided into multiple blocks, each block is divided into an equal number of thread access areas, and the access device, as a memory object, corresponds to an access thread, and each access device accesses multiple thread access areas with the same number;
[0012] The method comprises the following steps:
[0013] The storage engine stores persistent data in PMem in a block-aligned manner, loads blocks into the data mapping area of DRAM in sequence according to the block alignment size, and records the saved persistent data and the number of blocks in the meta information.
[0014] Record the created blocks in DRAM through a list structure;
[0015] Validity determination is performed on data entries in the thread access area in PMem, and for data entries that pass the determination, key-to-address mapping is performed on the data entries, and the key-to-address mapping corresponding to each data entry is recorded in DRAM through an index to support query through a search tree structure;
[0016] The free data entries released on PMem are organized into an implicit linked list freelist, and operation management operations are performed on the data entries, wherein the operation management includes allocation, release and merging;
[0017] Create an accessor instance, receive read and write requests initiated by the upper-layer application through the storage instance, and assign a number to the accessor through the storage engine based on the read and write requests, wherein the number corresponds to a thread access area in PMem that the accessor can rewrite, wherein the thread access area stores data entries and meta-information, and performs data operations based on the read and write requests, wherein the data operations include read operations, write operations, and delete operations, wherein the write operations and delete operations are limited by the number of the thread access area, and can only perform write operations or delete operations on data entries with the same number, and the read operation is not limited.
[0018] Preferably, each thread access area is used to store data entries and a lock flag and a freelist local entry pointer;
[0019] The data entry is aligned with 8 bytes in PMem, including an 8-byte header field, a data field in the middle, and an 8-byte footer field at the end. The header field and the footer field are both used to record the length and valid identifier of the data entry. When the header and footer are inconsistent, it indicates that the data entry is invalid.
[0020] The lock mark is stored at the starting position of the corresponding thread access area, and is used to mark whether the corresponding thread access area is being rewritten by an accessor. The lock mark is based on the thread access area, and multiple accessors can be assigned the same number to rewrite data in different thread access areas.
[0021] Each thread access area stores the address of the first free data entry belonging to the thread access area, and the first free data entry address is used as a local entry of the freelist so that the newly released free data entry can be inserted into the freelist of the thread access area nearby.
[0022] Preferably, determining the validity of each data entry in the thread access area includes the following steps:
[0023] Verify the validity of the freelist pointers before and after the free data entry: If they are invalid, it means that the data entry pointed to has not been updated when the pointer is merged and the power is turned off, then it is directly modified to the starting address of the target data entry;
[0024] There are two consecutive free data entries: This indicates that the data entries are not written after allocation and the power is turned off, the contents of the data entries are discarded, and the latter data entries are merged forward;
[0025] Verify the consistency of header and footer: Since the header is updated earlier than the footer in the data entry, and the timing of updating the header is to complete the key value writing or prepare to release the data entry, the status of the header shall prevail during recovery. If the header indicates that the data entry is free, it means that the release is not completed, and the data entry continues to be released; if the data entry has been allocated, the footer is updated and regarded as a valid data entry, and added to the index as usual;
[0026] Verify whether the key of the current data entry already exists: According to the mechanism of alternating the headers of the new and old data entries in the insertion operation, the header and footer of the old data entry must be inconsistent before the old data entry is released, and the header is marked as released. During recovery, the release operation can be completed according to the rules for verifying the consistency of the header and footer. In the case where the power is cut off when the header of the old data entry is not marked as released, the footer of the new data entry has not been modified. Therefore, during recovery, the data entry whose footer has not been modified will be used as the basis, and the other one will be deleted.
[0027] Preferably, the data field of the free data entry is filled with forward and backward pointers, and is organized into a freelist with other free entries, and the forward and backward pointers of the free data entry are stored after the header;
[0028] The allocation operation for free data entries is as follows: select an entry of appropriate length from the freelist, intercept it as needed, modify the allocation identifier and the pointers of the previous and next entries, and then hand it over to the accessor to write the data;
[0029] The release operation for free data entries is: directly modify the allocation flag, insert into the freelist and add the front and back pointers;
[0030] The merging operation for free data entries is: merging two free blocks with adjacent physical addresses, which is triggered after the entry is released;
[0031] Among them, when the free data entries are reallocated, they are traversed according to the front and back pointers. All the free data entries in the thread access area corresponding to the same number form the same freelist and are provided to the same accessor for access. The freelists between the thread access areas corresponding to different numbers do not overlap to ensure thread isolation.
[0032] As a preference,
[0033] When the accessor performs a write operation, it accesses the thread access area with the corresponding number in the storage engine, traverses the freelist to obtain free data entries of appropriate length, cuts off the excess parts and adds them back to the freelist, writes the key-value pair content into the data entry, marks the length and allocation flag, modifies the pointers of the predecessor and successor elements of the freelist, and finally inserts the new data entry into the index;
[0034] When the accessor performs a read operation, it directly searches the data entry address of the target key-value pair in PMem from the index and then reads it from PMem. If the search fails, it means that the key-value pair does not exist.
[0035] Since the thread access area is aligned by address, the number to which the data entry belongs can be directly determined by the thread access area address. Before modifying or deleting the old value, the number can be determined by a read operation, and then the number is assigned to the accessor to perform the operation;
[0036] When a known numbered accessor deletes a key-value pair in the corresponding area, it directly marks the data entry as free, inserts the local entry into the freelist, and removes the key-value pair from the index;
[0037] When a release operation causes two data entries adjacent to each other in physical addresses to be free, the two data entries are merged into one free entry, and the front and back pointers are modified accordingly.
[0038] In a second aspect, the present invention provides a persistent memory data storage system for a distributed database, which is applied to a storage engine and a plurality of access devices, and is used to implement storage of persistent memory data through a persistent memory data storage method for a distributed database as described in any one of the first aspects;
[0039] The system comprises:
[0040] A data storage module, the data storage module is used to sequentially store persistent data in PMem in a block-aligned manner through a storage engine, load the blocks into the data mapping area of DRAM in sequence according to the block alignment size, and record the saved persistent data and the number of blocks in the meta information; record the created blocks in the DRAM through a list structure; and perform validity determination on data entries in the thread access area in PMem;
[0041] A space management module, the space management module is used to organize the free data entries released on the PMem into an implicit linked list freelist, and perform operation management operations on the data entries, the operation management including allocation, release and merging;
[0042] A key-value index module, which is used to map the key to the address of the data entry. The DRAM records the key-to-address mapping corresponding to each data entry through an index to support querying through a search tree structure; and is used to locate the data entry and query through the search tree structure;
[0043] A data access interface, the data access interface is used to accept read and write requests from upper-layer applications, and based on the read and write requests, a number is assigned to the accessor through the storage engine, the number corresponds to a thread access area that the accessor can rewrite, the thread access area stores data entries and meta-information, and is used to create accessor instances to perform data operations, the data operations include read operations, write operations, and delete operations, the write operations and delete operations are limited by the number of the thread access area, and can only write or delete data entries within the same number, and the read operation is not limited.
[0044] Preferably, the data storage module is used to divide the storage space of PMem into blocks aligned with fixed-size addresses, and divide each block into equal thread access areas, the size of each thread access area depends on the concurrency degree of concurrent access supported and the average size of key-value data, and each thread access area is used to store data entries and a lock flag and a freelist local entry pointer;
[0045] The data entry is aligned with 8 bytes in PMem, including an 8-byte header field, a data field in the middle, and an 8-byte footer field at the end. The header field and the footer field are both used to record the length and valid identifier of the data entry. When the header and footer are inconsistent, it indicates that the data entry is invalid.
[0046] The lock mark is stored at the starting position of the corresponding thread access area, and is used to mark whether the corresponding thread access area is being rewritten by an accessor. The lock mark is based on the thread access area, and multiple accessors can be assigned the same number to rewrite data in different thread access areas.
[0047] Each thread access area stores the address of the first free data entry belonging to the thread access area, and the first free data entry address is used as a local entry of the freelist so that the newly released free data entry can be inserted into the freelist of the thread access area nearby.
[0048] The DRAM reflects the content in PMem by mmap mapping, and records the currently created blocks by list structure. When there is not enough space in the accessor to write new data entries, the data storage module is used to create new blocks and expand the storage space. When the storage system is restored, it is used to read the PMem mapping file, determine the first address of the storage space, and read the subsequent blocks in sequence and remap them.
[0049] The data storage module is used to determine the validity of each data item in the thread access area through the following steps:
[0050] Verify the validity of the freelist pointers before and after the free data entry: If they are invalid, it means that the data entry pointed to has not been updated when the pointer is merged and the power is turned off, then it is directly modified to the starting address of the target data entry;
[0051] There are two consecutive free data entries: This indicates that the data entries are not written after allocation and the power is turned off, the contents of the data entries are discarded, and the latter data entries are merged forward;
[0052] Verify the consistency of header and footer: Since the header is updated earlier than the footer in the data entry, and the timing of updating the header is to complete the key value writing or prepare to release the data entry, the status of the header shall prevail during recovery. If the header indicates that the data entry is free, it means that the release is not completed, and the data entry continues to be released; if the data entry has been allocated, the footer is updated and regarded as a valid data entry, and added to the index as usual;
[0053] Verify whether the key of the current data entry already exists: According to the mechanism of alternating the headers of the new and old data entries in the insertion operation, the header and footer of the old data entry must be inconsistent before the old data entry is released, and the header is marked as released. During recovery, the release operation can be completed according to the rules for verifying the consistency of the header and footer. In the case where the power is cut off when the header of the old data entry is not marked as released, the footer of the new data entry has not been modified. Therefore, during recovery, the data entry whose footer has not been modified will be used as the basis, and the other one will be deleted.
[0054] Preferably, the data field of the free data entry is filled with forward and backward pointers, and is organized into a freelist with other free entries, and the forward and backward pointers of the free data entry are stored after the header;
[0055] The allocation operation traverses and selects an entry of suitable length from the freelist, intercepts it as needed, modifies the allocation identifier and the pointers of the previous and next entries, and then hands it over to the accessor to write data;
[0056] The release operation directly modifies the allocation flag, inserts into the freelist, and adds the front and back pointers;
[0057] The merge operation is used to merge two free blocks with adjacent physical addresses, which is triggered after the entry is released;
[0058] Among them, when the free data entries are reallocated, the space management module is used to execute: traverse according to the front and back pointers, all the free data entries in the thread access area corresponding to the same number form the same freelist, and are provided to the same access device for access; the freelists between the thread access areas corresponding to different numbers do not overlap to ensure thread isolation.
[0059] Preferably, the key-value index module is used to rebuild by scanning when the system is restarted.
[0060] Preferably, the data access interface provides three operation interfaces, namely a read interface, a write interface and a delete interface, and each operation interface is allocated with an accessor to undertake read and write operations;
[0061] The read interface includes single read and range read, and the ordered access within the range is realized by pre-order traversal search of the radix tree index;
[0062] When calling the accessor to perform a read operation, the data entry address of the target key-value pair in PMem is directly searched from the index, and then read from PMem. If the search fails, it means that the key-value pair does not exist;
[0063] The write interface is used to call the allocation operation of the space management module to allocate entries for new data and write data, and to determine whether there is an old version of the corresponding key through a pre-read operation, and to perform an additional deletion operation on the old version;
[0064] When the accessor is called to perform a write operation, it accesses the thread access area with the corresponding number in the storage engine, traverses the freelist to obtain free data entries of appropriate length, cuts off the excess parts and adds them back to the freelist, writes the key-value pair content to the data entry, marks the length and allocation flag, modifies the pointers of the predecessor and successor elements of the freelist, and finally inserts the new data entry into the index;
[0065] The deletion interface uses a lightweight logical deletion method to call the release operation and merge operation of the space management module to reclaim the storage space;
[0066] Since the thread access area is aligned by address, when performing a delete operation, the number to which the data entry belongs can be directly determined by the thread access area address. Before modifying or deleting the old value, the number can be determined by a read operation, and then the number is assigned to the accessor to perform the operation;
[0067] When a known numbered accessor deletes a key-value pair in the corresponding area, it directly marks the data entry as free, inserts the local entry into the freelist, and removes the key-value pair from the index.
[0068] The persistent memory data storage method and system for distributed databases of the present invention have the following advantages:
[0069] 1. Improve the write operation and concurrent access performance of persistent memory through data storage, space management and index modules;
[0070] 2. The non-log validity determination scheme is used to ensure the transactional nature of write operations and reduce the write overhead in non-fault states;
[0071] 3. The reading efficiency of key-value pairs is improved through ordered indexes, and range reading is supported. BRIEF DESCRIPTION OF THE DRAWINGS
[0072] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0073] The present invention is further described below in conjunction with the accompanying drawings.
[0074] Figure 1 This is a schematic diagram of a persistent memory key-value storage scenario;
[0075] Figure 2 This is a schematic diagram of the key-value storage engine principle;
[0076] Figure 3 This is a schematic diagram of the freelist principle;
[0077] Figure 4 This is a schematic diagram of the improved concurrent control system of this embodiment. DETAILED DESCRIPTION
[0078] The present invention is further described below in conjunction with the accompanying drawings and specific embodiments so that those skilled in the art can better understand the present invention and implement it. However, the embodiments are not intended to limit the present invention. In the absence of conflict, the embodiments of the present invention and the technical features in the embodiments may be combined with each other.
[0079] The embodiments of the present invention provide a persistent memory data storage method and system for a distributed database, which are used to solve the technical problem.
[0080] Embodiment 1:
[0081] The present invention provides a method for persistent memory data storage for distributed databases. The method mainly includes two types of participants: (1) a storage engine that stores and manages PMem data (hereinafter referred to as the engine). (2) an initiator of data access operations (hereinafter referred to as the accessor). The engine is composed of Figure 2 As shown in the figure, it consists of two parts: DRAM and PMem. The DRAM part contains three types of volatile data: data mapping area, access control and index. The PMem part contains persistent data stored sequentially in a block-aligned manner, and organizes the free space into an implicit linked list (freelist). Each block is divided into an equal number of thread access areas. Multiple accessors simultaneously access access areas with different numbers, and interleave in the address space to improve the efficiency of parallel access. An accessor is a memory object that corresponds to an access thread and accesses multiple access areas with the same number.
[0082] The method comprises the following steps:
[0083] S100, sequentially storing the persistent data in PMem in a block-aligned manner through a storage engine, sequentially loading the blocks into a data mapping area of DRAM according to the block alignment size, and recording the saved persistent data and the number of blocks in the meta information;
[0084] Record the created blocks in DRAM through a list structure;
[0085] The validity of data entries in the thread access area in PMem is determined. For data entries that pass the determination, the key-to-address mapping is performed on the data entries. The key-to-address mapping corresponding to each data entry is recorded in DRAM through an index to support query through a search tree structure.
[0086] S200, forming the free data entries released on PMem into an implicit linked list freelist, and performing operation management operations on the data entries, wherein the operation management includes allocation, release and merging;
[0087] S300, create an accessor instance, receive read and write requests initiated by the upper-layer application through the storage instance, and assign a number to the accessor through the storage engine based on the read and write requests. The number corresponds to a thread access area in PMem that the accessor can rewrite. The thread access area stores data entries and meta-information, and performs data operations based on the read and write requests. The data operations include read operations, write operations, and delete operations. The write operations and delete operations are limited by the number of the thread access area, and can only write or delete data entries within the same number. The read operation is not limited.
[0088] In this embodiment, each thread access area is used to store data entries, a lock flag, and a freelist local entry pointer.
[0089] Data entries are aligned with 8 bytes in PMem, including an 8-byte header field, a data field in the middle, and an 8-byte footer field at the end. Both the header field and the footer field are used to record the length and valid identifier of the data entry. When the header and footer are inconsistent, it means that the data entry is invalid.
[0090] The lock flag is saved at the beginning of the corresponding thread access area to mark whether there is an accessor in the corresponding thread access area that is being rewritten. Since the lock flag is in units of areas, multiple accessors can be assigned the same number to rewrite data in different areas.
[0091] At the same time, each thread access area stores the address of the first free data entry belonging to the thread access area, and the first free data entry address is used as a local entry of the freelist so that the newly released free data entry can be inserted into the freelist of the thread access area nearby.
[0092] In this embodiment, determining the validity of each data entry in the thread access area includes the following steps:
[0093] (1) Verify the validity of the freelist pointers before and after the freelist of the idle data entry: If it is invalid, it means that the data entry pointed to has not been updated when the pointer is merged and the power is turned off, then it is directly modified to the starting address of the target data entry;
[0094] (2) There are two consecutive free data entries: This indicates that the data entry is not written after allocation and the power is turned off, the content of the data entry is discarded, and the latter data entry is merged forward;
[0095] (3) Verify the consistency of the header and footer: Since the header is updated earlier than the footer in the data entry, and the header is updated when the key value is written or the data entry is ready to be released, the status of the header is used as the basis for recovery. If the header indicates that the data entry is free, it means that the release is not completed, and the data entry continues to be released; if the data entry has been allocated, the footer is updated and regarded as a valid data entry and added to the index as usual;
[0096] (4) Verify whether the key of the current data entry already exists: According to the mechanism of alternating the headers of the new and old data entries in the insertion operation, the header and footer of the old data entry must be in a state of inconsistency before the old data entry is released, and the header is marked as released. During recovery, the release operation can be completed according to the rules for verifying the consistency of the header and footer. In the case where the header of the old data entry is not marked as released and the power is cut off, the footer of the new data entry has not been modified. Therefore, during recovery, the data entry whose footer has not been modified shall prevail and the other one shall be deleted.
[0097] In this embodiment, the data field of the free data entry is filled with forward and backward pointers, and is organized into a freelist with other free entries. The forward and backward pointers of the free data entry are saved after the header; the allocation operation of the free data entry is: traverse and select an entry of appropriate length from the freelist, intercept it as needed, modify the allocation identifier and the pointers of the forward and backward entries, and then hand it over to the accessor to write the data.
[0098] The release operation of the free data entry is: directly modify the allocation mark, insert the freelist and add the front and back pointers. When the free data entry is reallocated, it is traversed according to the front and back pointers. All the free data entries in the thread access area corresponding to the same number form the same freelist and are provided to the same accessor for access; the freelists between the thread access areas corresponding to different numbers do not overlap to ensure thread isolation.
[0099] The merging operation for free data entries is: merging two free blocks with adjacent physical addresses, which is triggered after the entry is released;
[0100] In this embodiment, when the accessor performs a write operation, it accesses the thread access area with the corresponding number in the storage engine, traverses the freelist to obtain free data entries of appropriate length, cuts off the excess parts and adds them back to the freelist, writes the key-value pair content into the data entry, marks the length and allocation identifier, modifies the pointers of the predecessor and successor elements of the freelist, and finally inserts the new data entry into the index.
[0101] When the accessor performs a read operation, it directly searches for the data entry address of the target key-value pair in PMem from the index, and then reads it from PMem. If the search fails, it means that the key-value pair does not exist.
[0102] Since the thread access area is aligned by address, the number to which the data entry belongs can be determined directly by the thread access area address. Before modifying or deleting the old value, the number can be determined by a read operation and then assigned to the accessor to perform the operation.
[0103] When a known numbered accessor deletes a key-value pair in the corresponding area, it directly marks the data entry as free, inserts the local entry into the freelist, and removes the key-value pair from the index.
[0104] When a release operation causes two data entries adjacent to each other in physical addresses to be free, the two data entries are merged into one free entry, and the front and back pointers are modified accordingly.
[0105] Embodiment 2:
[0106] The present invention provides a persistent memory data storage system for distributed databases, which is applied to a storage engine and multiple access devices, including a data storage module, a space management module, a key value index module and a data access interface, and is used to implement the storage of persistent memory data through the method disclosed in Example 1.
[0107] The data storage module is used to store persistent data in PMem in a block-aligned manner through the storage engine, load the blocks into the data mapping area of DRAM in sequence according to the block alignment size, and record the saved persistent data and the number of blocks in the meta-information; record the created blocks in DRAM through a list structure; and determine the validity of data entries in the thread access area in PMem.
[0108] In this embodiment, it includes two parts, DRAM and PMem, wherein the PMem part sequentially stores PMem in a block-aligned manner, and the DRAM part maps the contents in PMem, while recording the saved data and the number of blocks in the meta-information. Specifically, the PMem part divides the PMem storage space into blocks aligned at fixed-size addresses, and each block is further divided into equal access areas, the size of which depends on the degree of concurrency supported for concurrent access and the average size of key-value data; each access area contains a lock flag, a freelist local entry pointer, and the rest of the part stores data entries. The DRAM part can timely reflect the persistent content on PMem and save read overhead through mmap mapping; the list structure is used to record the currently created blocks, and when the access device does not have enough space to write new entries, the module is responsible for creating new blocks and expanding the storage space. When the system is restored, the PMem mapping file is read first to determine the first address of the storage space, and then the subsequent blocks are read in sequence and remapped.
[0109] As a specific implementation, the data storage module is used to divide the storage space of PMem into blocks aligned with fixed-size addresses, and divide each block into an equal number of thread access areas. The size of each thread access area depends on the concurrency level supported for concurrent access and the average size of the key-value data. Each thread access area is used to store data entries as well as a lock flag and a freelist local entry pointer.
[0110] Data entries are aligned to 8 bytes in PMem, including an 8-byte header field, a data field in the middle, and an 8-byte footer field at the end. Both the header field and the footer field are used to record the length and valid identifier of the data entry. When the header and footer are inconsistent, it means that the data entry is invalid.
[0111] The lock flag is saved at the starting position of the corresponding thread access area and is used to mark whether there is any accessor in the corresponding thread access area that is being rewritten. The lock flag is based on the thread access area, and multiple accessors can be assigned the same number to rewrite data in different thread access areas.
[0112] Each thread access area stores the address of the first free data entry belonging to the thread access area, and the first free data entry address is used as a local entry of the freelist so that the newly released free data entry can be inserted into the freelist of the thread access area nearby.
[0113] DRAM reflects the contents in PMem through mmap mapping and records the currently created blocks through a list structure. When there is not enough space in the accessor to write new data entries, the data storage module is used to create new blocks and expand the storage space. When the storage system is restored, it is used to read the PMem mapping file, determine the first address of the storage space, and read subsequent blocks in sequence and remap them.
[0114] The data storage module is used to determine the validity of each data item in the thread access area through the following steps:
[0115] (1) Verify the validity of the freelist pointers before and after the freelist of the idle data entry: If it is invalid, it means that the data entry pointed to has not been updated when the pointer is merged and the power is turned off, then it is directly modified to the starting address of the target data entry;
[0116] (2) There are two consecutive free data entries: This indicates that the data entry is not written after allocation and the power is turned off, the content of the data entry is discarded, and the latter data entry is merged forward;
[0117] (3) Verify the consistency of the header and footer: Since the header is updated earlier than the footer in the data entry, and the header is updated when the key value is written or the data entry is ready to be released, the status of the header is used as the basis for recovery. If the header indicates that the data entry is free, it means that the release is not completed, and the data entry continues to be released; if the data entry has been allocated, the footer is updated and regarded as a valid data entry and added to the index as usual;
[0118] (4) Verify whether the key of the current data entry already exists: According to the mechanism of alternating the headers of the new and old data entries in the insertion operation, the header and footer of the old data entry must be in a state of inconsistency before the old data entry is released, and the header is marked as released. During recovery, the release operation can be completed according to the rules for verifying the consistency of the header and footer. In the case where the header of the old data entry is not marked as released and the power is cut off, the footer of the new data entry has not been modified. Therefore, during recovery, the data entry whose footer has not been modified shall prevail and the other one shall be deleted.
[0119] The space management module is used to organize the free data entries that have been released on the PMem into an implicit linked list freelist, and perform operation management operations on the data entries, and the operation management includes allocation, release and merging.
[0120] In this embodiment, the data fields of the free data entry are filled with forward and backward pointers, and are organized into a freelist with other free entries. The forward and backward pointers of the free data entry are stored after the header.
[0121] The allocation operation traverses the freelist and selects an entry of appropriate length, intercepts it as needed, modifies the allocation identifier and the pointers to the previous and next entries, and then hands it over to the accessor to write the data.
[0122] The release operation is similar to logical deletion, which directly modifies the allocation flag, inserts into the freelist, and adds the front and back pointers;
[0123] The merge operation is used to merge two free blocks with adjacent physical addresses and is triggered after the entry is released.
[0124] When the system is restored, each field of the entry is used as the basis for determining integrity.
[0125] Among them, when the free data entries are reallocated, the space management module is used to execute: traverse according to the front and back pointers, all the free data entries in the thread access area corresponding to the same number form the same freelist, and are provided to the same access device for access; the freelists between the thread access areas corresponding to different numbers do not overlap to ensure thread isolation.
[0126] The key-value index module is used to map the key to the address of the data entry. The DRAM records the key-to-address mapping corresponding to each data entry through the index to support query through the search tree structure; it is used to locate the data entry and query through the search tree structure.
[0127] This embodiment adopts a radix tree index, which can support an operation of sequentially scanning a given key range.
[0128] The data access interface is used to accept read and write requests from upper-level applications. Based on the read and write requests, a number is assigned to the accessor through the storage engine. The number corresponds to a thread access area that the accessor can rewrite. The thread access area stores data entries and meta-information and is used to create accessor instances to perform data operations. Data operations include read operations, write operations, and delete operations. The write operations and delete operations are limited by the number of the thread access area, and can only write or delete data entries within the same number. The read operation is not limited.
[0129] The accessor is the medium for the upper-layer application to access the storage system, and its function is similar to that of the client. The accessor is an object in memory, which needs to save an ID field, an operation descriptor, and the key to be accessed; for write operations, the value to be written is saved additionally, and for range scan operations, the end key is saved in addition to the start key.
[0130] The application of the system of this embodiment in the multi-version concurrent control system of the open source database CockroachDB is taken as an example. Figure 4 As shown in the figure, the system realizes the persistent memory storage of intent in the concurrent control system. The implementation principle is: by reconstructing the two key-value pairs of the intent, they are stored in the PMem key-value storage engine in the form of a single key-value pair; by improving the reading, writing, and intention resolution logic of the original multi-version concurrent control system, the intent is transferred to the PMem device for storage, and only after successful submission is the intent converted into a formal version and stored on the disk, which greatly reduces the disk access frequency of the intent.
[0131] The present invention is shown and described in detail above through the accompanying drawings and preferred embodiments. However, the present invention is not limited to these disclosed embodiments. Based on the above multiple embodiments, those skilled in the art can know that the code review methods in the above different embodiments can be combined to obtain more embodiments of the present invention, and these embodiments are also within the protection scope of the present invention.
Claims
1. A persistent memory data storage method for distributed databases, It is characterized in that Applied to a storage engine and multiple accessors, the storage engine includes DRAM and PMem, the DRAM is used to map the content in PMem, and records the saved data and the number of blocks in the meta information, the storage space of PMem is divided into multiple blocks, each block is divided into equal thread access areas, the accessor is a memory object, corresponding to an access thread, and each accessor accesses multiple thread access areas with the same number; The method comprises the following steps: The storage engine stores persistent data in PMem in a block-aligned manner, loads blocks into the data mapping area of DRAM in sequence according to the block alignment size, and records the saved persistent data and the number of blocks in the meta information. Record the created blocks in DRAM through a list structure; Validity determination is performed on data entries in the thread access area in PMem, and for data entries that pass the determination, key-to-address mapping is performed on the data entries, and the key-to-address mapping corresponding to each data entry is recorded in DRAM through an index to support query through a search tree structure; The free data entries released on PMem are organized into an implicit linked list freelist, and operation management operations are performed on the data entries, wherein the operation management includes allocation, release and merging; Create an accessor instance, receive read and write requests initiated by the upper-layer application through the storage instance, and assign a number to the accessor through the storage engine based on the read and write requests, wherein the number corresponds to a thread access area in PMem that the accessor can rewrite, wherein the thread access area stores data entries and meta-information, and performs data operations based on the read and write requests, wherein the data operations include read operations, write operations, and delete operations, wherein the write operations and delete operations are limited by the number of the thread access area, and can only perform write operations or delete operations on data entries within the same number, and the read operation is not limited.
2. The method for storing persistent memory data in a distributed database according to claim 1, It is characterized in that Each thread access area is used to store data entries as well as a lock flag and a freelist local entry pointer; The data entry is aligned with 8 bytes in PMem, including an 8-byte header field, a data field in the middle, and an 8-byte footer field at the end. The header field and the footer field are both used to record the length and valid identifier of the data entry. When the header and footer are inconsistent, it indicates that the data entry is invalid. The lock mark is stored at the starting position of the corresponding thread access area, and is used to mark whether the corresponding thread access area is being rewritten by an accessor. The lock mark is based on the thread access area, and multiple accessors can be assigned the same number to rewrite data in different thread access areas. Each thread access area stores the address of the first free data entry belonging to the thread access area, and the first free data entry address is used as a local entry of the freelist so that the newly released free data entry can be inserted into the freelist of the thread access area nearby.
3. The method for storing persistent memory data in a distributed database according to claim 2, It is characterized in that Determine the validity of each data entry in the thread access area, including the following steps: Verify the validity of the freelist pointers before and after the free data entry: If they are invalid, it means that the data entry pointed to has not been updated when the pointer is merged and the power is turned off, then it is directly modified to the starting address of the target data entry; There are two consecutive free data entries: This indicates that the data entries are not written after allocation and the power is turned off, the contents of the data entries are discarded, and the latter data entries are merged forward; Verify the consistency of header and footer: Since the header is updated earlier than the footer in the data entry, and the timing of updating the header is to complete the key value writing or prepare to release the data entry, the status of the header shall prevail during recovery. If the header indicates that the data entry is free, it means that the release is not completed, and the data entry continues to be released; if the data entry has been allocated, the footer is updated and regarded as a valid data entry, and added to the index as usual; Verify whether the key of the current data entry already exists: According to the mechanism of alternating the headers of the new and old data entries in the insertion operation, the header and footer of the old data entry must be inconsistent before the old data entry is released, and the header is marked as released. During recovery, the release operation can be completed according to the rules for verifying the consistency of the header and footer. In the case where the power is cut off when the header of the old data entry is not marked as released, the footer of the new data entry has not been modified. Therefore, during recovery, the data entry whose footer has not been modified will be used as the basis, and the other one will be deleted.
4. The method for storing persistent memory data in a distributed database according to claim 2, It is characterized in that The data field of the free data entry is filled with forward and backward pointers, and is organized into a freelist with other free entries. The forward and backward pointers of the free data entry are stored after the header; The allocation operation for free data entries is as follows: select an entry of appropriate length from the freelist, intercept it as needed, modify the allocation identifier and the pointers of the previous and next entries, and then hand it over to the accessor to write the data; The release operation for free data entries is: directly modify the allocation flag, insert into the freelist and add the front and back pointers; The merging operation for free data entries is: merging two free blocks with adjacent physical addresses, which is triggered after the entry is released; Among them, when the free data entries are reallocated, they are traversed according to the front and back pointers. All the free data entries in the thread access area corresponding to the same number form the same freelist and are provided to the same accessor for access. The freelists between the thread access areas corresponding to different numbers do not overlap to ensure thread isolation.
5. The method for storing persistent memory data for a distributed database according to any one of claims 1 to 4, It is characterized in that When the accessor performs a write operation, it accesses the thread access area with the corresponding number in the storage engine, traverses the freelist to obtain free data entries of appropriate length, cuts off the excess parts and adds them back to the freelist, writes the key-value pair content into the data entry, marks the length and allocation flag, modifies the pointers of the predecessor and successor elements of the freelist, and finally inserts the new data entry into the index; When the accessor performs a read operation, it directly searches the data entry address of the target key-value pair in PMem from the index and then reads it from PMem. If the search fails, it means that the key-value pair does not exist. Since the thread access area is aligned by address, the number to which the data entry belongs can be directly determined by the thread access area address. Before modifying or deleting the old value, the number can be determined by a read operation, and then the number is assigned to the accessor to perform the operation; When a known numbered accessor deletes a key-value pair in the corresponding area, it directly marks the data entry as free, inserts the local entry into the freelist, and removes the key-value pair from the index; When a release operation causes two data entries adjacent to each other in physical addresses to be free, the two data entries are merged into one free entry, and the front and back pointers are modified accordingly.
6. A persistent memory data storage system for distributed databases, It is characterized in that Applied to a storage engine and a plurality of accessors, and used to implement storage of persistent memory data by using the persistent memory data storage method for a distributed database as described in any one of claims 1 to 5; The system comprises: A data storage module, the data storage module is used to sequentially store persistent data in PMem in a block-aligned manner through a storage engine, load the blocks into the data mapping area of DRAM in sequence according to the block alignment size, and record the saved persistent data and the number of blocks in the meta information; record the created blocks in the DRAM through a list structure; and perform validity determination on data entries in the thread access area in PMem; A space management module, the space management module is used to organize the free data entries released on the PMem into an implicit linked list freelist, and perform operation management operations on the data entries, the operation management including allocation, release and merging; A key-value index module, which is used to map the key to the address of the data entry. The DRAM records the key-to-address mapping corresponding to each data entry through an index to support querying through a search tree structure; and is used to locate the data entry and query through the search tree structure; A data access interface, the data access interface is used to accept read and write requests from upper-layer applications, and based on the read and write requests, a number is assigned to the accessor through the storage engine, the number corresponds to a thread access area that the accessor can rewrite, the thread access area stores data entries and meta-information, and is used to create accessor instances to perform data operations, the data operations include read operations, write operations, and delete operations, the write operations and delete operations are limited by the number of the thread access area, and can only write or delete data entries within the same number, and the read operation is not limited.
7. The persistent memory data storage system for distributed databases according to claim 6, It is characterized in that The data storage module is used to divide the storage space of PMem into blocks aligned according to fixed-size addresses, and divide each block into equal thread access areas, the size of each thread access area depends on the concurrency degree of concurrent access supported and the average size of key-value data, and each thread access area is used to store data entries and a lock flag and a freelist local entry pointer; The data entry is aligned with 8 bytes in PMem, including an 8-byte header field, a data field in the middle, and an 8-byte footer field at the end. The header field and the footer field are both used to record the length and valid identifier of the data entry. When the header and footer are inconsistent, it indicates that the data entry is invalid. The lock mark is stored at the starting position of the corresponding thread access area, and is used to mark whether the corresponding thread access area is being rewritten by an accessor. The lock mark is based on the thread access area, and multiple accessors can be assigned the same number to rewrite data in different thread access areas. Each thread access area stores the address of the first free data entry belonging to the thread access area, and the first free data entry address is used as a local entry of the freelist, so as to insert the newly released free data entry into the freelist of the thread access area nearby; The DRAM reflects the content in PMem by mmap mapping, and records the currently created blocks by list structure. When there is not enough space in the accessor to write new data entries, the data storage module is used to create new blocks and expand the storage space. When the storage system is restored, it is used to read the PMem mapping file, determine the first address of the storage space, and read the subsequent blocks in sequence and remap them. The data storage module is used to determine the validity of each data item in the thread access area through the following steps: Verify the validity of the freelist pointers before and after the free data entry: If they are invalid, it means that the data entry pointed to has not been updated when the pointer is merged and the power is turned off, then it is directly modified to the starting address of the target data entry; There are two consecutive free data entries: This indicates that the data entries are not written after allocation and the power is turned off, the contents of the data entries are discarded, and the latter data entries are merged forward; Verify the consistency of header and footer: Since the header is updated earlier than the footer in the data entry, and the timing of updating the header is to complete the key value writing or prepare to release the data entry, the status of the header shall prevail during recovery. If the header indicates that the data entry is free, it means that the release is not completed, and the data entry continues to be released; if the data entry has been allocated, the footer is updated and regarded as a valid data entry, and added to the index as usual; Verify whether the key of the current data entry already exists: According to the mechanism of alternating the headers of the new and old data entries in the insertion operation, the header and footer of the old data entry must be inconsistent before the old data entry is released, and the header is marked as released. During recovery, the release operation can be completed according to the rules for verifying the consistency of the header and footer. In the case where the power is cut off when the header of the old data entry is not marked as released, the footer of the new data entry has not been modified. Therefore, during recovery, the data entry whose footer has not been modified will be used as the basis, and the other one will be deleted.
8. The persistent memory data storage system for distributed databases according to claim 6, It is characterized in that The data field of the free data entry is filled with forward and backward pointers, and is organized into a freelist with other free entries. The forward and backward pointers of the free data entry are stored after the header; The allocation operation traverses and selects an entry of suitable length from the freelist, intercepts it as needed, modifies the allocation identifier and the pointers of the previous and next entries, and then hands it over to the accessor to write data; The release operation directly modifies the allocation flag, inserts into the freelist, and adds the front and back pointers; The merge operation is used to merge two free blocks with adjacent physical addresses, which is triggered after the entry is released; Among them, when the free data entries are reallocated, the space management module is used to execute: traverse according to the front and back pointers, all the free data entries in the thread access area corresponding to the same number form the same freelist, and are provided to the same access device for access; the freelists between the thread access areas corresponding to different numbers do not overlap to ensure thread isolation.
9. The persistent memory data storage system for distributed databases according to claim 6, It is characterized in that The key-value index module is used to rebuild by scanning when the system restarts.
10. The persistent memory data storage system for distributed databases according to claim 6, It is characterized in that The data access interface provides three operation interfaces, namely, a read interface, a write interface and a delete interface, and each operation interface is allocated with an accessor to undertake read and write operations; The read interface includes single read and range read, and the ordered access within the range is realized by pre-order traversal search of the radix tree index; When calling the accessor to perform a read operation, the data entry address of the target key-value pair in PMem is directly searched from the index, and then read from PMem. If the search fails, it means that the key-value pair does not exist; The write interface is used to call the allocation operation of the space management module to allocate entries for new data and write data, and to determine whether there is an old version of the corresponding key through a pre-read operation, and to perform an additional deletion operation on the old version; When the accessor is called to perform a write operation, it accesses the thread access area with the corresponding number in the storage engine, traverses the freelist to obtain free data entries of appropriate length, cuts off the excess parts and adds them back to the freelist, writes the key-value pair content to the data entry, marks the length and allocation flag, modifies the pointers of the predecessor and successor elements of the freelist, and finally inserts the new data entry into the index; The deletion interface uses a lightweight logical deletion method to call the release operation and merge operation of the space management module to reclaim the storage space; Since the thread access area is aligned by address, when performing a delete operation, the number to which the data entry belongs can be directly determined by the thread access area address. Before modifying or deleting the old value, the number can be determined by a read operation, and then the number is assigned to the accessor to perform the operation; When a known numbered accessor deletes a key-value pair in the corresponding area, it directly marks the data entry as free, inserts the local entry into the freelist, and removes the key-value pair from the index.
Citation Information
Patent Citations
Distributed block storage data processing method and device, equipment and storage medium
CN114443364A
Key Value Store Snapshot in a Distributed Memory Object Architecture
US20200042496A1