Separate key-value storage system, storage control method thereof, and electronic device

By implementing the space reuse method and pipeline writing strategy in the separated key-value storage system, the problem of memory space waste in LSM-Tree is solved, the memory utilization and writing efficiency are improved, the reading performance is optimized, and the stability and consistency of the system are achieved.

CN118860276BActive Publication Date: 2025-10-03HUAZHONG UNIV OF SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410819674.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-24
Publication Date
2025-10-03
Estimated Expiration
2044-06-24

AI Technical Summary

Technical Problem

In the existing technology, the separated memory architecture has the problems of memory space waste and low utilization under write-intensive loads, especially in the LSM-Tree key-value storage system, where invalid data occupies space due to old values ​​not being overwritten in time.

Method used

Memory nodes are accessed through the RDMA network, and a spatial multiplexing method is implemented. Invalid data is merged and recorded in the invalid data mapping table. The invalid data mapping table is used for spatial multiplexing. Combined with the vLog write cache and the Data Block write cache, asynchronous RDMA unilateral write operation and pipeline write strategy are adopted to optimize the data writing and reading processes.

Benefits of technology

It effectively reduces memory space waste, improves memory utilization, enhances write efficiency and concurrent processing capabilities, optimizes read performance and system stability, and ensures data consistency and integrity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118860276B_ABST
    Figure CN118860276B_ABST
Patent Text Reader

Abstract

The present application belongs to the technical field related to information storage, and discloses a separate key-value storage system and a method for controlling its storage, as well as an electronic device. The control method performs spatial reuse on memory nodes during data writing, including: saving the Key and vLog pointers that locate invalid Values ​​in an invalid data mapping table; when writing a new key-value pair, determining whether there is an available Chunk with an available space size that meets the expected size, and if so, overwriting the useless data with the Value according to the invalid data mapping table; selecting a Chunk with an invalid Value for rewriting according to the invalid data mapping table, continuously storing the valid Value in the selected Chunk into a new Chunk, and modifying the SST file in situ to release the original Chunk. Achieving spatial reuse of invalid data by the above method can avoid waste of memory space and improve the utilization of memory space.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the technical field related to information storage, and more specifically, relates to a separate key-value storage system and a storage control method and electronic device thereof. Background Art

[0002] The resource separation architecture can abstract the different components of the system into independent resources, such as computing resources, memory resources, and storage resources, thereby improving the maintainability and scalability of the system. The memory disaggregation architecture is a typical resource separation architecture and an emerging trend in modern data centers. It can support the independent and elastic expansion of computing resources and memory resources. The core idea of ​​the memory disaggregation architecture is to decouple computing nodes and memory nodes. In the resource separation architecture, especially the memory disaggregation architecture, the computing resource pool and the storage or memory resource pool are often connected using the RDMA high-speed network. The application layer reads, writes, and controls the memory of the memory node by calling various primitives (verbs) provided by the RDMA protocol.

[0003] LSM-Tree (Log-Structured Merge-Tree) is a commonly used key-value storage structure, widely used in distributed and local key-value storage systems. A characteristic of LSM-Tree is that when data is written, it is first appended to a sequentially written log file (called a write-ahead log, or WAL) rather than directly written to the tree structure. This fully utilizes the efficient performance of sequential writes. As write operations proceed, data gradually accumulates in memory. When a certain threshold is reached, a data merge operation is triggered. In this merge operation, the data in memory is merged and sorted with the data on disk to generate a new ordered data file. This merge sort process effectively reduces the number of random writes to the disk and improves write performance. For write-intensive workloads, LSM-Tree-based KV storage has become mainstream.

[0004] With the development of information technology, data in different fields are showing a rapid growth trend. How to reduce space waste and improve memory utilization is a technical problem that needs to be solved urgently. Summary of the Invention

[0005] In response to the above-mentioned defects or improvement needs of the prior art, the present application provides a separate key-value storage system and its storage control method and electronic device, the purpose of which is to reduce space waste and improve memory utilization.

[0006] To achieve the above objectives, according to one aspect of the present application, a method for manipulating a separate key-value store is provided, which is executed on a computing node. The method includes accessing a memory node through an RDMA network, writing the value in a write instruction key-value pair {Key, Value} to a value log file of the memory node, and writing the key and vLog pointer used to locate the value into an ordered string table of the memory node. The memory node is spatially reused during data writing. The steps of performing spatial reuse specifically include:

[0007] Merge multiple ordered string tables of memory nodes, save the key and vLog pointer of the invalid value in the invalid data mapping table and delete it from the ordered string table. The invalid value is the old value in the stored key-value pair {Key, Value} that needs to be overwritten by the new value;

[0008] When writing new key-value pairs to the memory node in batches, determine the size of the available space in each granularity Chunk of the value log file, including the space occupied by invalid Values, and determine whether there is an available Chunk with the expected available space size in the existing Chunk. If so, write the Value in the key-value pair into the available space in the available Chunk to overwrite the useless data according to the invalid data mapping table. Otherwise, write the Value in the key-value pair into a new Chunk; store the Key and vLog pointer that locate the newly written Value into the merged ordered string table;

[0009] According to the invalid data mapping table, select the chunk with invalid value to rewrite, store the valid value in the selected chunk continuously into the new chunk, modify the vLog pointer of the corresponding key in the ordered string table, and release the original chunk to store the value of the new key-value pair.

[0010] In some embodiments, the ordered string table includes a plurality of Data Blocks, each DataBlock storing a plurality of Keys and associated vLog pointers;

[0011] The process of writing data includes:

[0012] Allocate vLog write cache and Data Block write cache;

[0013] Store the key-value pair {Key, Value} into the memory table ImmTable;

[0014] The following refresh operations are performed regularly through the background thread:

[0015] Select ImmTable and perform key-value separation on the key-value pairs stored in it;

[0016] Separate Values ​​are stored sequentially in the vLog write cache, and separate Keys are stored sequentially in the Data Block write cache.

[0017] Whenever the number of newly added values ​​in the vLog write cache reaches the preset batch size, an asynchronous RDMA unilateral write operation is triggered to sequentially write the values ​​in the vLog write cache to the value log file of the memory node;

[0018] The key and the associated vLog pointer in the Data Block write cache are written sequentially to the ordered string table of the memory node. After all the data in the Data Block write cache is written to the ordered string table, a synchronous blocked polling check is performed to ensure that all the data in the Data Block write cache and the vLog write cache are written to the memory node.

[0019] In some embodiments, the ordered string table includes a plurality of Data Blocks, each DataBlock storing a plurality of Keys and associated vLog pointers;

[0020] The control method further includes reading key-value pairs within a specified key range from a memory node in a scanning manner according to a read instruction. The data reading process includes:

[0021] Allocate Data Block read cache and vlog read cache;

[0022] Call the Seek operation to locate the scan start key of the ordered string table;

[0023] Read the data of the ordered string table from the starting key through RDMA asynchronous IO and store it into the Data Block read cache;

[0024] Whenever the user calls the Next method to obtain the next key-value pair {Key, Value} within the specified Key range, it first determines whether the data block to which the Key belongs in the Data Block read cache is ready. If not, it blocks and waits until the corresponding Data Block is transferred from the ordered string table to the Data Block read cache. If it is ready, it gradually stores the Value in the value log file into the vlog read cache through RDMA asynchronous IO according to the location pointed to by the Data Block read cache, and determines whether the Value corresponding to the Key in the vlog read cache is ready. If not, it blocks and waits until the Value corresponding to the Key is transferred from the value log file to the vLog read cache. The found Value and the corresponding Key form a key-value pair and return it to the user.

[0025] In some embodiments, the ordered string table includes a filter block, an index block, multiple data blocks, and a perfect hash. The filter block is used to determine whether the required key exists in the ordered string table to which it belongs. The index block is used to locate the data block where the required key is located. The perfect hash is used to locate the specific position of the required key in the data block.

[0026] The control method further includes reading the required key-value pairs from the memory node in a single-point query manner according to the read instruction, and the data reading process includes:

[0027] Query the Filter Block to locate the ordered string table where the required key is located;

[0028] Perform a binary search on the Index Block in the located ordered string table to locate the Data Block in the ordered string table;

[0029] Query the perfect hash in the located ordered string table to locate the required key in the data block;

[0030] Trigger an asynchronous RDMA unilateral read operation, read the corresponding value from the value log file based on the key located by the data block and the associated vLog_ptr, and return the found value and the corresponding key as a key-value pair to the user.

[0031] In some embodiments, the writing process adopts the following pipeline mode:

[0032] Writing batches of key-value pairs corresponding to each Data Block consists of four steps: the first step is to write the value log file, the second step is to write the Data Block data to the ordered string table, the third step is to build the perfect hash of the current Data Block data, and the fourth step is to release the Data Block write cache;

[0033] The key-value pairs are written into the memory nodes in batches according to different Data Blocks. When the nth step of writing is executed for the previous batch of key values, the n-1th step of writing is executed for the next batch of key-value pairs, where n=2, 3, and 4.

[0034] In some embodiments, the memory space of the memory node is divided into a plurality of granularity pools ChunkPool according to a set granularity, and each ChunkPool has a plurality of Chunks;

[0035] The control method also includes: when data needs to be written into a new Chunk in the memory node, first applying for a ChunkPool from the memory node, then managing the Chunks in the applied ChunkPool, and writing the new data.

[0036] In some embodiments, the control method further includes performing the following RDMA connection management:

[0037] A separate QP connection pool is used to manage all QP connections that implement memory node access. When any thread initiates an RDMA request, the current thread applies for F QP connections from the QP connection pool:

[0038] If the number of available QP connections in the QP connection pool is greater than or equal to F, then F available QP connections are directly selected from the QP connection pool;

[0039] If the number of available QP connections in the QP connection pool is less than F, and the number of global QP connections is less than the connection limit, all available QP connections are selected from the QP connection pool and new QP connections are established. The total number of newly created QP connections and available QP connections selected from the QP connection pool is F.

[0040] If the number of available QP connections in the QP connection pool is less than F and the number of global QP connections reaches the connection upper limit, wait for the QP connections already requested by other threads to complete their tasks and return to the QP connection pool until the number of available QP connections in the QP connection pool increases to F, and then select F available QP connections from the QP connection pool.

[0041] In some embodiments, the control method further comprises:

[0042] When an RDMA request is initiated, each RDMA request is split according to the number of QP connections applied for and the actual IO size, so that the split RDMA requests can transmit data in parallel through multiple QP connections applied for.

[0043] According to another aspect of the present application, an electronic device is provided, comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of any of the above methods when executing the computer program.

[0044] According to another aspect of the present application, a separate key-value storage system is provided, comprising a computing node and a memory node connected via an RDMA network, wherein the computing node is the above-mentioned electronic device.

[0045] In general, the above technical solutions conceived by this application have the following beneficial effects compared with the prior art:

[0046] 1. The control method of the separated key-value storage provided by the present invention, during the data writing period, records invalid data by merging the SST files and records them in the invalid data mapping table, and the space occupied by the invalid data in the vLog file can be reused as available space. On the one hand, based on the invalid data mapping table, the space size of the invalid data in each Chunk in the vLog file can be quickly judged. When writing the key-value pair again, if there is available space that meets the conditions in the existing Chunk, the Value is directly written to the position of the invalid data to achieve the reuse of the invalid data space. On the other hand, based on the invalid data mapping table, the valid data and invalid data in the Chunk can be quickly determined, the valid data is written into the new Chunk and the SST file is modified in situ to release the space of the original Chunk to achieve space reuse. The reuse of invalid data space is achieved by the above method, which can avoid the waste of memory space and improve the utilization rate of memory space.

[0047] 2. Furthermore, combining vLog write cache and Data Block write cache to implement data writing can significantly improve the system's write efficiency and concurrent processing capabilities; the key-value separation strategy reduces the latency of foreground writes and optimizes background write performance through asynchronous RDMA unilateral write operations and batch processing strategies; at the same time, through the thread-specific vLog cache mechanism, competition between threads is avoided, further improving write efficiency; synchronous polling checks ensure the consistency and integrity of SST files and vLog data during the refresh process; Chunk management of different granularities enables the system to achieve a balance between memory utilization and metadata maintenance, effectively improving the overall performance and stability of the system.

[0048] 3. Furthermore, the use of RDMA asynchronous I / O enables phased pre-reading of data blocks and vLogs, reducing blocking during the read process. This phased pre-reading not only effectively utilizes network bandwidth resources but also avoids unnecessary bandwidth waste during small-scale scans by limiting the number of pre-read key-value pairs and the total value size. Adjacent vLog_ptrs are consolidated and pre-fetched as larger read / write units, improving network resource utilization. A retry mechanism verifies vLog data inconsistencies caused by garbage collection, ensuring the accuracy and consistency of read data.

[0049] 4. Furthermore, the use of a pipelined write method can improve the overall data writing efficiency of the system. By dividing the writing process into multiple steps and adopting a pipeline mode, each data block immediately triggers the previous processing step of the next data block when it completes one step, processing multiple tasks in parallel, thus maximizing the utilization of the system's computing and transmission resources. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] Figure 1 1 is a schematic diagram of the structure of a separate key-value storage system in one embodiment of the present invention;

[0051] FIG2( a ) is a flowchart of the steps of the spatial multiplexing method according to one embodiment of the present invention;

[0052] FIG2( b ) is a flow chart of SST file merging to generate invalid data in one embodiment of the present invention;

[0053] FIG2( c ) is a flowchart of a garbage collection process according to an embodiment of the present invention;

[0054] FIG3( a ) is a schematic diagram of key-value separation during refresh in one embodiment of the present invention;

[0055] FIG3( b ) is a schematic diagram of key-value separation during refresh in one embodiment of the present invention;

[0056] Figure 4 is a schematic diagram of realizing data reading by scanning in one embodiment of the present invention;

[0057] FIG5( a ) is a schematic diagram of constructing an SST file in one embodiment of the present invention;

[0058] FIG5( b ) is a schematic diagram of a single-point query method based on perfect hashing in one embodiment of the present invention;

[0059] Figure 6 is a schematic diagram of writing in pipeline mode in one embodiment of the present invention;

[0060] Figure 7Schematic diagram of memory partitioning of memory nodes in one embodiment of the present invention;

[0061] Figure 8 Schematic diagram of an RDMA connection management method in one embodiment of the present invention;

[0062] Figure 9 Schematic diagram of an RDMA IO management method in one embodiment of the present invention. DETAILED DESCRIPTION

[0063] In order to make the purpose, technical solutions and advantages of this application more clearly understood, this application is further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are merely for explaining this application and are not intended to limit this application. In addition, the technical features involved in the various embodiments of this application described below may be combined with each other as long as they do not conflict with each other.

[0064] In order to facilitate understanding of the present invention, the separated key-value storage system rLSM mentioned in the present invention is first introduced. Figure 1 The figure shows a structural diagram of a separate key-value storage system in an embodiment, which includes a computing node and a memory node. The computing node has a memory table MemTable for temporarily storing newly written key-value pair data {Key, Value}. The MemTable in the computing node is locked as ImmTable after the data reaches a certain threshold. When data is refreshed, the data of ImmTable is refreshed to the memory node. The memory node stores the value log file (referred to as vLog file) and the ordered string table (SortedString Table, referred to as SST file); the computing node is connected to the memory node through a remote direct memory access network (RDMA) for writing and reading data. When writing data, the computing node writes the Value in {Key, Value} into the vLog file of the memory node according to the write instruction and writes the Key and vLog pointer vLog_ptr used to locate the Value into the SST file of the memory node. vLog_ptr is a pointer to the value corresponding to each key in the vLog; when reading data, the computing node reads the Value from the vLog file of the memory node and the Key from the SST file of the memory node according to the read instruction, and constructs the key-value pair data {Key, Value} and feeds it back to the user.

[0065] Among them, the SST file generally contains a filter block, an index block, and multiple data blocks. The filter block is used to determine whether the required key exists in the ordered string table to which it belongs. The index block is used to locate the data block where the required key is located. Each data block stores multiple groups of keys and associated vLog pointers.

[0066] When a value is stored in a vLog file, it is usually stored as an entry along with the corresponding key and metadata. Metadata stores the storage size of the key and value and is used to parse the key and value when reading. Therefore, a vLog file has multiple entries.

[0067] Based on the above separated key-value storage system, the present invention proposes the following control method to optimize the scheduling and management of the system data.

[0068] Example 1

[0069] In a key-value storage system with key-value separation, after the data is written into the memory node, the level L K The SST files are merged, compressed and passed to L K+1 , achieving multi-layer data storage. If a user ID stores key-value pairs multiple times, with the same key but different values, the latest value is generally considered valid data, while the older value corresponding to the same key is invalid data. Invalid data remains in the vLog file, occupying memory space. Excessive invalid data can cause serious waste of memory space. Based on this, the present invention proposes the following space reuse method.

[0070] FIG2( a ) is a flowchart of the steps of the spatial multiplexing method in one embodiment of the present invention, which mainly includes steps S11 to S13 . Each step is introduced below.

[0071] Step S11: Merge multiple ordered string tables of memory nodes, save the Key and vLog pointers locating invalid Values ​​in the invalid data mapping table and delete them from the ordered string table. The invalid Value is the old Value in the stored key-value pair {Key, Value} that needs to be overwritten by the new Value.

[0072] As shown in Figure 2(b), this is a flowchart of invalid data generated by SST file merging. Invalid data is automatically generated when performing multi-way merging of multiple SST files. Specifically, when recording the key of each key-value pair, the stored sequence number seq is also recorded. When merging files, for the same key, only the one with the largest sequence number seq is retained. The remaining keys are invalid keys, and the corresponding values ​​are invalid values. Although invalid keys are deleted through compression and merging in the SST file, the corresponding invalid values ​​still exist in the vLog file, resulting in space waste. In a specific embodiment, a value is stored in the vLog file as an entry together with the corresponding key and MetaData. Therefore, the invalid space is actually an invalid entry.

[0073] In order to efficiently manage these invalid data, the present invention will actively record these invalid data during the merging process, and specifically save their corresponding vLog pointers (vLog_ptr) in a global invalid data mapping table (map). By querying the invalid data mapping table, the position of the invalid value in the vLog file can be located.

[0074] Step S12: When writing new key-value pairs to the memory node in batches, determine the size of the available space in each granularity Chunk of the value log file. The available space includes the space occupied by invalid Value. Determine whether there is an available Chunk with the expected available space size in the existing Chunk. If so, write the Value in the key-value pair into the available space in the available Chunk according to the invalid data mapping table to overwrite the useless data. Otherwise, write the Value in the key-value pair into a new Chunk; store the Key and vLog pointer that locate the newly written Value into the merged ordered string table.

[0075] For recorded invalid space, the system treats it as available space when writing new values. In addition to accurately recording the location of invalid space, the system also records the percentage of invalid data within a specific granularity, Chunk, to speed up the search for matching appropriately sized space. This granularity is in Chunks by default.

[0076] During key-value separation, if free space is needed, the system will prioritize searching the chunk with the highest amount of invalid data. If suitable free space is found within the specified number of searches, it will be immediately used for writing. If no suitable free space is found, a new chunk will be written.

[0077] When a new value overwrites an invalid value, the invalid data mapping table can be updated at the same time, and the key and vLog pointer of the invalid value overwritten in the invalid mapping table can be deleted, indicating that the location is already valid data.

[0078] Step S13: Select the Chunk with invalid Value to rewrite according to the invalid data mapping table, store the valid Value in the selected Chunk continuously into the new Chunk, modify the vLog pointer of the corresponding Key in the SST file, and release the original Chunk to store the Value of the new key-value pair.

[0079] This process is called the invalid space recovery process, also known as the garbage collection process.

[0080] As shown in Figure 2(c), this is a flowchart of the garbage collection process in one embodiment of the present invention. Specifically, when the background garbage collection thread performs garbage collection on the selected Chunk, it will traverse the Entry in the Chunk and determine its validity based on the invalid data mapping table. If it is determined to be invalid, there is no need to query the LSM-Tree. If it is valid, a new Chunk is written. Since the position of the Entry in the vLog file changes, the vLog pointer used for positioning in the SST file also needs to be adjusted. Therefore, the LSM-Tree is queried through the key and seq information recorded in the Entry to obtain the position of the corresponding vLog pointer in the SST file and modify the vLog pointer of the SST file in place without rewriting the LSM-Tree.

[0081] Combined with the invalid data mapping table, the need to query the LSM-Tree when checking the validity of key-value pairs during garbage collection can be effectively reduced, thereby improving the overall performance and efficiency of the system.

[0082] It should be noted that the execution order of step S12 and step S13 may not be particular.

[0083] The above method, during data writing, records invalid data by merging SST files and records them in the invalid data mapping table. The space occupied by invalid data in the vLog file can be reused as available space. On the one hand, based on the invalid data mapping table, the space size of invalid data in each Chunk in the vLog file can be quickly judged. When writing the key-value pair again, if there is available space that meets the conditions in the existing Chunk, the Value is directly written to the position of the invalid data to achieve the reuse of the invalid data space. On the other hand, based on the invalid data mapping table, the valid data and invalid data in the Chunk can be quickly determined, the valid data is written to the new Chunk and the SST file is modified in situ to release the space of the original Chunk to achieve space reuse. The reuse of invalid data space is achieved by the above method, which can avoid the waste of memory space and improve the utilization rate of memory space.

[0084] Example 2

[0085] Based on Example 1, this embodiment further provides a specific data writing process, which is as follows:

[0086] Allocate vLog write cache and Data Block write cache;

[0087] Store the key-value pair {Key, Value} into the memory table ImmTable;

[0088] The following refresh operations are performed regularly through the background thread:

[0089] Select ImmTable and perform key-value separation on the key-value pairs stored in it;

[0090] Separate Values ​​are stored sequentially in the vLog write cache, and separate Keys are stored sequentially in the Data Block write cache.

[0091] Whenever the number of newly added values ​​in the vLog write cache reaches the preset batch size, an asynchronous RDMA unilateral write operation is triggered to sequentially write the values ​​in the vLog write cache to the value log file of the memory node;

[0092] The key and the associated vLog pointer in the Data Block write cache are written sequentially to the ordered string table of the memory node. After all the data in the Data Block write cache is written to the ordered string table, a synchronous blocked polling check is performed to ensure that all the data in the Data Block write cache and the vLog write cache are written to the memory node.

[0093] FIG3( a ) is a schematic diagram of key-value separation during refresh in an embodiment of the present invention, and FIG3( b ) is a schematic diagram of asynchronous writing to a vLog file in an embodiment of the present invention.

[0094] Key-value separation means storing the key and value in the key-value store separately, and the separated values ​​are stored in a separate vLog file. The present invention adopts a separation-at-refresh approach. On the one hand, the key-value separation process and vLog writing are handed over to the background thread, reducing the foreground write latency. On the other hand, when faced with intensive update loads, conflicts can be resolved in the MemTable for key-value pairs with the same key, significantly reducing the repeated writing of vLogs with conflicting keys. Key-value separation at refresh also facilitates space management. When the foreground read operation hits the MemTable or the immutable memory table (ImmTable), it can directly obtain the value without reading from the memory node.

[0095] The refresh thread is responsible for refreshing the ImmTable to the memory node and building the SST file. When building the SST file, the value is asynchronously written to the vLog of the memory node. The vLog_ptr pointing to the Entry in the vLog is saved in the SST file.

[0096] Data refreshes are performed using thread-specific vLog caches: Each thread has its own vLog cache area, typically set to a certain number of chunks (16 chunks by default). During write operations, the system first writes the value to the vLog cache of the respective thread. This design avoids conflicts and improves write efficiency.

[0097] Asynchronous RDMA writes to the value log (vLog) cache: When an SST file needs to update a data block or the vLog cache has accumulated sufficient data, the system triggers an asynchronous RDMA unilateral write operation. To optimize write performance, when the batch size is less than a set threshold, the write request is marked as SIGNAL, which helps minimize write latency.

[0098] Write synchronization polling check: To ensure that a single flush completes the SST write operation and all associated vLog records, the system performs a synchronous blocking polling check (poll_cq) on all issued unilateral RDMA write requests at the end of the SST write. This step ensures data consistency and integrity and ensures that all write operations have been successfully written to the disk.

[0099] Similar to SST files, vLogs are managed at the chunk level. New chunks are requested and managed by compute nodes from a chunkpool allocated from remote memory nodes. Considering the additional metadata maintenance overhead during garbage collection, different chunk granularities can be used for value log files and SST files, balancing memory space utilization and metadata maintenance overhead. The vLogs generated by key-value separation by a single refresh thread are written sequentially to the vLog in the order of key-value pairs, thereby accelerating the step-by-step read-ahead performance during range queries.

[0100] In this embodiment, the above writing method is adopted to significantly improve the writing efficiency and concurrent processing capabilities of the system. The key-value separation strategy reduces the latency of foreground writing and optimizes the background writing performance through asynchronous RDMA unilateral write operations and batch processing strategies. At the same time, through the thread-specific vLog cache mechanism, competition between threads is avoided, further improving the writing efficiency. Synchronous polling checks ensure the consistency and integrity of SST files and vLog data during the refresh process. Chunk management of different granularities enables the system to achieve a balance between memory utilization and metadata maintenance, effectively improving the overall performance and stability of the system.

[0101] Example 3

[0102] Based on Example 1, this embodiment further provides a specific data reading process, which uses a scanning method to read the required key-value pairs within a specified range from the memory node. The process is as follows:

[0103] Allocate Data Block read cache and vlog read cache;

[0104] Call the Seek operation to locate the scan start key of the ordered string table;

[0105] Read the data of the ordered string table from the starting key through RDMA asynchronous IO and store it into the Data Block read cache;

[0106] Whenever the user calls the Next method to obtain the next key-value pair {Key, Value} within the specified Key range, it first determines whether the data block to which the Key belongs in the Data Block read cache is ready. If not, it blocks and waits until the corresponding Data Block is transferred from the ordered string table to the Data Block read cache. If it is ready, it gradually stores the Value in the value log file into the vlog read cache through RDMA asynchronous IO according to the location pointed to by the Data Block read cache, and determines whether the Value corresponding to the Key in the vlog read cache is ready. If not, it blocks and waits until the Value corresponding to the Key is transferred from the value log file to the vLog read cache. The found Value and the corresponding Key form a key-value pair and return it to the user.

[0107] like Figure 4 Figure 2 shows a schematic diagram of data reading using a scanning method in one embodiment of the present invention. This embodiment implements a range query acceleration method with step-by-step pre-reading to implement data reading. This method utilizes RDMA asynchronous I / O to pre-read the DataBlock and the vLog it points to in steps, reducing the blocking caused by pre-reading on the critical path.

[0108] Specifically, taking backward Scan as an example, when the Seek operation is first called to locate the starting key of the Scan, subsequent Data Blocks are read through RDMA asynchronous IO, and subsequent keys and vLog_ptr are loaded into the Data Block read cache of the computing node. When performing the Next operation, if there are unloaded values ​​in the Data Block read cache, RDMA asynchronous IO is also used to pre-read the memory area pointed to by vLog_ptr in batches and load them into the cache.

[0109] To reduce waste of RDMA network resources, step-by-step prefetching limits the number of key-value pairs and the total value size to be prefetched. This prevents bandwidth waste and preemption caused by prefetching during small-scale scan operations. Furthermore, because key-value pairs in the vLog corresponding to the same SST file are often stored in logical address order, step-by-step prefetching aggregates vLog_ptrs with adjacent addresses into larger read / write units for prefetching, improving network resource utilization.

[0110] Since vLog may be garbage collected, the vLog data pre-read by Scan may not be accurate. In order to avoid complex concurrent access control, the present invention adopts a simple retry to solve the problem of cache inconsistency. The data obtained by pre-reading will be compared with the checksum recorded in vLog_ptr. If there is any inconsistency, the pre-read data will be discarded and a retry will be performed.

[0111] In this embodiment, the above reading method can effectively improve data reading efficiency. The use of RDMA asynchronous IO realizes the step-by-step pre-reading of Data Block and vLog, reducing the blocking during the reading process. This step-by-step pre-reading not only effectively utilizes network bandwidth resources, but also avoids unnecessary bandwidth waste during small-scale scanning by limiting the number of pre-read key-value pairs and the total value size. Adjacent vLog_ptrs are integrated and pre-fetched as larger read and write units, which improves network resource utilization. For vLog data inconsistencies caused by garbage collection, a retry mechanism is used to verify the accuracy and consistency of the read data.

[0112] Example 4

[0113] As shown in Figure 5(a), this is a schematic diagram of the construction of an SST file in one embodiment of the present invention. The SST file includes a filter block, an index block, multiple data blocks and a perfect hash. The filter block is used to determine whether the required key exists in the ordered string table to which it belongs. The index block is used to locate the data block where the required key is located. The perfect hash is used to locate the specific position of the required key in the data block.

[0114] Perfect hashing is a specialized hash function design whose key characteristic is that, given a unique combination, it can map each input to a unique hash value without collisions. Since it only applies to static data, it aligns with the SST's inherent nature as a static data set, allowing for the acceleration of single-point queries. After key-value separation, the values ​​stored in the SST file become pointers to the vLog. While the size of a single SST file remains constant, the number of key-value pairs in the SST increases significantly, leading to a corresponding increase in the binary search time for data blocks during single-point queries. However, since index blocks are built for data blocks, the number of data items in an index block does not increase significantly while the size of a single SST file remains constant. Therefore, constructing a perfect hash for a single data block can accelerate queries against that data block. By default, the build process is performed in the background to minimize its impact on refresh and merge processes.

[0115] The construction of the perfect hash is initiated by the refresh thread and the merge thread, and is constructed when a new SST file is generated. When the refresh thread or the merge thread is in progress, the Data Block will be constructed one by one when constructing the SST, and the Index Block and the Filter Block will be constructed at the same time during the construction of the Data Block. Since the construction process of the perfect hash needs to know all the input keys, it needs to be performed after all the inputs are completed in a Data Block. When the Block configuration is large, it may result in a long construction time. Therefore, the present invention puts the construction process of the perfect hash in the background by default, removes it from the critical path of creating the new SST file, and reduces the impact of the construction of the perfect hash on writing.

[0116] Based on Example 1, this embodiment provides a single-point query acceleration method to achieve single-point data reading. FIG5( b ) is a schematic diagram of a single-point query method based on perfect hashing in one embodiment of the present invention. The specific process is as follows:

[0117] Query the filter block to locate the ordered string table containing the required key: Based on the metadata recorded in the RDMA file system, read the filter block of the SST file. Use the Bloom filter to determine whether the key is likely to exist in the current SST file. If not, return directly; if it is likely to exist, proceed to the next step.

[0118] Perform a binary search on the index block in the located ordered string table to locate the data block in the ordered string table: Perform a binary search on the index block. Perform a binary search on the index block based on the query key to locate the target data block and its index idx;

[0119] Query the perfect hash in the located ordered string table to locate the required key in the data block: according to the sequence number of the SST file and the subscript idx of the data block, obtain its perfect hash and query the target key;

[0120] Trigger an asynchronous RDMA unilateral read operation, read the corresponding value from the value log file based on the key located in the data block and the associated vLog_ptr, and return the found value and the corresponding key as a key-value pair to the user: initiate a unilateral RDMA read based on the calculated offset and key-value pair size. If the key-value pair is not separated, the single-point query ends. If the key-value pair is separated, initiate a new unilateral RDMA read based on the vLog_ptr stored in the value to obtain the actual value.

[0121] The traditional LSM query process requires reading the Filter Block, Index Block, and Data Block in a single SST one by one. Since the Index Block and Data Block need to perform a large number of key comparisons in the binary search, not only does the CPU consumption increase, but the characteristics of the binary search itself will cause repeated cache line misses, further reducing the query speed, which is especially obvious when the key is large or the BlockSize is large. Since the present invention adopts a key-value separation strategy, the size of the value stored in the SST will be greatly reduced, resulting in an increase in the number of data items in a Data Block. Therefore, the present invention uses perfect hashing to speed up the query process of the Data Block during single-point query, thereby speeding up single-point query.

[0122] Example 5

[0123] Based on Example 4, this embodiment provides a pipeline mode writing process, such as Figure 6 FIG. 1 is a schematic diagram of writing in a pipeline mode according to an embodiment of the present invention, specifically including:

[0124] The writing of batch key-value pairs corresponding to each Data Block includes four steps: the first step completes the writing of the value log file, the second step completes the writing of the Data Block data in the ordered string table, the third step completes the construction of the perfect hash of the current Data Block data, and the fourth step releases the Data Block write cache; the key-value pairs are written to the memory node in batches according to different Data Blocks. When the nth step of writing is executed for the previous batch of key values, the n-1th step of writing is executed for the next batch of key-value pairs, where n = 2, 3, and 4.

[0125] The perfect hash construction process adopts a pipeline approach. The data in the Data Block, which is the main body of the construction, will flow through three processes: key-value separation, SST file writing, and perfect hash construction.

[0126] The newly constructed perfect hash will be cached in the compute node memory by default and asynchronously written to the memory node for storage. The local perfect hash cache can be managed according to the specific memory space usage limit and caching strategy, thereby reducing the memory usage of the compute node.

[0127] In this embodiment, the above pipeline writing method can improve the overall data writing efficiency of the system. By dividing the writing process into multiple steps and adopting a pipeline mode, each data block immediately triggers the previous processing step of the next data block when it completes one step, processing multiple tasks in parallel, thus maximizing the utilization of the system's computing and transmission resources.

[0128] Example 6

[0129] The memory space of the memory node is divided into multiple granularity pools ChunkPool according to the set granularity. Each ChunkPool has multiple Chunks. Based on Example 1, this embodiment provides a memory management method, which specifically includes: when data needs to be written to a new Chunk in the memory node, first apply for a ChunkPool from the memory node, then manage the Chunks in the applied ChunkPool, and write new data.

[0130] like Figure 7The figure shows a memory node memory partition diagram in one embodiment of the present invention. The present invention divides the memory space of the memory node into ChunkPools according to a certain granularity (the default is 1GB). Each ChunkPool is an RDMA memory domain, which is used to reduce the metadata management pressure on the memory domain in the RDMA smart network card. It is further divided into Chunks according to the specified granularity (the default is 1MB), which is the minimum unit of space allocation. ChunkPool is allocated on demand and dynamically expanded. When a computing node needs to write new data, it will apply for space from the memory node at the ChunkPool granularity. The space division and use within the ChunkPool are organized by the computing node itself.

[0131] Example 7

[0132] Based on Example 1, this embodiment provides an RDMA connection management method, which specifically includes: using a separate QP connection pool to manage all QP connections that implement memory node access. When any thread initiates an RDMA request, the current thread applies for F QP connections from the QP connection pool:

[0133] If the number of available QP connections in the QP connection pool is greater than or equal to F, then F available QP connections are directly selected from the QP connection pool;

[0134] If the number of available QP connections in the QP connection pool is less than F, and the number of global QP connections is less than the connection limit, all available QP connections are selected from the QP connection pool and new QP connections are established. The total number of newly created QP connections and available QP connections selected from the QP connection pool is F.

[0135] If the number of available QP connections in the QP connection pool is less than F and the number of global QP connections reaches the connection upper limit, wait for the QP connections already requested by other threads to complete their tasks and return to the QP connection pool until the number of available QP connections in the QP connection pool increases to F, and then select F available QP connections from the QP connection pool.

[0136] like Figure 8Figure 1 shows a schematic diagram of an RDMA connection management method according to an embodiment of the present invention. This method utilizes a separate QP connection pool to manage all connections. When a foreground or background thread needs to initiate an RDMA request, during a complete operation, such as a complete merge, the current thread requests a certain number of QP connections from the RDMA Manager. These connections are exclusive to the current thread, and its Send Queue (SQ) and Complete Queue (CQ) are guaranteed to be empty. After the thread completes all operations and waits for all events to complete, it returns the requested QP connection to the QP connection pool. If no QP connection is available when a thread requests a connection, and the number of already requested connections globally is less than a threshold, a new QP connection is established between the compute node and the memory node. If the number of already requested connections is greater than the threshold, an already requested connection is reused.

[0137] Example 8

[0138] Based on Example 1, this embodiment provides an RDMA IO management method, which specifically includes: when initiating an RDMA request, splitting each RDMA request according to the number of applied QP connections and the actual IO size, so that the split RDMA requests transmit data in parallel through multiple applied QP connections.

[0139] like Figure 9 Figure 2 shows a schematic diagram of an RDMA IO management method in one embodiment of the present invention. To decouple upper-layer RDMA network resource usage from complex RDMA operation primitives and efficiently utilize RDMA network resources for diverse read and write requirements, the present invention manages all upper-layer RDMA read and write requests through a separate RDMA IO manager. Leveraging optimizations such as asynchronous IO and multiple QPs, the method automatically splits upper-layer read and write operations. This splitting is primarily based on the number of QPs requested by the thread and the actual IO size. While ensuring that each split RDMA read and write request exceeds the minimum threshold required to maximize RDMA network bandwidth performance, it also fully utilizes multiple QPs to reduce data transmission latency for large IOs.

[0140] Example 9

[0141] The present application also relates to an electronic device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the above method when executing the computer program.

[0142] The electronic device may be a desktop computer, a notebook, a PDA, a cloud server, or other electronic devices. The processor may be a central processing unit (CPU), or other general-purpose processors, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The memory may be used to store computer programs and / or modules, and the processor may perform various functions of the electronic device by running or executing the computer programs and / or modules stored in the memory, and calling the data stored in the memory.

[0143] Example 10

[0144] The present application also relates to a separated key-value storage system, characterized in that it includes a computing node and a memory node connected via an RDMA network, and the computing node is the electronic device in Example 9.

[0145] The technical features of the above-described embodiments can be combined in any combination. To simplify the description, not all possible combinations of the technical features in the above-described embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification. It should be noted that the phrases "in one embodiment," "for example," "and another example," etc. in this application are intended to illustrate this application and are not intended to limit this application.

[0146] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the patent application. It should be noted that a person skilled in the art may make various modifications and improvements without departing from the spirit of the present application, and such modifications and improvements are all within the scope of protection of the present application.

Claims

1. A method for controlling a separate key-value store, characterized in that: Executed on a computing node, the manipulation method includes accessing a memory node through an RDMA network, writing the value in the write instruction key-value pair {Key, Value} into the value log file of the memory node, and writing the key and vLog pointer used to locate the value into the ordered string table of the memory node. During the data writing process, the memory node is spatially reused. The spatial reuse steps specifically include: Merge multiple ordered string tables of memory nodes, save the key and vLog pointer of the invalid value in the invalid data mapping table and delete it from the ordered string table. The invalid value is the old value in the stored key-value pair {Key, Value} that needs to be overwritten by the new value; When writing new key-value pairs to the memory node in batches, determine the size of the available space in each granularity Chunk of the value log file, including the space occupied by invalid Values, and determine whether there is an available Chunk with the expected available space size in the existing Chunk. If so, write the Value in the key-value pair into the available space in the available Chunk to overwrite the useless data according to the invalid data mapping table. Otherwise, write the Value in the key-value pair into a new Chunk; store the Key and vLog pointer that locate the newly written Value into the merged ordered string table; According to the invalid data mapping table, select the chunk with invalid value to rewrite, store the valid value in the selected chunk continuously into the new chunk, modify the vLog pointer of the corresponding key in the ordered string table, and release the original chunk to store the value of the new key-value pair.

2. The method for controlling a separate key-value store according to claim 1, wherein: The ordered string table includes multiple data blocks, each of which stores multiple sets of keys and associated vLog pointers; The process of writing data includes: Allocate vLog write cache and Data Block write cache; Store the key-value pair {Key, Value} into the memory table ImmTable; The following refresh operations are performed regularly through the background thread: Select ImmTable and perform key-value separation on the key-value pairs stored in it; Separate Values ​​are stored sequentially in the vLog write cache, and separate Keys are stored sequentially in the Data Block write cache. Whenever the number of newly added values ​​in the vLog write cache reaches the preset batch size, an asynchronous RDMA unilateral write operation is triggered to sequentially write the values ​​in the vLog write cache to the value log file of the memory node; The key and the associated vLog pointer in the Data Block write cache are written sequentially to the ordered string table of the memory node. After all the data in the Data Block write cache is written to the ordered string table, a synchronous blocked polling check is performed to ensure that all the data in the Data Block write cache and the vLog write cache are written to the memory node.

3. The method for controlling a separate key-value store according to claim 1, wherein: The ordered string table includes multiple data blocks, each of which stores multiple sets of keys and associated vLog pointers; The control method further includes reading key-value pairs within a specified key range from a memory node in a scanning manner according to a read instruction. The data reading process includes: Allocate Data Block read cache and vlog read cache; Call the Seek operation to locate the scan start key of the ordered string table; Read the data of the ordered string table from the starting key through RDMA asynchronous IO and store it into the Data Block read cache; Whenever the user calls the Next method to obtain the next key-value pair {Key, Value} within the specified Key range, it first determines whether the data block to which the Key belongs in the Data Block read cache is ready. If not, it blocks and waits until the corresponding Data Block is transferred from the ordered string table to the Data Block read cache. If it is ready, it gradually stores the Value in the value log file into the vlog read cache through RDMA asynchronous IO according to the location pointed to by the Data Block read cache, and determines whether the Value corresponding to the Key in the vlog read cache is ready. If not, it blocks and waits until the Value corresponding to the Key is transferred from the value log file to the vLog read cache. The found Value and the corresponding Key form a key-value pair and return it to the user.

4. The method for controlling a separate key-value store according to claim 2, wherein: The ordered string table includes a filter block, an index block, multiple data blocks, and a perfect hash. The filter block is used to determine whether the required key exists in the ordered string table to which it belongs. The index block is used to locate the data block where the required key is located. The perfect hash is used to locate the specific position of the required key in the data block. The control method further includes reading the required key-value pairs from the memory node in a single-point query manner according to the read instruction, and the data reading process includes: Query the Filter Block to locate the ordered string table where the required key is located; Perform a binary search on the Index Block in the located ordered string table to locate the Data Block in the ordered string table; Query the perfect hash in the located ordered string table to locate the required key in the data block; Trigger an asynchronous RDMA unilateral read operation, read the corresponding value from the value log file based on the key located by the data block and the associated vLog_ptr, and return the found value and the corresponding key as a key-value pair to the user.

5. The method for controlling a separate key-value store according to claim 4, wherein: The writing process of writing new key-value pairs in batches to the memory nodes adopts the following pipeline mode: Writing batches of key-value pairs corresponding to each Data Block consists of four steps: the first step is to write the value log file, the second step is to write the Data Block data to the ordered string table, the third step is to build the perfect hash of the current Data Block data, and the fourth step is to release the Data Block write cache; The key-value pairs are written into the memory nodes in batches according to different Data Blocks. When the nth step of writing is executed for the previous batch of key values, the n-1th step of writing is executed for the next batch of key-value pairs, where n=2, 3, or 4.

6. The method for controlling a separate key-value store according to claim 1, wherein: The memory space of the memory node is divided into multiple granularity pools ChunkPool according to the set granularity, and each ChunkPool has multiple Chunks; The control method also includes: when data needs to be written into a new Chunk in the memory node, first applying for a ChunkPool from the memory node, then managing the Chunks in the applied ChunkPool, and writing the new data.

7. The method for controlling a separate key-value store according to claim 1, wherein: The control method further includes performing the following RDMA connection management: A separate QP connection pool is used to manage all QP connections that implement memory node access. When any thread initiates an RDMA request, the current thread applies for F QP connections from the QP connection pool: If the number of available QP connections in the QP connection pool is greater than or equal to F, then F available QP connections are directly selected from the QP connection pool; If the number of available QP connections in the QP connection pool is less than F, and the number of global QP connections is less than the connection limit, all available QP connections are selected from the QP connection pool and new QP connections are established. The total number of newly created QP connections and available QP connections selected from the QP connection pool is F. If the number of available QP connections in the QP connection pool is less than F and the number of global QP connections reaches the connection upper limit, wait for the QP connections already requested by other threads to complete their tasks and return to the QP connection pool until the number of available QP connections in the QP connection pool increases to F, and then select F available QP connections from the QP connection pool.

8. The method for controlling a separate key-value store according to claim 7, wherein: The control method further includes: When an RDMA request is initiated, each RDMA request is split according to the number of QP connections applied for and the actual IO size, so that the split RDMA requests can transmit data in parallel through multiple QP connections applied for.

9. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 8 are implemented.

10. A separate key-value storage system, characterized in that: It includes a computing node and a memory node connected via an RDMA network, and the computing node is the electronic device according to claim 9.

Citation Information

Patent Citations

  • High-performance and easy-to-expand key value storage method utilizing differential index mechanism

    CN110825748A

  • Key value storage system based on SCM and SSD and read-write request processing method

    CN110968269A