Data-driven hash indexing method based on persistent CPU cache
By building the directory-segment-bucket three-layer hash structure and loop queue design in the persistent CPU cache, the read and write amplification problem of hash indexes is solved, efficient data layout and cache optimization are achieved, and index performance and throughput are improved.
Patent Information
- Application Number
- CN202510280792.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-03-11
AI Technical Summary
In the prior art, the hash index based on persistent CPU cache has read and write amplification problems when frequent access to metadata and scattered hash buckets, and fails to fully utilize the optimization potential of persistent CPU cache.
Adopting a hybrid DRAM-PM architecture, the three-layer expansion hash structure of the directory-segment-bucket layer is built, and metadata and key-value pairs are stored separately, and multiple cyclic queues are set up in DRAM. Through the cyclic queue, data is processed in batches to reduce read and write amplification.
It effectively eliminates read and write amplification of PM hash indexes, improves index throughput and persistence efficiency, optimizes CPU cache utilization, adapts to the size of XPLine, and reduces PM access overhead.
Smart Images

Figure CN120255799A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of computer technology. More specifically, it relates to a data-driven hash index method based on persistent CPU cache. Background Art
[0002] Hash-based index structures are widely used in in-memory databases and in-memory key-value stores. With the development of emerging storage technologies, hash indexes have evolved to better adapt to the characteristics of new hardware. Persistent Memory (PM) technology, with advantages such as large capacity, byte-addressability, persistence, and low latency close to Dynamic Random Access Memory (DRAM), has become a new generation of storage media. Existing PM technologies include PCM, ReRAM, 3D XPoint, and MRAM, etc. To make full use of the advantages of PM, many studies have optimized hash indexes on PM. Hash indexes on PM can retain the index structure during system crashes. Therefore, when optimizing hash indexes, the crash consistency problem of PM must be solved to ensure that the correct state of the index can be restored after a system crash. On the premise of ensuring crash consistency, improving the throughput of the index is the core goal of the design.
[0003] Initially, when the concept of PM was first proposed, researchers used DRAM-based PM simulators such as Quartz to explore the feasibility of deploying hash indexes on PM. It was found that DRAM-based PM simulators have the following characteristics: (1) The access latency of the simulator is between DRAM and HDD; (2) The write latency is higher than the read latency; (3) Special instructions, such as cache line flush instruction (CLFLUSH) and memory fence instruction (MFENCE), are required to persist data from the CPU cache to PM to ensure data consistency. During this period, minimizing the read and write to PM was the focus of index design. The representative work CCEH (Cache-Conscious Extendible Hashing) implemented an optimized extendible hash scheme on a PM simulator. While reasonably using CLFLUSH and MFENCE instructions to ensure data consistency, it reduced the number of cache line accesses, thereby reducing the frequency of read and write accesses and improving the throughput of index operations. It is worth noting that based on the directory-bucket (hash bucket) two-layer structure of extendible hashing, CCEH proposed the concept of a directory-segment-bucket three-layer extendible hash, which has been widely used in subsequent research.
[0004] In recent years, a commercial PM supporting the Asynchronous DRAM Refresh (ADR) technology, called the Optane Data Center Persistent Memory Module (DCPMM), has been introduced in this field. The first-generation DCPMM ensures the persistence of data in the PM and has the following distinct features: (1) The minimum access unit of the PM (referred to as the XPLine) is 256 bytes, which determines that when writing data less than 256 bytes in an XPLine, write amplification will occur; (2) The write bandwidth of the PM is one-third of the read bandwidth and is five times lower than the DRAM write bandwidth; (3) It is indeed necessary to use CLFLUSH and MFENCE in the PM to ensure the consistency between the CPU cache and PM data. However, these instructions have a relatively high write latency and will occupy the limited PM write bandwidth. To fully adapt to these features, researchers further improved based on the hash index technology developed previously on the PM simulator. First, to adapt to the size of the XPLine, the index of this period stores a certain amount of key-value pairs and their corresponding metadata in the same 256-byte block, enabling the simultaneous reading and writing of a pair of metadata and key-value pairs by accessing a single XPLine. For example, Dash is optimized based on the three-layer scalable hash structure of CCEH, modifying the hash bucket size to the XPLine size, and storing the key-value pairs and their corresponding metadata in the same hash bucket. This setting has become the general setting for hash indexes on the PM. Second, to better utilize the limited bandwidth, some studies store unimportant or recoverable data in the DRAM, thus transferring some of the read and write operations on the PM to the DRAM. Finally, the use of CLFLUSH and MFENCE instructions is still necessary. When designing the index structure, reducing the number of uses of these instructions is regarded as a major point.
[0005] In recent years, the second-generation DCPMM, which can support the enhanced Asynchronous DRAM Refresh (eADR) technology, has also been introduced in this field. Under the eADR feature, when a power failure occurs in the machine, the backup power supply of the CPU can transfer data from the CPU cache to the persistent memory (PM), enabling the CPU cache to have persistence. The eADR feature ensures the equivalence of data visibility and persistence, eliminating the possibility of data inconsistency between the CPU cache and the PM. Therefore, on machines supporting the eADR feature, it is no longer necessary to explicitly use CLFLUSH and MFENCE instructions to persist data from the CPU cache to the PM.
[0006] Since the introduction of the eADR feature, index designs based on this feature have been committed to utilizing persistent CPU cache to mitigate the impact of PM write amplification on index performance. These designs mainly focus on the following aspects to improve the performance of indexes deployed on PM: (1) Consider the mismatch between CPU cache lines and XPLine granularity in the design, that is, when cache line eviction occurs, the data written to PM is in units of 64 bytes of the CPU cache line, which is smaller than the 256 bytes of a PM line, causing write amplification; (2) Utilize lock-free technology, or integrate other technologies such as hardware transactional memory (HTM) and cache allocation technology (Intel CAT), combined with persistent cache to enhance index performance. However, these studies still use the data layout designed under the ADR feature, and the optimization of persistent CPU cache is not sufficient. This application believes that these hash indexing works have the following two defects.
[0007] The data layout is not optimized for frequently accessed metadata. As mentioned above, storing key-value pairs and their corresponding metadata in a hash bucket of the same XPLine size is a common setting for hash indexes on PM. However, the access frequency of metadata is significantly higher than that of key-value pairs, resulting in a large amount of read-write amplification when accessing metadata. With the support of the eADR feature, a better data layout needs to be designed.
[0008] The CPU cache is not optimized for the scattered access to hash buckets. Ideally, for a hash bucket, after as many insert requests as possible are completed in the CPU cache, the data of the hash bucket is persisted from the CPU cache to the PM, so that multiple insertions can be completed with one PM write access. However, since the data inserted into the index is randomly distributed, the data will hit different hash buckets sporadically, causing a large number of cache line evictions and a very low cache hit rate. This causes the data in the same hash bucket to be evicted in advance before accumulation, resulting in a large amount of write amplification.
[0009] The above defects in the prior art cause read-write amplification of the PM hash index. How to effectively eliminate the read-write amplification of the PM hash index is a technical problem that needs to be solved urgently in the field. Summary of the invention
[0010] In view of the defects of the prior art, the purpose of this application is to optimize the data layout of frequently accessed metadata and the CPU cache optimization for scattered hash bucket accesses, aiming to solve the problem of read-write amplification of PM hash index in the prior art.
[0011] To achieve the above objectives, in a first aspect, the present application provides a data-driven hash indexing method based on a persistent CPU cache, the method comprising: Adopt a DRAM-PM hybrid architecture, construct a three-layer scalable hash structure of directory-segment-bucket in the persistent memory PM, and configure a physical separation layout of metadata and key-value pairs within the segment. The physical separation layout includes storing metadata at the head of the segment and key-value pairs at the tail of the segment; Set multiple circular queues in the dynamic random access memory DRAM, collect and classify requests for index operations through the circular queues, set a threshold for each circular queue, and when the data volume in the circular queue reaches the threshold, trigger batch processing of the data in the corresponding circular queue. The batch processing is used to persist the data in batches to the PM.
[0012] It can be understood that by designing the hash index part in the PM based on the three-layer scalable hash structure of CCEH and designing a data layout that separates metadata from key-value pairs in the segment and bucket, the read-write amplification caused by frequently accessed metadata can be effectively reduced.
[0013] By setting multiple circular queues in the DRAM, collecting and classifying requests for index operations in the DRAM, and then batch processing the data in the requests after the requests accumulate to a certain scale, the write amplification caused by scattered hash bucket access can be effectively reduced.
[0014] Therefore, through the above optimization of the data layout for frequently accessed metadata and the CPU cache optimization for scattered hash bucket access, the read-write amplification of the PM hash index can be effectively eliminated.
[0015] In a possible implementation, the intra-segment area includes a segment head, a segment middle, and a segment tail, and the local depth and redundant buckets are provided in the segment middle.
[0016] It can be understood that by setting redundant buckets composed of a small number of key-value pairs and their corresponding metadata in the middle of the segment space, the splitting of the segment can be delayed. At the same time, by storing the local depth variable used to manage the scalable hash hierarchy in the structure of the redundant bucket, the segment structure can be adapted to the size of the XPLine.
[0017] In a possible implementation, the above collection and classification of requests for index operations through circular queues include: Classify the key-value pairs into a circular queue according to the most significant bits MSBs of the key and store them in the empty slots at the tail of the circular queue; Among them, multiple circular queues are allocated in a continuous memory block, and the slots for storing key-value pairs in the circular queue are located by offsets.
[0018] It can be understood that if a pointer array is used to manage multiple queues, additional pointer dereferences will be introduced. In this application, a continuous memory block is allocated for all queues in DRAM, and the slots storing key-value pairs in the queue are located through offsets, which can effectively eliminate the overhead caused by pointers.
[0019] In a possible implementation manner, the above-mentioned batch persistence of data to PM includes: When the amount of data in the target circular queue reaches the threshold, find the directory structure in PM related to the target circular queue, build a directory snapshot in DRAM, and copy the found directory structure into the directory snapshot; Insert the data in the target circular queue into the directory snapshot; Insert the data in the directory structure in PM related to the target circular queue into the directory snapshot; Persist the directory snapshot to PM.
[0020] It can be understood that by pre-building the index structure of PM, the directory snapshot merges the circular queue data and the index data together, reducing the access to PM and making the batch write become sequential access, thus improving the persistence efficiency and throughput.
[0021] In a possible implementation manner, it further includes: During cold start, merge the logic of multiple circular queues into a single queue, and set the initial threshold according to the scale of the merged queue; Gradually split the queue as the directory splits, and control the threshold to increase dynamically with the split progress until the merged queue is split completely, and then configure the threshold to a fixed value.
[0022] It can be understood that in the above-mentioned batch persistence scheme based on circular queues and directory snapshots, there is a phenomenon that the PM bandwidth is not fully utilized. Specifically, first, in the cold start stage, it is necessary to wait for enough data to accumulate in the queue before entering the request processing stage of index construction; in addition, if the threshold is simply lowered, multiple queues persisting at the same time will cause PM bandwidth contention. These two problems exacerbate the tail latency. To solve these problems, this application dynamically adjusts the merging and splitting of queues during the index construction process, thereby improving the early PM bandwidth utilization rate and reducing the tail latency.
[0023] In a possible implementation manner, the above-mentioned control of the threshold to increase dynamically with the split progress until the merged queue is split completely includes: The previous times when triggering the request processing stage, the rd threshold is , , represents the number of circular queues, Indicates the amount of data persisted each time, ; Starting from the th trigger of the request processing phase, the configured threshold is a fixed value.
[0024] It should be noted that for the threshold , if a fixed value is adopted, when a certain original queue in the merge queue triggers a directory snapshot, the data accumulated in the merge queue can be as high as times that of a single original queue threshold at most, and the problem of tail latency cannot be effectively alleviated. In order to maintain a roughly constant amount of data during persistence, this application designs an algorithm for dynamically adjusting the threshold according to the above formula to ensure that the amount of data persisted each time is approximately , effectively avoiding waste of PM bandwidth and reducing tail latency.
[0025] In a possible implementation, it further includes: Storing the offset of the secondary directory in the directory block through the primary directory, storing the offset of the segment in the segment block through the secondary directory, and restoring the data structure in PM by using the offset; Recording the key-value pairs collected in DRAM but not yet persisted in the write-ahead log WAL, and during the recovery process, re-executing the unfinished data operations in WAL to restore the data structure in DRAM.
[0026] It can be understood that remapping the physical address in PM will cause the pointer to become invalid during recovery. This application manages the hash index structure by recording the offset, and the structure in PM can be instantaneously restored by using the offset. For the data structure of DRAM, during the recovery process, by re-executing the unfinished data operations in WAL, data consistency can be ensured.
[0027] In a possible implementation, it further includes: Configuring a segment space allocator, and the segment space allocator serves as a recycle queue; When allocating a new segment, first try to allocate a recycled block from the recycle queue. If the allocation fails, calculate the offset within the segment block and allocate an unused block from the segment block; Scanning the directory in PM and recording the active segments; Recycling the unused segments into the recycle queue.
[0028] It can be understood that the segment space allocator only requires a small amount of memory, and this memory can reside in the CPU cache, thereby minimizing the PM access overhead.
[0029] In a second aspect, the present application provides an electronic device / image signal generator / network device / transmitter / terminal / base station / industrial control computer, including: at least one memory for storing programs; at least one processor for executing the programs stored in the memory, and when the programs stored in the memory are executed, the processor is used to execute the method described in the first aspect or any possible implementation manner of the first aspect.
[0030] In a third aspect, the present application provides a computer-readable storage medium storing a computer program, and when the computer program runs on a processor, it causes the processor to execute the method described in the first aspect or any possible implementation manner of the first aspect.
[0031] It can be understood that the beneficial effects of the above second aspect to the third aspect can refer to the relevant descriptions in the first aspect, and will not be elaborated here.
[0032] Generally speaking, compared with the prior art by the above technical solutions conceived by the present application, the following beneficial effects are achieved: (1) Through the data layout optimization for frequently accessed metadata and the CPU cache optimization for scattered hash bucket access, the read / write amplification of the PM hash index can be effectively eliminated.
[0033] (2) By setting redundant buckets composed of a small number of key-value pairs and their corresponding metadata in the middle of the segment space, the splitting of the segment can be delayed. At the same time, by storing the local depth variable for managing the scalable hash hierarchy in the structure of the redundant bucket, the segment structure can be adapted to the size of the XPLine.
[0034] (3) A continuous memory block is allocated for all queues in the DRAM, and the slots for storing key-value pairs in the queue are located by offsets, which can effectively eliminate the overhead caused by pointers.
[0035] (4) Through pre-building the index structure of the PM in the directory snapshot, the merging of the circular queue data and the index data is concentrated in the directory snapshot, reducing the access to the PM and making the batch write become sequential access, thus improving the efficiency and throughput of persistence.
[0036] (5) Dynamically adjust the merging and splitting of the queue during the index construction process, thereby improving the early PM bandwidth utilization rate and reducing the tail latency. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Figure 1 is the overall process framework diagram provided by the embodiment of the present application; Figure 2 is the schematic diagram of the segment structure with metadata separation mode implemented in the PM provided by the embodiment of the present application; Figure 3 It is a schematic diagram of a circular queue implemented in DRAM provided by an embodiment of the present application; Figure 4 It is a flowchart on the circular queue provided by an embodiment of the present application; Figure 5 It is a schematic diagram of a method for directory snapshot and persistent data implemented in DRAM provided by an embodiment of the present application; Figure 6 It is a flowchart on the directory snapshot provided by an embodiment of the present application; Figure 7 It is a schematic diagram of a method for optimizing PM bandwidth utilization during cold start provided by an embodiment of the present application; Figure 8 It is a flowchart after introducing an optimization method when persisting the directory snapshot provided by an embodiment of the present application; Figure 9 It is a schematic diagram of a persistent memory space allocator provided by an embodiment of the present application; Figure 10 It is a schematic diagram of the throughput comparison between the present application and other PM hash indexes under different numbers of threads; Figure 11 It is a schematic diagram of the insertion throughput of the present application and other PM hash indexes on the ADR platform and the eADR platform; Figure 12 It is a schematic diagram of the change in load rate as data is inserted for the present application and other PM hash indexes; Figure 13 It is a schematic diagram of the structure of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0038] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0039] In the embodiments of the present application, words such as "exemplary" or "for example" are used to represent examples, illustrations or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the embodiments of the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of words such as "exemplary" or "for example" is intended to present relevant concepts in a specific manner.
[0040] In the description of the embodiments of the present application, unless otherwise specified, the meaning of "a plurality of" refers to two or more. For example, a plurality of processing units refers to two or more processing units, etc.; a plurality of elements refers to two or more elements, etc.
[0041] The technical solution adopted in this application includes the following six steps.
[0042] Step 1, when designing the hash index, a DRAM-PM hybrid structure is adopted.
[0043] The DRAM-PM hybrid structure is a hybrid memory system. A hybrid memory system usually has multiple memory layers or integrates two or more heterogeneous memory technologies. The DRAM-PM hybrid structure can specifically be a hybrid memory system constructed using two memory technologies, namely DRAM (such as DDR4 DRAM) and PM (such as byte-addressable persistent memory PM).
[0044] Step 2, in order to reduce the read-write amplification caused by frequently accessed metadata, based on the directory-segment-bucket three-layer scalable hash structure of CCEH, the hash index part in PM is designed, and a data layout (hereinafter referred to as "metadata separation") in which metadata and key-value pairs are separated is designed in segments and buckets.
[0045] Step 3, in order to reduce the write amplification caused by scattered hash bucket accesses, the requests for index operations are collected and classified in DRAM, and after the requests accumulate to a certain scale, the data in the requests is processed in batches.
[0046] Step 4, in order to improve the persistence efficiency on the basis of Step 3, when processing requests in batches, first pre-construct the index structure in DRAM, and then persist the structure to PM in a batch write manner.
[0047] Step 5, in order to improve the bandwidth utilization rate at the initial stage of index construction, the conditions for triggering batch processing of requests are reduced through structural improvements.
[0048] Step 6, in order to ensure the crash consistency of PM, a suitable PM space allocator is designed.
[0049] The above Step 1 is implemented through the following solution: Manage PM through the App-Direct mode, and let the application program manage the usage timing of DRAM and PM.
[0050] The App-Direct mode is an access mode of the Optane Data Center Persistent Memory Module (DC Persistent Memory Module, DCPMM), which allows direct access to persistent memory and the data has persistence.
[0051] The above Step 2 is implemented through the following solution: (1) The key-value pairs and their corresponding metadata are no longer stored in the same hash bucket. Only the key-value pairs are stored in each hash bucket, which is placed at the tail of the segment space, while the metadata in each bucket is stored separately at the head of the segment space. (2) In the middle of the segment space, redundant buckets consisting of a small number of key-value pairs and their corresponding metadata are set up to delay the splitting of the segment. To make the structure adapt to the size of XPLine, the local depth variable used to manage the scalable hash hierarchy is stored in the structure of the redundant bucket.
[0052] The above step 3 is implemented through the following scheme: (1) Design several circular queues in DRAM to collect and classify the requests of index operations; (2) Set a threshold for each circular queue. When the amount of data in a certain queue reaches the threshold, batch processing of the data in that queue is triggered.
[0053] The above step 4 is implemented through the following scheme: (1) When the amount of data in a certain circular queue reaches the threshold, find the directory structure in PM related to that queue, build a directory snapshot in DRAM, and copy the found directory structure into the directory snapshot; (2) Insert the data in the circular queue into the directory snapshot; (3) Insert the data in the directory structure in PM related to that queue into the directory snapshot; (4) Persist the directory snapshot to PM.
[0054] The above step 5 is implemented through the following scheme: (1) In the stage of index initialization, combine multiple original queues logically into a merged queue; (2) For each original queue, set a dynamic threshold when the queue is in the merged state; (3) The merged queue is gradually split as the corresponding directory in the index splits, and the threshold gradually increases during the splitting process; (4) When the merged queue is split back to the original queue, the threshold becomes a fixed value.
[0055] The above step 6 is implemented through the following scheme: (1) Design directory blocks and segment blocks, and manage the space by recording offsets; (2) Design a segment space allocator to allocate and recycle segments during persistence.
[0056] The embodiments of the present application will be described below with reference to the accompanying drawings in the embodiments of the present application.
[0057] Refer to Figure 1 As shown, it is the overall process framework diagram of the present application.
[0058] First, the overall structure of the index is introduced. In the PM, a hash index based on the three-layer scalable hash structure of directory-segment-bucket of CCEH is designed, and the idea of metadata separation is implemented in the data layout of the segment. In the DRAM, several circular queues continuously collect and classify the insertion requests of key-value pairs. After the requests accumulate to a certain number, the directory snapshot assists in completing the insertion into the index structure in the PM.
[0059] Next, the overall process of building the index is introduced. Figure 1 The red part in [Figure] is the main process of the index. For a key-value pair requested to be inserted, the processing process can be divided into two stages. The first stage is the request collection stage. According to the most significant bit of the key ( ), the key-value pair is collected and classified into a circular queue in the DRAM and stored in the empty slot of the queue. When the amount of data in the queue reaches the set threshold, all requests in the queue enter the second stage. The second stage is the request processing stage. First, find the directory in the PM corresponding to the queue, and then create a structure snapshot for the directory in the DRAM. Next, insert the data in the queue into the structure snapshot according to the insertion method of the scalable hash, and then insert the data in the directory into the snapshot. After all data is inserted, the directory snapshot is persistently written to the hash index in the PM in a batch write manner, and finally the snapshot is destroyed. Figure 1 The other parts in [Figure] are the optimized designs of the index. First, in order to manage the segment space opened up and recycled when persisting the queue, a segment space allocator is designed in this paper. Second, in order to maximize the utilization of the PM bandwidth at the initial stage of index construction, a mechanism for merging and splitting queues is designed in this paper. At the initial stage of index construction, multiple queues are logically merged into an overall entity, and as the amount of data increases, these queues are logically split.
[0060] Refer to Figure 2 As shown, it is the segment structure implemented in the PM of this application, with a data layout in the metadata separation mode.
[0061] In the ADR era, the CPU cache has not been incorporated into the persistent domain. Therefore, in order to reduce PM writes while ensuring data consistency, existing PM hash structures usually store metadata and key-value pairs in the same hash bucket. However, this design is not applicable in the eADR era because it fails to fully utilize the characteristics of the persistent cache. This application observes that metadata occupies little space and has a very high access frequency, so it is very suitable for storage in the cache. Based on this observation, this application proposes a metadata separation design, which stores metadata and key-value pairs separately, and redesigns the data layout in the segment structure to make it more suitable for building an index on the eADR platform.
[0062] Figure 2 Shows the data layout of the middle section structure of this application. In each section, the metadata of the section, key-value pairs, and local depth are stored. The metadata is stored at the beginning of the section, the local depth is in the middle, and the key-value pairs are stored at the end.
[0063] Specifically, most key-value pairs are stored in units of hash buckets. Each section contains 16 hash buckets, each bucket has 16 slots, and each slot occupies 16 bytes of space to store a key-value pair. When dividing the buckets in a section, this application uses the least significant bit of the key ( ) to divide, so 4 bits are used to distinguish 16 hash buckets. . For example, assume that a key-value pair with a key of 1110110 is stored in a section. Then this key-value pair is stored in the 6th hash bucket of this section. The storage order of metadata items corresponds one by one to the order of key-value pairs in the hash bucket. For example, for the key-value pair in the first slot of the first bucket, its metadata item is stored in the first position of the metadata area. Each metadata item includes: (1) a 14-bit fingerprint (for accelerating search); (2) a 1-bit valid bit (1 indicates the item is valid, 0 indicates invalid); (3) a 1-bit used to mark whether the value or pointer is stored at the corresponding position (1 indicates the value is directly stored in the bucket, 0 indicates a pointer to the value is stored in the bucket). Among them, the fingerprint is used to accelerate search, and the valid bit is used to identify whether the data is valid at the position of the key-value pair corresponding to this metadata.
[0064] When the data volume in any bucket reaches the bucket capacity, the next insertion operation for that bucket will cause a section split, which involves rehashing the data in the entire section and will result in a lower space utilization rate. To reduce the frequency of section splits, the newly inserted key-value pairs to be inserted are first stored in the redundant bucket until the capacity of this area is insufficient and then a section split is performed. Therefore, a small number of key-value pairs are stored in the redundant bucket. This application sets that the redundant bucket in the section can store at most 28 key-value pairs and their corresponding metadata, and the space reserved for each key-value pair and metadata is also 16 bytes and 2 bytes. In addition, the local depth is used to manage the levels of extensible hashing and occupies 2 bytes. The position of this variable is arranged in the redundant bucket to adapt to the size of XPLine.
[0065] In the entire section structure, 16 hash buckets altogether occupy 16 XPLines, the metadata corresponding to these hash buckets altogether occupies 2 XPLines, and the redundant bucket occupies 2 XPLines. Therefore, such a data layout makes better use of the access characteristics of PM on the premise of separating metadata.
[0066] Refer to Figure 3 As shown, it is the circular queue implemented in DRAM of this application, which completes the collection and classification of insertion requests; refer to Figure 4As shown, it is the flowchart of this application on the circular queue.
[0067] In previous PM-based hash index work, the mismatch between the key-value pair size and the XPLine size led to severe write amplification. This application believes that by leveraging the persistent CPU cache, the problem can be alleviated by collecting key-value pairs within a single XPLine and writing them back to PM in one persistence operation. However, experiments show that the hit rate of newly inserted key-value pairs in the CPU cache is very low, and only a very small number of key-value pairs can be persisted together with other key-value pairs, so the write amplification problem remains unsolved. Therefore, this application uses DRAM to collect and classify data, and then performs batch writes, thereby reducing the write amplification problem and the write frequency of PM.
[0068] This application creates multiple circular queues in DRAM to serve the request collection phase of the index construction process. The start and end positions of the i-th queue are respectively denoted as and . In the request collection phase, the data newly inserted into the queue is stored at the end position of the queue, and it causes to move back 1 position. When the process enters the request processing phase, each data is persisted to PM, which causes to move back one position. The capacity of each queue is a fixed value, denoted as .
[0069] In the request collection phase, according to the of the key, the key-value pairs are classified into one of the queues and stored in the empty slots at the tail of the queue. Using an array of pointers to manage multiple queues will introduce additional pointer dereferencing. Therefore, to eliminate the overhead caused by pointers, this application allocates a continuous memory block in DRAM for all queues, and the slots storing key-value pairs in the queue are located through offsets. Figure 3 The process of locating the slots is illustrated by an example. In this example, assume that each queue can hold at most data, and the blue squares represent the occupied slots. For the key-value pair with the key 110110, assume has a length of 2, that is, the queue uses the first 2 bits for classification. Then, according to , the key-value pair is classified into the fourth queue . Since , the queue is not full and insertion can be performed. Since , the position of the key-value pair in the DRAM block is , that is, the position of the green square in the figure. Summarizing this process, the flowchart of operations in the circular queue is as shown in Figure 4 .
[0070] The total capacity of the circular queue structure is much larger than the order of magnitude of the CPU cache capacity and has a fixed size. When collecting data, it can accommodate a large amount of data without exceeding the established capacity.
[0071] As shown in Figure 5 is the directory snapshot implemented in DRAM in this application and the method of persisting data; as shown in Figure 6 is the flowchart of this application on the directory snapshot.
[0072] In the circular queue, each queue sets a threshold, called the persistence threshold T. When the amount of data in the queue reaches T, it enters the request processing stage of the index construction process. In this stage, the data in the PM index should be merged with the data in the DRAM circular queue. However, in-place merging may cause space contention for multi-threaded write operations and reduce concurrency performance. Therefore, this application arranges the merging process in the directory snapshot and adopts the copy-on-write technology in the directory snapshot. In addition, this application introduces a two-level directory design to more effectively implement recovery.
[0073] As previously described, the hash index of PM is a three-layer structure of directory-segment-bucket. On this basis, in order to associate the circular queue in DRAM with the directory entries in the PM index to facilitate optimizing the operations in the request processing stage, this application divides the directory into two levels. The first-level directory is a static directory, and each directory entry corresponds to a queue one by one. The second-level directory is several dynamic directories, where each second-level directory corresponds to a first-level directory entry, and each second-level directory continues to be split into several second-level directory entries.
[0074] When the queue reaches the persistence threshold T, the method of this application passes through the of the queue to find the directory entry corresponding to this queue in the first-level directory of PM. Then, according to the pointer stored in this directory entry, further determine the storage location of the corresponding second-level directory and copy this second-level directory to the directory snapshot. To reduce the number of accesses to XPLine, only the structure of the second-level directory and its referenced segments are copied, and the data within the segments is not copied. After constructing the structure snapshot, the method of this application inserts the items in the queue into the snapshot in order from to and records which segments have data insertions during this process. After all the data in the queue is inserted, find all the segments in the PM index corresponding to the segments modified in the directory snapshot and merge the data in these segments into the segments of the directory snapshot. During the insertion of the queue data and the PM index data into the directory snapshot, segment splitting and directory doubling may be triggered, so the dynamic directory may change. Finally, persist the segments and dynamic directories modified in the directory snapshot to PM, and then update to indicate the end of persistence. The number of shifted positions is the number of key-value pairs that are persisted from the queue to the PM this time.
[0075] Figure 5 The example in illustrates the operations in the request processing stage. Assume , and continue using the example of the key-value pair with the key 110110. This key-value pair has , so the queue and PM index drawn in the figure both start with 11, and the first 2 bits of each data are omitted in the figure. The operation steps in the example are as follows: (1) Read the dynamic directory in the persistent memory corresponding to and create a directory snapshot in DRAM; (2) Insert the items in Figure 5 into the directory snapshot; (3) For the segments modified in the directory snapshot in the previous step, read the corresponding key-value pairs from the PM index and merge them into the directory snapshot; (4) Persist the directory snapshot to the PM and destroy the snapshot. In Figure 6 , the green area represents the part that reads from the PM, and the blue area represents the part that writes to the PM. Summarizing this process, the flowchart of the operations in the directory snapshot is as shown in
[0076] By pre-building the index structure of the PM, the directory snapshot centralizes the merging of the circular queue data and the index data, reducing the access to the PM and making the batch writes sequential access, thus improving the throughput.
[0077] Referring to Figure 7 shown, it is the method for optimizing the PM bandwidth utilization rate during the cold start period of this application; referring to Figure 8 shown, it is the final flowchart after introducing the optimization method when persisting the directory snapshot in this application.
[0078] In the batch persistence scheme based on the circular queue and directory snapshot mentioned above, there is a phenomenon that the PM bandwidth is not fully utilized. First, in the cold start stage, it is necessary to wait for enough data to accumulate in the queue before entering the request processing stage of index construction; in addition, if the threshold is simply lowered, multiple queues persisting at the same time will cause PM bandwidth contention. These two problems exacerbate the tail latency. To solve these problems, this application dynamically adjusts the merging and splitting of the queues during the index construction process, thereby improving the early PM bandwidth utilization rate and reducing the tail latency.
[0079] In the initialization stage of the system, M queues are logically merged. Each independent queue is called the original queue, and the logically merged queue is called the merged queue. The corresponding first-level directories of the queues are also logically merged accordingly. At this time, assume that the number of bits used to divide the queues is ( indicating the number of bits obtained ), the number of bits used to identify the merged queue of M queues is . If directory doubling occurs during the insertion of the directory snapshot, the merged queue and its corresponding first-level directory are logically split. Specifically, when the data volume of a certain queue reaches the threshold T and enters the request processing stage, first check whether the directory entry of the first-level directory corresponding to this queue is empty. If it is empty, it means that the current queue is an original queue in a certain merged queue. In this case, in order to find the head of the merged queue, each time is subtracted by 1, and check whether the pointer in the first-level directory entry found with this number of bits is valid until the first valid pointer is found, and then find the pointed second-level directory. Then, continue to follow the Figure 5 directory snapshot workflow in, first copy the second-level directory structure in PM in the directory snapshot, then insert the data in the circular queue and PM index into the snapshot in turn, and finally persist the modification of the directory snapshot to PM and destroy the snapshot. It should be noted that for the snapshot of a merged queue, if directory doubling occurs, it means that the merged queue and its corresponding first-level directory are split. Therefore, some first-level directory entries should be modified after persistence, that is, the first-level directory entry should point to the correct second-level directory.
[0080] For the threshold T, if a fixed value is adopted, when a certain original queue in the merged queue triggers a directory snapshot, the accumulated data in the merged queue can be as high as nearly M times the threshold of a single original queue at most, and the problem of tail latency cannot be effectively alleviated. In order to keep the data volume roughly constant during the execution of persistence, this application designs an algorithm for dynamically adjusting the threshold. When initializing the queue, set the of each original queue to the second slot of the queue, and the first slot is empty. For each original queue, in the first k times of triggering the request processing stage, the threshold for the i-th time ( ) is . After the i-th persistence is completed, the total amount of data persisted in the first i times is . This application sets k to , then after the k-th persistence is completed, the total amount of data that has been persisted is , and the last key-value pair to be persisted is the data at the last position in the queue. Under the current total amount of persisted data, it can exactly ensure that the first-level directory corresponding to the original queue is completely split, that is, the original queue is no longer in a merged state. Next, when the request processing stage is triggered for the (k + 1)-th time, the first slot is no longer empty. Therefore, whether the original queue is a merged queue can be determined by whether the first slot is empty. Starting from the (k + 1)-th trigger of the request processing stage, the persistence threshold T is set to a fixed value slightly less than s.
[0081] Figure 7 shows an example of queue merging and splitting. Suppose , , and the threshold T of all circular queues at the current moment is 4. When inserting the key-value pair with the key 110110, the operation steps on the directory snapshot are as follows: (1) When inserting the key-value pair, it is found that the queue reaches the threshold, and this queue triggers the request processing stage; (2) Locate the directory corresponding to the merged queue, find the first-level directory entry , and then find the secondary directory pointed to by ; (3) Copy this secondary directory to the directory snapshot; (4) Cause segment splitting; (5) Persist the directory snapshot to the PM, change the pointer of the first-level directory, and destroy the directory snapshot. Summarizing this process, after adding the optimization of queue merging and splitting, one more thing is added to the persistence of the directory snapshot, that is, changing the pointer of the first-level directory. Specifically, it is necessary to determine whether a new secondary directory is generated. If so, the corresponding pointer in the first-level directory needs to be changed. After introducing the optimization method when persisting the directory snapshot, the final flowchart is as Figure 8 shown.
[0082] By following this strategy, this application ensures that the amount of data persisted each time is approximately s, effectively avoiding the waste of PM bandwidth.
[0083] Refer to Figure 9 shown, which is the persistent memory space allocator designed by this application.
[0084] The method of this application is equipped with a dispenser in the PM and restores the data structures in the PM and DRAM in the following ways. (1) Restore the data structure of the PM. Remapping the physical address in the PM causes the pointers to become invalid during restoration. Therefore, the method of this application manages the hash index structure by recording the offsets. Specifically, the first-level directory stores the offset of the second-level directory in the directory block, and the second-level directory stores the offset of the segment in the segment block. Thus, the structure in the PM can be restored immediately by using the offsets. (2) Restore the data structure of the DRAM. The key-value pairs that are collected in the DRAM but not yet persisted will be recorded in the write-ahead log (WAL). During the restoration process, the unfinished data operations in the WAL will be re-executed to ensure data consistency.
[0085] The method of this application also introduces a segment space dispenser for allocating and recycling segments during persistence, as Figure 9 shown. The segment block contains used blocks, unused blocks, and recycled blocks. The segment space dispenser is a recycling queue that records the blocks allocated and recycled in the segment block (recycled blocks) in the queue. When a new segment needs to be allocated, it first tries to allocate a recycled block from the recycling queue. If the allocation fails, it calculates the offset within the segment block and allocates an unused block from the segment block. When the system crashes, a new segment that is about to be persisted may be lost, resulting in memory leakage. To ensure memory safety, the directories in the PM are scanned and the active segments are recorded. Any unused segments identified during this process will be recycled into the dispenser queue.
[0086] The PM dispenser only requires a small amount of memory, which can reside in the CPU cache, thus minimizing the PM access overhead.
[0087] In summary, due to the adoption of the above technical solutions, the beneficial effects of this application are as follows.
[0088] (1) The high-performance scalable hash index method designed in this application achieves a throughput higher than that of hash indexes on other PMs. YCSB is a tool that can generate workloads of index operation requests, Figure 10 showing the throughput comparison between this application (Ours) and other PM hash indexes (Dash, Pea, Spash, Level, CCEH, Clevel) under different numbers of threads for 6 workloads of YCSB. Among the 6 workloads used in the experiment, (a) to (c) are uniformly distributed workloads, and (d) to (f) are skewed distributed workloads. (a) and (d) are write workloads, (b) and (e) are read-write mixed workloads, and (c) and (f) are read workloads. As Figure 10 can be seen, this application overall outperforms other comparison indexes under various workloads and performs better under write-intensive workloads and a large number of threads.
[0089] (2)The high-performance scalable hash index method designed in this application can make better use of the eADR feature compared with hash indexes on other PMs. Figure 11 The insertion throughput of the method of this application and other comparison indexes under the ADR platform and the eADR platform is compared. Figure 11 It can be seen that the method of this application can not only obtain the maximum throughput, but also achieve the highest performance improvement under the eADR feature.
[0090] (3)The hash index designed in this application performs comparably with hash indexes on other PMs in terms of key performance indicators (such as load factor and recovery time). Table 1 shows the recovery time of this application. It can be seen from Table 1 that the recovery time of this application can complete index recovery in milliseconds, slightly faster than the existing hash indexes. Figure 12 shows the change in the load factor of this application as data is inserted. Figure 12 It can be seen that the load factor of this application shows a fluctuating state, and the overall data as well as the highest and lowest load factors are comparable to those of other hash indexes.
[0091] Table 1 Comparison Table of Recovery Time
[0092] Based on the method in the above embodiments, an embodiment of this application provides an electronic device. As Figure 13 shown, the electronic device may include: a processor (Processor) 810, a communication interface (Communications Interface) 820, a memory (Memory) 830, and a communication bus 840. Among them, the processor 810, the communication interface 820, and the memory 830 communicate with each other through the communication bus 840. The processor 810 can call the logical instructions in the memory 830 to execute the method in the above embodiments.
[0093] In addition, when the logical instructions in the above-mentioned memory 830 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions to enable a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of this application.
[0094] Based on the method in the above embodiments, an embodiment of the present application provides a computer-readable storage medium. The computer-readable storage medium stores a computer program. When the computer program runs on a processor, the processor is caused to execute the method in the above embodiments.
[0095] Based on the method in the above embodiments, an embodiment of the present application provides a computer program product. When the computer program product runs on a processor, the processor is caused to execute the method in the above embodiments.
[0096] It can be understood that the processor in the embodiments of the present application may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. The general-purpose processor may be a microprocessor or any conventional processor.
[0097] The method steps in the embodiments of the present application may be implemented in a hardware manner or by a processor executing software instructions. The software instructions may be composed of corresponding software modules. The software modules may be stored in a random access memory (RAM), flash memory, read-only memory (ROM), programmable ROM (PROM), erasable PROM (EPROM), electrically erasable PROM (EEPROM), registers, hard disk, removable hard disk, CD-ROM, or any other form of storage medium well-known in the art. An exemplary storage medium is coupled to the processor so that the processor can read information from the storage medium and write information to the storage medium. Of course, the storage medium may also be a component of the processor. The processor and the storage medium may be located in an ASIC.
[0098] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted through the computer-readable storage medium. The computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired manner (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or a wireless manner (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or a data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, a hard disk, a magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)), etc.
[0099] It can be understood that the various numerical numbers involved in the embodiments of the present application are only for the convenience of description and are not used to limit the scope of the embodiments of the present application.
[0100] Those skilled in the art can easily understand that the above are only the preferred embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present application shall be included in the protection scope of the present application.
Claims
1. A data-driven hash index method based on persistent CPU cache, characterized in that including: Adopt a DRAM-PM hybrid architecture, build a three-layer scalable hash structure of directory-segment-bucket in the persistent memory PM, and configure a physical separation layout of metadata and key-value pairs within the segment. The physical separation layout includes storing metadata at the head of the segment and key-value pairs at the tail of the segment; Set multiple circular queues in the dynamic random access memory DRAM, collect and classify requests for index operations through the circular queues, set a threshold for each circular queue, and when the data volume in the circular queue reaches the threshold, trigger batch processing of the data in the corresponding circular queue. The batch processing is used to batch-persist the data to PM.
2. The data-driven hash index method based on persistent CPU cache according to claim 1, wherein The area within the segment includes the segment head, the segment middle, and the segment tail. The local depth and redundant buckets are provided in the segment middle.
3. The data-driven hash index method based on persistent CPU cache according to claim 1, wherein The collecting and classifying requests for index operations through the circular queue includes: Classify the key-value pairs into a circular queue according to the most significant bits MSBs of the key, and store them in the empty slots at the tail of the circular queue; Among them, multiple circular queues are allocated in a continuous memory block, and the slots storing key-value pairs in the circular queue are located by offsets.
4. The data-driven hash index method based on persistent CPU cache according to claim 1, wherein The batch-persisting the data to PM includes: When the data volume in the target circular queue reaches the threshold, find the directory structure in PM related to the target circular queue, build a directory snapshot in DRAM, and copy the found directory structure to the directory snapshot; Insert the data in the target circular queue into the directory snapshot; Insert the data in the directory structure in PM related to the target circular queue into the directory snapshot; Persist the directory snapshot to PM.
5. The data-driven hash index method based on persistent CPU cache according to claim 1, wherein Also include: During cold start, logically merge multiple circular queues into a single queue, and set the initial threshold according to the scale of the merged queue; Gradually split the queue as the directory splits, control the threshold to increase dynamically with the split progress until the merged queue is split completely, and then configure the threshold as a fixed value.
6. The data-driven hash index method based on persistent CPU cache according to claim 5, wherein The controlling the threshold to increase dynamically with the split progress until the merged queue is split completely includes: Previous When processing the previous trigger request, the threshold for the th time is ; represents the number of circular queues, represents the amount of data persisted each time, ; Starting from the th trigger request processing phase, the configuration threshold is set to a fixed value.
7. The data-driven hash index method based on persistent CPU cache according to any one of claims 1-6, characterized in that, Also include: Store the offset of the secondary directory in the directory block through the primary directory, store the offset of the segment in the segment block through the secondary directory, and restore the data structure in PM by using the offset; Record the key-value pairs collected in DRAM but not yet completed for persistence in the write-ahead log WAL. During the recovery process, re-execute the uncompleted data operations in WAL to restore the data structure in DRAM.
8. The data-driven hash index method based on persistent CPU cache according to any one of claims 1-6, characterized in that Also include: Configure a segment space allocator, and the segment space allocator serves as a recycling queue; When allocating a new segment, first try to allocate a recycled block from the recycling queue. If the allocation is unsuccessful, calculate the offset within the segment block and allocate an unused block from the segment block; Scan the directory in PM and record the active segments; Recycle the unused segments to the recycling queue.
9. An electronic device, characterized in that, including: At least one memory for storing computer programs; At least one processor for executing the program stored in the memory. When the program stored in the memory is executed, the processor is used to execute the method according to any one of claims 1-8.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program runs on the processor, the processor is caused to execute the method according to any one of claims 1-8.
Citation Information
Patent Citations
Dynamic data storage method and device
CN104516912A
Separate storage method oriented to temporary metadata
CN107659626A
High performance scalable hash index based on hybrid storage
CN117112557A
Separated key-value pair storage system
CN118550474A
NUMA (Non Uniform Memory Access) perceived key value storage system based on hybrid memory and operation method
CN118779255A