Direct connection persistent memory accelerator-oriented data access method, controller and system

By designing a multi-level hash table and thread bundle collaborative execution mechanism on a direct-connected persistent memory accelerator system, the high concurrency and crash consistency problems are solved, and efficient data management and system performance improvement are achieved.

CN120045318APending Publication Date: 2025-05-27HUAZHONG UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510090742.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

When the prior art realizes high concurrent persistent hash tables on a direct-connected persistent memory accelerator system, it faces the performance degradation caused by unaware execution of thread bundles, unmerged memory access problems, and high overhead caused by ensuring data crash consistency.

Method used

A hash table for directly connected persistent memory accelerator is designed to handle hash collisions and improve memory efficiency through the sharing mechanism of multi-level hash tables and inter-level hash buckets. The thread bundle collaborative execution mechanism is adopted to realize lock-free hash table operation through atomic operations and slot states, and reduce access to persistent memory through cache based on the freezing mechanism.

Benefits of technology

It realizes efficient data management, makes full use of the high parallelism of the accelerator, reduces the overhead of write overhead and crash consistency guarantee, and improves the overall performance of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120045318A_ABST
    Figure CN120045318A_ABST
Patent Text Reader

Abstract

The invention discloses a direct connection persistent memory accelerator-oriented data access method, controller and system, and belongs to the field of key value pair storage, and the method comprises the following steps: establishing a hash table containing N levels in a persistent memory, sharing one hash bucket of each level by two hash buckets of the previous level, the storage position of each key value pair in the top layer is calculated by K hash functions; when the persistent memory is accessed, a hash operation is allocated to each thread in the thread bundle, and the threads in the thread bundle are activated in sequence; when the activated thread executes the allocated Hash operation, the rest threads in the thread bundle calculate the mapped Hash bucket according to the key corresponding to the allocated Hash operation, and transmit the mapped Hash bucket to the activated thread through the communication primitive between the threads, so that the activated thread determines a target slot corresponding to the Hash operation to be executed at present, and the target slot corresponds to the Hash operation to be executed at present. And executing a hash operation. The high parallelism of the accelerator can be fully utilized, the write overhead is reduced, and the overall performance of the system is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of key-value pair storage, and more specifically, relates to a data access method, a controller, and a system for a direct-attached persistent memory accelerator. Background Art

[0002] In the era of big data, the computing throughput and memory bandwidth of accelerators (e.g., GPUs) have increased significantly, and accelerators have generally empowered application programs in various fields. These application programs need to process a large amount of data. To ensure the reliability of data, applications usually store data on large-capacity, persistent storage devices. Existing accelerators need to rely on the processor (i.e., CPU) on the host to access the data on these persistent storage devices. This process results in high data transmission overhead and requires the participation of the processor, which may interfere with other programs.

[0003] An accelerator that directly connects to persistent memory through unified virtual address technology allows applications on the accelerator to directly access persistent memory at the byte granularity. Hash tables have been widely used in many data management systems, so application programs can utilize hash tables for accelerators that directly connect to persistent memory to achieve efficient data management. Currently, there is no hash table specifically designed for accelerators that directly connect to persistent memory. Transplanting existing persistent memory-based and accelerator-based hash tables to an accelerator system that directly connects to persistent memory is an intuitive implementation method. However, achieving a high-concurrency persistent hash table on an accelerator system that directly connects to persistent memory faces the following challenges:

[0004] First, the warp-agnostic execution mode will lead to a serious performance degradation. Generally, a hash table uses multiple threads to concurrently execute operations to achieve high throughput, where each thread independently executes hash table operations. However, this warp-agnostic execution mode has serious warp divergence and uncoalesced memory access problems. Specifically, when threads in the same warp execute different operations without intra-warp thread communication, they will passively encounter warp divergence caused by branch instructions. In addition, most hash tables only store pointers to keys (or key-value pairs) to reduce storage overhead. However, when threads in a warp access the queried keys in parallel, the addresses of these keys will be scattered, resulting in uncoalesced memory access. In addition, some hash tables use a lock-based design, which exacerbates the contention among thousands of concurrent threads, and even threads in the same warp may cause deadlocks when trying to obtain the same lock.

[0005] Secondly, ensuring crash consistency of data incurs high overhead. Since the hash table of the accelerator directly connected to persistent memory manages data in persistent memory, ensuring crash consistency of data is very important but not easy. The atomic memory write size of persistent memory is limited by the memory bus width (for example, the memory bus width of a 64-bit processor is 8 bytes). Therefore, if a system failure occurs before writing data larger than 8 bytes is completed, the data will be corrupted. Logging and copy-on-write techniques are widely used to ensure crash consistency of data larger than 8 bytes. When using logging, the index stores the old data or new data in the log and then writes the new data in place. When using the copy-on-write technique, the old data is copied to a newly created space and updates are performed on the copy, and then an 8-byte pointer is atomically modified to point to the new data. However, these techniques introduce high write overhead.

[0006] In addition, the huge bandwidth gap between persistent memory and the accelerator limits the utilization of the high parallelism of the accelerator. There is an obvious bandwidth gap between persistent memory and accelerator memory. Therefore, when a large number of concurrent hash table operations are performed, the persistent memory with limited bandwidth cannot effectively handle a large number of concurrent accesses from the bandwidth-intensive accelerator cores. Such a huge bandwidth gap between persistent memory and the accelerator hinders the full utilization of the high parallelism of the accelerator. Moreover, some schemes, such as cuckoo hashing and linked list-based hashing, generate additional persistent memory accesses to handle hash conflicts, which further exacerbates the bandwidth problem.

[0007] Overall, existing key-value store solutions lack a hash table and corresponding operation mechanism suitable for the accelerator directly connected to persistent memory, and the overall system performance needs to be further improved. Summary of the Invention

[0008] In view of the deficiencies and improvement requirements of the prior art, the present invention provides a data access method, a controller, and a system for an accelerator directly connected to persistent memory, aiming to propose a suitable hash table construction method and corresponding operation mechanism for the accelerator directly connected to persistent memory, so as to fully utilize the high parallelism of the accelerator and reduce write overhead, thereby providing an efficient data management method for applications and improving the overall performance of the system.

[0009] To achieve the above object, according to one aspect of the present invention, a data access method for an accelerator directly connected to persistent memory is provided, where the persistent memory is used to store key-value pair data, and the data access method includes:

[0010] Create a hash table in persistent memory; the hash table contains N levels, each level contains multiple hash buckets, each hash bucket contains M slots, and each slot is used to store a key-value pair; in two adjacent levels, the number of hash buckets in the upper level is twice that of the lower level, and each hash bucket in the lower level is shared by two hash buckets in the upper level; the storage location of each key-value pair in the top level is calculated by K hash functions;

[0011] When accessing data in persistent memory, allocate a hash operation for each thread in the warp, and activate the threads in the warp in sequence, so that the activated threads execute their assigned hash operations until all the hash operations assigned to the warp are completed; when the activated thread executes its assigned hash operation, the remaining threads in the warp calculate the mapped hash bucket according to the key corresponding to the assigned hash operation, and transfer the information stored in each slot in the hash bucket to the activated thread through the inter-thread communication primitive, so that the activated thread determines the target slot corresponding to the currently to-be-executed hash operation based on the received information, and executes the hash operation;

[0012] Among them, M, N, and K are all preset positive integers, and satisfy M×N×K = P, where P is the number of threads in the warp.

[0013] Furthermore, the information stored in each slot also includes: an 8-byte slot status; the slot status includes: idle, inserting, and inserted; and:

[0014] If the key length of the key-value pair is fixed and does not exceed 8 bytes, the key-value pair information stored in each slot includes: a pointer to the value in the key-value pair, and after the key-value pair information is successfully inserted, the slot status is set to the key of the key-value pair to indicate that the status of the slot is inserted;

[0015] If the key length of the key-value pair is fixed and exceeds 8 bytes, the key-value pair information stored in each slot includes: the key in the key-value pair and a pointer to the value in the key-value pair, and after the key-value pair information is successfully inserted, the slot status is set to the hash value of the key in the key-value pair to indicate that the status of the slot is inserted;

[0016] If the key length of the key-value pair is not fixed, the key-value pair information stored in each slot includes: a pointer to the key-value pair, and after the key-value pair information is successfully inserted, the slot status is set to the hash value of the key in the key-value pair to indicate that the status of the slot is inserted.

[0017] Furthermore, if the currently to-be-executed hash operation is an insert operation, the activated thread determines the target slot corresponding to the currently to-be-executed hash operation based on the received information, and executes the hash operation, including:

[0018] S1: The key of the key-value pair to be inserted by the activated thread is the target key k i , compare the target key k i with the keys stored in each slot. If the target key k i does not exist, obtain an idle slot as the target slot, and use the CAS primitive to change the status of the target slot from idle to being inserted;

[0019] S2: If the status change fails, go to step S3; otherwise, go to step S4;

[0020] S3: Notify the remaining threads in the warp to calculate the mapped hash bucket according to the key corresponding to the assigned hash operation, and pass the information stored in each slot in the hash bucket to the activated thread through the inter-thread communication primitive, and then go to step S1;

[0021] S4: Insert the key-value pair information of the key-value pair to be inserted into the target slot, and then set the status of the target slot to inserted, and the insertion operation ends.

[0022] Further, if in step S1, the target key k i does not exist, if the acquisition of the idle target slot fails, perform the expansion operation and then obtain an idle slot as the target slot;

[0023] The expansion operation includes:

[0024] Allocate a new level above the top level of the hash table as the top level;

[0025] Scan the bottom layer of the hash table to obtain the key-value pair information stored in the bottom layer, and re-insert the key-value pairs stored in the bottom layer into the hash table through the insertion operation;

[0026] After all the key-value pair information stored in the bottom layer is re-inserted into the hash table, delete the bottom layer.

[0027] Further, if the currently to-be-executed hash operation is a deletion operation, the activated thread determines the target slot corresponding to the currently to-be-executed hash operation based on the received information, and performs the hash operation, including:

[0028] The activated thread uses the key of the key-value pair to be deleted as the target key k d , compare the target key k d with the keys stored in each slot to locate all the slots storing the target key k d as the target slots;

[0029] The activated thread sets the slot status of each target slot to idle through the CAS primitive, and the deletion operation ends.

[0030] Further, if the currently to-be-executed hash operation is an update operation, the activated thread determines the target slot corresponding to the currently to-be-executed hash operation based on the received information, and executes the hash operation, including:

[0031] The activated thread uses the key of the key-value pair to be deleted as the target key k u , and compares the target key k u with the keys stored in each slot to locate all the slots storing the target key k u , and takes one of them as the target slot and the rest as duplicate slots;

[0032] The activated thread sets the slot status of each duplicate slot to idle through the CAS primitive, and updates the key-value pair information in the target slot to the new key-value pair information through the CAS primitive.

[0033] Further, if the currently to-be-executed hash operation is a query operation, the activated thread determines the target slot corresponding to the currently to-be-executed hash operation based on the received information, and executes the hash operation, including:

[0034] T1: The activated thread uses the key of the key-value pair to be deleted as the target key k s , and compares the target key k s with the keys stored in each slot to locate all the slots storing the target key k s ;

[0035] T2: If the number of slots storing the target key k s is 0, return query failure; otherwise, go to T3;

[0036] T3: The activated thread takes one of all the slots storing the target key k s as the target slot and the rest as duplicate slots;

[0037] T4: The activated thread obtains the value pointer according to the key-value pair information stored in the target slot, reads the corresponding value and returns it.

[0038] Further, the data access method for a direct-connected persistent memory accelerator provided by the present invention further includes:

[0039] Create a cache in the memory of the accelerator;

[0040] Regularly identify the hash buckets storing hot data in the hash table as hot hash buckets, and load the identified hot hash buckets into the cache;

[0041] Moreover, after each thread in the thread bundle calculates the mapped hash bucket according to the key corresponding to the assigned hash operation, it further includes:

[0042] Access the cache to find the slot storing the target key;

[0043] And, if the cache hit occurs and the currently to-be-executed hash operation is a query operation, it further includes:

[0044] Obtain the key-value pair information from the slot storing the target key in the cache, access the value according to the key-value pair information and return it, and end the query operation;

[0045] And, if the cache hit occurs and the currently to-be-executed hash operation is not a query operation, the remaining threads in the thread bundle will transfer the information stored in each slot in the hash bucket to the activated thread through inter-thread communication primitives, so that the activated thread can determine the target slot corresponding to the currently to-be-executed hash operation based on the received information, and after executing the hash operation, it further includes:

[0046] Make corresponding modifications to the hash bucket in the cache according to the execution result of the hash operation.

[0047] According to another aspect of the present invention, there is provided a storage controller for a direct-attached persistent memory accelerator, including:

[0048] A computer-readable storage medium for storing a computer program;

[0049] And a processor for reading the computer program stored in the computer-readable storage medium and executing the above data access method for the direct-attached persistent memory accelerator provided by the present invention.

[0050] According to another aspect of the present invention, there is provided a key-value storage system, including: persistent memory, a direct-attached persistent memory accelerator, and the above storage controller for the direct-attached persistent memory accelerator provided by the present invention.

[0051] Generally speaking, through the above technical solutions conceived by the present invention, the following beneficial effects can be achieved:

[0052] (1) The hash table established by the present invention includes multiple levels. Through the sharing mechanism of hash buckets between levels, hash collisions can be effectively handled, achieving higher memory efficiency. At the same time, only the hash buckets at the top level can be located by the hash function, and the product of the number of levels, the slot correlation (the number of slots in the hash bucket), and the number of hash functions used to locate the top-level hash bucket is equal to the number of threads in a warp. Based on this design, after each thread is assigned a hash operation, the corresponding associated hash bucket can be calculated based on the relevant key. When accessing data in persistent memory through hash operations, the hash operations are allocated at the thread granularity. Each warp is activated in turn to execute the corresponding operations. When executing each hash operation, the target slot is located based on the hash buckets associated with all threads within the warp. All threads within the thread count transfer data and synchronize through communication primitives. That is to say, the present invention allocates hash operations at the thread granularity, and each hash operation is executed at the warp granularity, thereby implementing a warp cooperative execution mechanism, which can minimize warp divergence, fully utilize the high parallelism of the accelerator, and reduce write overhead, thus providing an efficient data management method for applications and improving the overall performance of the system.

[0053] (2) In the hash table established by the present invention, each slot also stores an 8-byte slot status, and the format of the information stored in the slot is set according to the length of the key in the key-value pair. When the length of the key is fixed and does not exceed 8 bytes, the key is directly stored in the slot as the slot status to indicate the inserted state, and a value pointer is stored in the slot, thereby improving memory utilization and quickly completing slot positioning. When the length of the key is fixed and exceeds 8 bytes, the hash value of the key is stored in the slot as the slot status, and the key and value are stored in the slot at the same time, thereby quickly completing slot positioning.

[0054] (3) When the present invention executes the insertion operation, it judges whether the slot is empty, being inserted, or already has data through the slot status, and uses atomic operations to update the slot status and data pointer, supporting lock-free operations. When a crash occurs, the hash table is restored through the slot status, so there is no need to write a log, which reduces the overhead of ensuring crash consistency of data while ensuring reliability, further improving the performance of the system.

[0055] (4) The present invention further implements a cache based on a freezing mechanism. Specifically, a cache is created in the memory of the direct-attached persistent memory accelerator, and the hash buckets storing hot data in the hash table are cached. The cached hash buckets are updated regularly, and between two adjacent loading phases, the cached hash buckets remain unchanged. For all hash table operations, keys are preferentially read from the cache, which can reduce access to persistent memory and effectively accelerate the process of locating the target hash bucket. Description of the Drawings

[0056] Figure 1 Schematic diagram of the logical structure of the hash table provided by the embodiment of the present invention;

[0057] Figure 2 Schematic diagram of the implementation of the in-place key placement strategy provided by the embodiment of the present invention;

[0058] Figure 3 Schematic diagram of the lock-free and log-free insertion operation provided by the embodiment of the present invention;

[0059] Figure 4 Schematic diagram of the asynchronous cache loading method provided by the embodiment of the present invention. Detailed implementation manners

[0060] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0061] In the present invention, terms such as "first" and "second" in the present invention and the accompanying drawings (if any) are used to distinguish similar objects and do not necessarily need to describe a specific order or sequence.

[0062] In order to make full use of the high parallelism of the accelerator and reduce the write overhead, thereby providing an efficient data management method for applications and improving the overall performance of the system, the present invention provides a data access method, a controller and a system for a direct-attached persistent memory accelerator. The hash table for the direct-attached persistent memory accelerator can enable the accelerator to directly operate on the hash table through virtual address space technology. The hash table designed by the present invention for the direct-attached persistent memory accelerator uses warp as the execution granularity and coordinates each thread instruction through the communication primitive between warps, reducing the divergence of warps. On this basis, the hash table realizes lock-free hash table operations with crash consistency guarantee through atomic operations and slot states, effectively reducing the overhead of ensuring data crash consistency. In addition, a freeze-based mechanism is used to cache the hash buckets storing hot data to reduce the access to persistent memory and support an asynchronous cache loading strategy to achieve low-overhead cache updates. The present invention can make full use of the high parallelism of the accelerator, has high space utilization and low request latency.

[0063] The following are embodiments.

[0064] Embodiment 1:

[0065] A data access method for a direct-attached persistent memory accelerator, where the persistent memory is used to store key-value pair data, and the data access method includes:

[0066] Create a hash table in persistent memory; the hash table contains N levels, each level contains multiple hash buckets, each hash bucket contains M slots, and each slot is used to store a key-value pair; in two adjacent levels, the number of hash buckets in the upper level is twice the number of hash buckets in the lower level, and each hash bucket in the lower level is shared by two hash buckets in the upper level; the storage location of each key-value pair in the top layer is calculated by K hash functions; where M, N, and K are all preset positive integers, and M×N×K = P, and P is the number of threads in a warp.

[0067] As the load factor of the hash table increases, the hash table will continuously expand to a higher level to achieve capacity expansion. While expanding, the number of levels of the hash table remains unchanged. In this embodiment, the hash table is as Figure 1 shown, which specifically includes 4 levels, that is, N = 4. The 4 levels are the 4th layer to the 7th layer in sequence, where the i-th layer contains 2 i buckets.

[0068] Optionally, in this embodiment, M = 4, K = 2, P = 32; correspondingly, the hash table established in this embodiment adopts the following design decisions:

[0069] (1) Slot association. Each hash bucket in the hash table proposed in this embodiment contains 4 slots, and each slot stores the information of a key-value pair. The key-value pairs in the same hash bucket are always accessed simultaneously. Therefore, the hash table established in this embodiment has 4-way slot association. By utilizing slot association, each hash bucket can handle multiple hash collisions without any data movement and additional persistent memory writes. In addition, since multiple slots in the same hash bucket can be accessed simultaneously, slot association is very friendly to utilizing the parallelism of the accelerator.

[0070] (2) The hash table uses an inter-layer sharing design to handle more hash collisions to achieve higher memory efficiency. Only the hash buckets in the top layer can be located by the hash function, and the buckets in other layers are shared by several hash buckets in the top layer. Each hash bucket in the top layer has multiple shared hash buckets. For example, in Figure 1 , each hash bucket in the 5th layer is shared by 4 buckets in the 7th layer (i.e., the top layer), and each bucket in the 7th layer has 3 shared buckets (in the 4th, 5th, and 6th layers respectively). When inserting a key-value pair, the new key-value pair can be inserted into the located hash bucket and its shared hash buckets. With the help of inter-layer sharing, the hash table established in this embodiment can handle more hash collisions and improve the efficiency of load balancing, thereby achieving higher memory efficiency.

[0071] (3) Multiple hash positions. The hash table uses multiple hash functions to calculate multiple hash positions for each key. As Figure 1As shown, by using two hash functions, each new key-value pair has eight associated hash buckets into which it can be inserted, which further improves memory efficiency.

[0072] (4) Single warp access. By leveraging the parallelism of the warp, the hash table can access the slots of all associated hash buckets for a given key in one go. With a suitable configuration, the hash table can probe the slots of all candidate buckets through a single warp access, and can utilize the high parallelism of the accelerator to achieve both high memory efficiency and high performance simultaneously. As Figure 1 shown, the 32 threads in a warp can simultaneously access the slots of all 32 associated hash buckets for a given key.

[0073] Based on the above-established hash table structure, this embodiment proposes a warp cooperative execution mechanism that allocates hash operations at the thread granularity, and each hash operation is executed at the warp granularity. Specifically, when accessing data in persistent memory, a hash operation is allocated to each thread in the warp, and the threads in the warp are activated in sequence, so that the activated thread executes its allocated hash operation until all the hash operations allocated to the warp are completed; when the activated thread executes its allocated hash operation, the remaining threads in the warp calculate the mapped hash bucket according to the key corresponding to the allocated hash operation, and pass the information stored in each slot in the hash bucket to the activated thread through the inter-thread communication primitive, so that the activated thread determines the target slot corresponding to the currently to-be-executed hash operation based on the received information and executes the hash operation.

[0074] Optionally, in this embodiment, the 32 threads in the warp are numbered in sequence from 0 to 31, and the threads in the warp are activated in sequence according to the thread numbers. Based on the warp cooperative execution mechanism proposed in this embodiment, it is possible to minimize warp divergence on the basis of fully utilizing the high parallelism of the accelerator, thereby effectively improving system performance.

[0075] To reduce the overhead of ensuring data crash consistency, in this embodiment, the information stored in each slot further includes: an 8-byte slot status. The slot status can be idle, being inserted, and inserted. Based on the setting of this slot status, this embodiment can use atomic operations to update the slot status, and can also be restored through the slot status when a crash occurs, realizing lock-free and log-free hash operations.

[0076] Furthermore, to improve memory utilization and quickly implement slot positioning, this embodiment will determine the corresponding placement strategy according to the key length of the key-value pair, and preferentially choose to adopt the in-place key placement strategy, that is, directly store the key in the slot. Specifically, considering that the object size of the atomic operation is within 8 bytes, as Figure 1 shown, according to the different lengths of the key, the corresponding placement strategies are as follows:

[0077] (a) If the key length of the key-value pair is fixed and the length of the key does not exceed 8 bytes, the key-value pair information stored in each slot includes: a pointer to the value in the key-value pair. After the key-value pair information is successfully inserted, the slot status is set to the key of the key-value pair to indicate that the status of the slot is inserted.

[0078] (b) If the key length of the key-value pair is fixed and the length of the key exceeds 8 bytes, the key-value pair information stored in each slot includes: the key in the key-value pair and a pointer to the value in the key-value pair. After the key-value pair information is successfully inserted, the slot status is set to the hash value of the key in the key-value pair to indicate that the status of the slot is inserted.

[0079] (c) If the key length of the key-value pair is not fixed, the key-value pair information stored in each slot includes: a pointer to the key-value pair. After the key-value pair information is successfully inserted, the slot status is set to the hash value of the key in the key-value pair to indicate that the status of the slot is inserted.

[0080] As Figure 2 shown, the above placement strategies (a) and (b) are in-place key placement strategies, and the above placement strategy (c) is a pointer-based key placement strategy. When the in-place key placement strategy is adopted, when the threads in a warp access the keys in the same hash bucket, these accesses can be merged. Since the value is usually accessed by a dedicated thread in a warp, the hash table still stores the pointer to the value.

[0081] It is easy to understand that in the above placement strategies (a) and (b), the value will be stored in a dedicated storage area in persistent memory; in the above placement strategy (c), the entire key-value pair will be stored in a dedicated storage area in persistent memory.

[0082] In this embodiment, the hash operations include insert operation, delete operation, update operation, and query operation, and each operation can be executed without locks and without logging. The lock-free and log-free operation mechanism is as follows:

[0083] As Figure 3 shown, if the current hash operation to be executed is an insert operation, the activated thread determines the target slot corresponding to the current hash operation to be executed based on the received information, and executes the hash operation, including:

[0084] S1: The activated thread uses the key of the key-value pair to be inserted as the target key k i , and compares the target key k i with the keys stored in each slot. If the target key k i does not exist, obtain an idle slot as the target slot, and use the CAS primitive to change the status of the target slot from idle to being inserted;

[0085] In order to balance the load among hash buckets, in step S1 of this embodiment, when there are idle slots in all of the associated multiple hash buckets, an idle slot is preferentially obtained from the hash bucket with the lightest load;

[0086] S2: If the status change fails, which means the slot status has been changed by other threads, then proceed to step S3 to re - execute the insertion operation; otherwise, proceed to step S4;

[0087] S3: Notify the remaining threads in the warp to calculate the mapped hash bucket according to the key corresponding to the assigned hash operation, and transfer the information stored in each slot in the hash bucket to the activated thread through the inter - thread communication primitive, and then proceed to step S1;

[0088] S4: Insert the key - value pair information of the key - value pair to be inserted into the target slot, and then set the status of the target slot to inserted, and the insertion operation ends.

[0089] If the currently to - be - executed hash operation is a deletion operation, the activated thread determines the target slot corresponding to the currently to - be - executed hash operation based on the received information, and performs the hash operation, including:

[0090] The activated thread uses the key of the key - value pair to be deleted as the target key k d , and compares the target key k d with the keys stored in each slot to locate all the slots storing the target key k d as the target slots;

[0091] The activated thread sets the slot status of each target slot to idle through the CAS primitive, and the deletion operation ends; due to the atomicity of the CAS primitive, even in the case of a crash, the deletion operation will not introduce any invalid slot status.

[0092] If the currently to - be - executed hash operation is an update operation, the activated thread determines the target slot corresponding to the currently to - be - executed hash operation based on the received information, and performs the hash operation, including:

[0093] The activated thread uses the key of the key - value pair to be deleted as the target key k u , and compares the target key k u with the keys stored in each slot to locate all the slots storing the target key k u , take one of them as the target slot, and the rest as duplicate slots;

[0094] The activated thread sets the slot status of each duplicate slot to idle through the CAS primitive and updates the key-value pair information in the target slot to the new key-value pair information through the CAS primitive. It is easy to understand that before updating the key-value pair information in the target slot, the new value is written into the pre-allocated space. Due to the atomicity of the CAS primitive, if a crash occurs, the value pointer in the slot either points to the old value or the new value, and both are valid.

[0095] If the current hash operation to be executed is a query operation, the activated thread determines the target slot corresponding to the current hash operation to be executed based on the received information and executes the hash operation, including:

[0096] T1: The activated thread uses the key of the key-value pair to be deleted as the target key k s , and compares the target key k s with the keys stored in each slot to locate all the slots that store the target key k s .

[0097] T2: If the number of slots that store the target key k s is 0, return query failure; otherwise, go to T3;

[0098] T3: The activated thread takes one of all the slots that store the target key k s as the target slot and the rest as duplicate slots;

[0099] T4: The activated thread obtains the value pointer according to the key-value pair information stored in the target slot, reads the corresponding value and returns it;

[0100] Since this embodiment uses the atomicity of the CAS primitive to perform insertion, deletion, and update operations, a lock-free search operation can be easily implemented.

[0101] As the load factor of the hash table increases, more hash conflicts will occur in the hash index, resulting in performance degradation and insertion failure. Thanks to the single warp access method adopted in this embodiment, the hash table will not experience performance degradation due to more hash conflicts. However, the hash table still needs to handle insertion failures to avoid loss of key-value pairs. When performing an insertion operation, if no idle slot can be found to insert a new entry, the hash table needs to be expanded to avoid loss of key-value pairs due to inability to insert. The expansion operation specifically includes:

[0102] Allocate a new level above the top level of the hash table as the top level;

[0103] Scan the bottom layer of the hash table to obtain the key-value pair information stored in the bottom layer, and re-insert the key-value pairs stored in the bottom layer into the hash table through an insertion operation; when scanning the bottom layer of the hash table, a large number of accelerator threads can be used for parallel scanning, and the re-insertion of the key-value pairs stored in the bottom layer into the hash table is also carried out through a warp cooperative execution method to reduce warp divergence;

[0104] After all the key-value pair information stored in the bottom layer is re-inserted into the hash table, delete the bottom layer; after deleting the bottom layer, a layer above the bottom layer in the original hash table will become the new bottom layer.

[0105] To further improve the overall performance of the system, this embodiment further proposes a cache mechanism based on a freezing mechanism to reduce access to persistent memory and accelerate the positioning of target slots. The cache mechanism based on the freezing mechanism includes the following three aspects:

[0106] (1) Create a cache in the memory of the accelerator and cache part of the content in the hash table in units of hash buckets. To facilitate determining the mapping relationship between the hash buckets in the hash table and the hash buckets in the cache, store two mapping functions: F(x) = y, which means that bucket x in the hash table is mapped to bucket y in the cache; and G(y) = x, which means that the hash bucket y stored in the cache stores bucket x in the hash table.

[0107] (2) Regularly identify the hash buckets in the hash table that store hot data and load them into the cache, while dynamically adapting to changes in hot data and maintaining low cache management overhead. Based on the load characteristics, to identify the hash buckets that store hot data, cache algorithms such as LFU and LRU can be used. Since there are a large number of accelerator threads, it is very important to avoid competition between threads when managing the cache. Therefore, between two adjacent loading phases, the hash buckets in the cache remain unchanged.

[0108] (3) For all hash table operations, the hash table preferentially reads keys from the cache, which accelerates the process of locating the target hash bucket. After locating the target bucket, for a search operation, if the target hash bucket is cached, the hash table can directly read the value through the pointer in the cached hash bucket, thereby further reducing access to persistent memory. For insert, delete, and update operations, if the target hash bucket is cached, after performing the corresponding hash operation to modify the hash bucket in persistent storage, it is necessary to make the same modifications to the hash buckets in the cache that are mapped to these hash buckets.

[0109] This embodiment adopts an asynchronous cache loading method to load the hash buckets in persistent memory into the cache, such as Figure 4As shown in the figure. To load the hash bucket x in the persistent memory into the hash bucket y in the cache, the old cache mapping is invalidated by setting F(G(y)) as uncached. To avoid inconsistency in the ongoing operations on the hash bucket y, the hash table needs to wait for these operations to complete (i.e., wait for the reference count of the hash bucket y to become 0) before updating the hash bucket y. Then, the content of the hash bucket x is copied to the hash bucket y. Finally, the new cache mapping becomes effective by setting F(x) as y and G(y) as x. The mapping relationship between the hash buckets is updated atomically by using the CAS primitive, thus ensuring the correctness of concurrency. A concurrent cache loading mechanism is implemented by using another accelerator stream to load the hash buckets in parallel.

[0110] Generally speaking, this embodiment establishes a hash table structure adapted to the direct-attached persistent memory accelerator. On this basis, through the warp system execution mechanism, the warp divergence is minimized, effectively improving the system performance; the slot content is specially designed, and atomic operations are used to update the slot status and data pointer, and recovery is performed through the slot status when a crash occurs, realizing lock-free and log-free operations, reducing the overhead of ensuring data crash consistency; the cache mechanism based on the freeze mechanism reduces the number of accesses to the persistent memory and accelerates the positioning of the target slot, which can further improve the system performance.

[0111] Embodiment 2:

[0112] A storage controller for a direct-attached persistent memory accelerator, comprising:

[0113] A computer-readable storage medium for storing a computer program;

[0114] And a processor for reading the computer program stored in the computer-readable storage medium and executing the data access method for a direct-attached persistent memory accelerator provided in the above Embodiment 1.

[0115] Embodiment 3:

[0116] A key-value storage system, comprising: a persistent memory, a direct-attached persistent memory accelerator, and the storage controller for a direct-attached persistent memory accelerator provided in the above Embodiment 2.

[0117] Those skilled in the art can easily understand that the above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present invention should be included in the protection scope of the present invention.

Claims

1. A data access method for a direct-connected persistent memory accelerator, wherein the persistent memory is used to store key-value pair data, characterized in that: The data access method comprises: A hash table is established in the persistent memory; the hash table includes N levels, each level includes multiple hash buckets, each hash bucket includes M slots, and each slot is used to store a key-value pair; in two adjacent levels, the number of hash buckets included in the upper level is twice the number of hash buckets included in the lower level, and each hash bucket in the lower level is shared by two hash buckets in the upper level; the storage location of each key-value pair in the top level is calculated by K hash functions; When accessing data in the persistent memory, a hash operation is assigned to each thread in the thread warp, and the threads in the thread warp are activated in sequence, so that the activated thread executes its assigned hash operation until all the hash operations assigned to the thread warp are executed; when the activated thread executes its assigned hash operation, the remaining threads in the thread warp calculate the mapped hash bucket according to the key corresponding to the assigned hash operation, and transmit the information stored in each slot in the hash bucket to the activated thread through the communication primitive between threads, so that the activated thread determines the target slot corresponding to the hash operation to be executed currently based on the received information, and executes the hash operation; Wherein, M, N, and K are all preset positive integers and satisfy M×N×K=P, where P is the number of threads in the warp.

2. The data access method for a direct-connected persistent memory accelerator according to claim 1, characterized in that: The information stored in each slot also includes: 8 bytes of slot status; the slot status includes: idle, inserting and inserted; and: If the key length of the key-value pair is fixed and the key length does not exceed 8 bytes, the key-value pair information stored in each slot includes: a pointer to the value in the key-value pair, and after the key-value pair information is successfully inserted, the slot state is set to the key of the key-value pair to indicate that the state of the slot is inserted; If the key length of the key-value pair is fixed and the key length exceeds 8 bytes, the key-value pair information stored in each slot includes: the key in the key-value pair and the pointer to the value in the key-value pair, and after the key-value pair information is successfully inserted, the slot state is set to the hash value of the key in the key-value pair to indicate that the slot state is inserted; If the key length of the key-value pair is not fixed, the key-value pair information stored in each slot includes: a pointer to the key-value pair, and after the key-value pair information is successfully inserted, the slot state is set to the hash value of the key in the key-value pair to indicate that the state of the slot is inserted.

3. The data access method for a direct-connected persistent memory accelerator according to claim 2, characterized in that: If the hash operation to be executed is an insert operation, the activated thread determines the target slot corresponding to the hash operation to be executed based on the received information and executes the hash operation, including: S1: The activated thread takes the key of the key-value pair to be inserted as the target key k i , the target key k i Compare with the keys stored in each slot. If the target key k i If it does not exist, an idle slot is obtained as the target slot, and the CAS primitive is used to change the state of the target slot from idle to inserting; S2: If the state change fails, go to step S3; otherwise, go to step S4; S3: notify the remaining threads in the warp to calculate the mapped hash bucket according to the key corresponding to the assigned hash operation, and pass the information stored in each slot in the hash bucket to the activated thread through the communication primitive between threads, and then proceed to step S1; S4: Insert the key-value pair information of the key-value pair to be inserted into the target slot, and then set the target slot state to inserted, and the insertion operation ends.

4. The data access method for a direct-connected persistent memory accelerator according to claim 3, characterized in that: If in step S1, the target key k i If it does not exist, if the acquisition of the free target slot fails, the expansion operation is performed and then the free slot is acquired as the target slot; The capacity expansion operation includes: Allocating a new level above the top level of the hash table as a top level; Scan the bottom layer of the hash table to obtain the key-value pair information stored in the bottom layer, and reinsert the key-value pairs stored in the bottom layer into the hash table through an insert operation; After all the key-value pair information stored in the bottom layer is reinserted into the hash table, the bottom layer is deleted.

5. The data access method for a direct-connected persistent memory accelerator according to claim 2, characterized in that: If the hash operation to be executed is a delete operation, the activated thread determines the target slot corresponding to the hash operation to be executed based on the received information and executes the hash operation, including: The activated thread takes the key of the key-value pair to be deleted as the target key k d , the target key k d Compare with the keys stored in each slot to locate the key k stored in the target slot d All slots of as target slots; The activated thread sets the slot status of each target slot to idle through the CAS primitive, and the deletion operation ends.

6. The data access method for a direct-connected persistent memory accelerator according to claim 2, characterized in that: If the hash operation to be executed is an update operation, the activated thread determines the target slot corresponding to the hash operation to be executed based on the received information and executes the hash operation, including: The activated thread takes the key of the key-value pair to be deleted as the target key k u , the target key k u Compare with the keys stored in each slot to locate the key k stored in the target slot u Of all the slots in the , one is used as the target slot and the rest are used as repeat slots; The activated thread sets the slot state of each duplicate slot to idle through the CAS primitive, and updates the key-value pair information in the target slot to the new key-value pair information through the CAS primitive.

7. The data access method for a direct-connected persistent memory accelerator according to claim 2, characterized in that: If the hash operation to be executed is a query operation, the activated thread determines the target slot corresponding to the hash operation to be executed based on the received information and executes the hash operation, including: T1: The activated thread takes the key of the key-value pair to be deleted as the target key k s , the target key k s Compare with the keys stored in each slot to locate the key k stored in the target slot s All slots of; T2: If the target key k is stored s If the number of slots is 0, the query fails; otherwise, it goes to T3; T3: The activated thread stores the target key k s One of all the slots in is used as the target slot, and the rest are used as repeat slots; T4: The activated thread obtains the value pointer according to the key-value pair information stored in the target slot, reads the corresponding value and returns it.

8. The data access method for a direct-connected persistent memory accelerator according to any one of claims 1 to 7, characterized in that: Also includes: Creating a cache in a memory of the accelerator; Periodically identifying a hash bucket storing hotspot data in the hash table as a hotspot hash bucket, and loading the identified hotspot hash bucket into the cache; Furthermore, after each thread in the warp calculates the mapped hash bucket according to the key corresponding to the assigned hash operation, the method further includes: accessing the cache to find a slot storing a target key; Furthermore, if the cache hits and the hash operation to be executed is a query operation, it also includes: Obtain key-value pair information from the slot storing the target key in the cache, access the value according to the key-value pair information and return, and end the query operation; Furthermore, if the cache hits, and the hash operation to be executed currently is not a query operation, then the remaining threads in the warp transmit the information stored in each slot in the hash bucket to the activated thread through the inter-thread communication primitive, so that the activated thread determines the target slot corresponding to the hash operation to be executed currently based on the received information, and after executing the hash operation, it also includes: The hash bucket in the cache is modified accordingly according to the execution result of the hash operation.

9. A storage controller for a direct-connected persistent memory accelerator, characterized in that: include: A computer-readable storage medium for storing a computer program; And a processor, configured to read the computer program stored in the computer-readable storage medium and execute the data access method for a direct-connected persistent memory accelerator as described in any one of claims 1-8.

10. A key-value storage system, characterized in that: include: Persistent memory, a directly connected persistent memory accelerator, and a storage controller for the directly connected persistent memory accelerator as described in claim 9.