An efficient cache optimization method based on GPU B-Tree
By using GPU high-performance B-Tree and shared memory to cache hotspot data in CPU-GPU heterogeneous computing, the frequent data exchange and exchanging problems caused by limited GPU video memory are solved, and more efficient data access and query speed is achieved.
Patent Information
- Application Number
- CN202111555609.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-17
- Publication Date
- 2025-08-19
- Estimated Expiration
- 2041-12-17
AI Technical Summary
In CPU-GPU heterogeneous computing, it is difficult for the prior art to effectively utilize the high-performance characteristics of the GPU for indexing operations. Due to the limited GPU video memory, data is frequently swapped in and out, which increases the data communication overhead between the CPU and the GPU.
Using GPU high-performance B-Tree as the index data structure, combined with GPU shared memory, an efficient cache mechanism is designed to retain hotspot data on the GPU, reduce data swap out, cache hotspot data items through hash tables, set up a cache replacement mechanism, and improve cache hit rate.
Improve data access efficiency, reduce data communication overhead between CPU and GPU, and improve data query speed and cache hit rate.
Smart Images

Figure CN114253721B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to CPU-GPU heterogeneous computing, and in particular to a method for accelerating indexing operations using a GPU when the CPU is accessing large amounts of data. Specifically, it relates to a method for implementing efficient data caching on a GPU using a high-performance GPU B-Tree as an index data structure. Background Art
[0002] To address the ever-increasing volume of data and the demand for higher computing performance, CPU-GPU heterogeneous computing can be used for acceleration. Traditional CPUs are designed to handle logically complex transactional computations, spending more execution time on controlling instruction flows rather than performing calculations. This has become a bottleneck limiting the efficient service delivery of big data systems. GPUs boast more abundant computing resources, thread resources, and higher memory bandwidth than CPUs, making them effective for big data processing and graphics computing. In heterogeneous systems, GPUs act as coprocessors for the CPU, achieving acceleration by standardizing data communication and offloading some computations to the GPU.
[0003] The key to big data storage lies in efficient data retrieval and processing. The data that needs to be processed can be managed and used on the GPU. The key technologies for accelerating data access using GPUs include the following three aspects:
[0004] First, a concurrent index data structure suitable for the GPU is used to effectively organize the underlying data. Common index data structures, such as hash tables and B-Trees, organize data based on keywords. Indexing technology on the CPU is already relatively mature, and when building concurrent index data structures on the GPU, it is necessary to reasonably consider the characteristics of the GPU. Due to the large differences between the GPU and the CPU, it is necessary to avoid problems such as disordered memory access or execution divergence when accessing the data structure, so as to obtain the ideal acceleration effect. Currently, data structures that have been successfully constructed on the GPU include LSM, hash tables, B-Trees, etc. The index with B-Tree as the data structure can maintain the orderliness of the key values and can support exact matching point queries and range queries, so B-Tree is used as the GPU data index in the present invention.
[0005] The second approach is to use GPUs to accelerate the construction and traversal of index data structures. GPUs offer high-performance computing capabilities, high concurrency, and high access speeds. These features can be leveraged to improve the efficiency of index operations, including index construction, data addition, deletion, and modification, and a series of other operations. Currently, GPU high-performance B-Trees can support concurrent point queries, range queries, and subsequent queries, as well as dynamic updates, insertions, and deletions. Specific implementation methods include: implementing thread collaboration and parallel processing through the WCWS strategy, achieving dynamic updates through a locking mechanism, correctly handling node overflows through an active splitting strategy, and avoiding spins by restarting the build.
[0006] Third, CUDA's memory copy technology is used to enable data communication between the CPU and GPU. The operating system already handles page swaps between the CPU's hard drive and physical memory, so the graphics driver only needs to implement data transfer from physical memory to GPU memory. By default, data is transferred from the system's paged memory to page-locked memory before being transferred to GPU memory. GPU memory is much smaller than CPU memory, and the amount of data that can be stored on the GPU is limited. Therefore, the data swapping caused by large amounts of data access also imposes performance overhead on GPU processing. Summary of the Invention
[0007] This invention improves data access efficiency by using the current GPU high-performance B-Tree as an index data structure. At the same time, to address the limited amount of data placed on the GPU, a high-efficiency GPU cache mechanism is designed and implemented, retaining a small portion of hot data on the GPU, thereby reducing data swapping and lowering data communication overhead between the CPU and GPU. The main technical content includes two aspects: first, leveraging the high-performance computing characteristics of the GPU, the construction process of the index data structure and the corresponding data operations are all performed on the GPU, allowing it to perform these tasks instead of the CPU, thereby improving data search efficiency; second, utilizing the GPU shared memory to effectively cache frequently used hot data items in a hash storage manner, allowing a small portion of data to remain on the GPU for a long time, reducing the frequency of data page swapping between the CPU and GPU, and minimizing data transmission overhead.
[0008] The present invention provides a method for organizing video memory data based on B-Tree and using GPU shared memory for data caching. The method steps are as follows: Figure 1 ,include:
[0009] 1. Index construction phase
[0010] The present invention uses a GPU B-Tree as the index data structure for video memory data. First, the CPU transfers a certain amount of data to the GPU through memory copying, places it in the video memory, and offloads the data search and calculation tasks to the GPU. The GPU can then concurrently and quickly build an index based on this data for subsequent data searches and dynamic updates. When copying memory data to the GPU video memory, a batch build method is used to create a B-Tree. The specific process draws on the construction method of GPU high-performance B-Trees and is improved to the following steps:
[0011] (1) According to the GPU memory and cache line size, determine the B-Tree structure. First, the node size of each B-Tree
[0012] They are all set to be consistent with the GPU cache line. The data of the entire node can be transmitted in one access. 32 threads merge to access the tree nodes in the form of thread bundles. Secondly, the nodes on the same layer of the tree are linked by pointers. After being connected, each level of the B-Tree is a complete linked list. In addition, the B-Tree used in the present invention only stores key-value pairs on leaf nodes, and the middle nodes of the tree only store key-value information to judge branches. The size of the video memory determines the amount of data, and the number of key-value pairs determines the height of the tree. The final constructed B-Tree index is in the form of Figure 2 .
[0013] (2) Based on the batch of key-value pairs, a large number of threads are started to build the tree concurrently. First, sort the data and keep the zeroth
[0014] The key value is used as the root node, and the remaining nodes are organized from the leaf nodes in a hierarchical order from left to right, building in a bottom-up manner. Secondly, when the amount of data is small, half of the data capacity is placed in each node at a time to avoid the situation where a large number of nodes will be split when inserting data after the construction is completed.
[0015] (3) After the batch construction is completed, the tree is not full, and subsequent data needs to be inserted to fill the tree nodes.
[0016] After a basic B-Tree is constructed in batches, subsequent data needs to be inserted into the tree nodes to fill it up. This begins by traversing downward from the root node, finding the appropriate node based on the key values along the intermediate paths. If the node is not full, data can be inserted directly. If it is full, the node must be split. This split uses a locking mechanism to prevent changes to the node by other threads.
[0017] 2. Cache establishment phase
[0018] The present invention provides a caching solution that eliminates the time and space waste generated on the index traversal path and reduces data swapping overhead by improving the cache hit rate. First, the GPU extracts node data into the cache line for branch judgment every time data is accessed. This process wastes a certain amount of time and space. We take advantage of the locality of access to selectively cache data items to avoid the additional overhead caused by index traversal. Secondly, the cached data resides on the GPU for a longer period of time. Even if the data on the video memory is replaced, it can ensure that commonly used data can be accessed on the GPU at any time. The cache design scheme of the present invention specifically includes the following four parts:
[0019] (1) Allocate cache space. We use GPU shared memory as cache space. Shared memory has fast access speed and can be accessed by all threads in a thread block. However, it is small in size. We set the cache to the size of one shared memory on the current device and use the threads of one thread block at a time to operate, thereby allocating as much cache space as possible.
[0020] (2) Cache data items. Key-value pairs of data are directly stored in the cache, and the result can be directly returned if a hit occurs. Due to the limited cache size, all cached data needs to be swapped in and out in a reasonable manner to ensure that the data in the cache are frequently used hot items and the hit rate is guaranteed. To this end, the present invention designs a data organization scheme within the cache. The main technical ideas include the following points:
[0021] ●In the cache, key-value pairs are organized in a hash table. The key-value pairs in the cache do not have data association. To facilitate quick query, the hash tag can be calculated based on the key of the data. Each time the data is accessed, the target data can be directly queried. The format of the cache hash table is as follows: Figure 3 .
[0022] Use counting to distinguish hot data. When the cache is full, it is necessary to determine whether to swap other data from the video memory into the cache. A counter is bound to each key-value pair in the cache, which serves as a criterion for judging the access popularity of cached data. Each time a cache hit occurs, the counter value of the current data is incremented by 1. Data items with larger counter values are considered hot data, while data items with smaller counter values, due to their fewer accesses, are considered infrequently used data.
[0023] ● Handling counter overflows. The counter size is fixed. After a period of frequent access, the counters of some hot items may overflow. In this case, a random non-zero counter in the current cache is selected and decremented by 1. This random selection ensures the relative popularity of the data in the cache.
[0024] (3) Data replacement mechanism. When the cache misses, the data on the GPU memory will be further accessed, relying on the B-Tree traversal index that has been built. If the cache is not full at this time, the query result can be directly put into the cache; if the cache is full, it is necessary to track the access frequency of the uncached data. At this time, a temporary counter is established for the currently accessed data item. If the value of the temporary counter exceeds the specified threshold within a certain time range, the data item and the counter are swapped into the cache together. In the worst case, the GPU memory does not find the corresponding data. At this time, new data must be introduced from the CPU side. This situation is a situation that can be avoided as much as possible by establishing a cache.
[0025] (4) Ensure data consistency. When data on the CPU is changed (updated or deleted), the data on the GPU must be updated synchronously. If the data is already in the GPU memory, it is updated accordingly, including synchronous updates in the cache and memory, to ensure that the resident data on the GPU remains consistent with the original data on the CPU.
[0026] 3. Data query stage
[0027] After the index and cache are established in the first two stages, each time the host program initiates a data query, it will first access the GPU cache, and then perform subsequent operations based on the cache hit situation. The CPU side is the host side and the GPU side is the device side. The specific algorithm process is as follows:
[0028] (1) The device receives the target data item to be queried and allocates space in the video memory to retain the key of the data.
[0029] (2) The device queries the cache to determine whether the target item is in the cache. If it is, the key-value pair stored in the table is directly returned to the host; if not, it is a cache miss and the next step is performed.
[0030] (3) The device traverses the B-Tree index on the video memory to perform data query. When it reaches a leaf node, if the target key value is found, it determines whether to perform the replacement operation according to the set cache replacement mechanism and returns the result to the host. If the target item is not found, it proceeds to the next step.
[0031] (4) When the target item is not in the device-side video memory, it is necessary to delete the existing video memory data, copy the new data from the host side to the device side, rebuild the index, and return to the first step to restart the query. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Figure 1 :Flowchart of efficient cache optimization method based on B-Tree
[0033] Figure 2 : Schematic diagram of B-Tree index structure
[0034] Figure 3 : Cache hash intent DETAILED DESCRIPTION
[0035] The hardware environment of the present invention is mainly an experimental server. The server CPU is an Intel(R) Core(TM) i5-4570, 3.20GHz, with 4GB RAM, a Quadro K620 GPU, 2GB video memory, and 48KB shared memory. It also uses a 64-bit operating system.
[0036] The software implementation of the present invention is developed in C language under the CUDA framework on Linux platform. The CUDA driver version is 10.1 and the gcc version is 7.0.
[0037] The experimental data is an integer key-value pair.
[0038] The operation is mainly divided into two parts, the first part is to establish a cache hash table, and the second part is data query.
[0039] 1. Create a cache hash table
[0040] 1) Initialize the hash table
[0041] The process of initializing a hash table includes allocating space, selecting the data storage format, defining the hash element format, and specifying the hash tag calculation method.
[0042] The specific algorithm steps are as follows:
[0043] (1) Use CUDA Memlloc to allocate memory on GPU shared memory to ensure that this part of memory is not occupied or modified by other threads, and store cache data more permanently;
[0044] (2) Define the storage object HashMember, which contains not only key value data but also a bound counter;
[0045] (3) Define the size of the hash bucket, select the chain storage method, and hash the data;
[0046] (4) Specify the index calculation method. Initially, the hash table is empty. When the hash table is not full, elements are continuously inserted according to the data access situation. When the hash table is full, a replacement strategy is adopted.
[0047] 2) Hash query
[0048] Algorithm input: hashtable, key
[0049] Algorithm output: result
[0050] Description: hashtable is a hash table, its format is as follows Figure 3 As shown in the example, key is the target key and result is the search result. If the key value is found, it returns the key value; if not, it returns NOT_FOUND.
[0051] Algorithm steps:
[0052] (1) Calculate the hash tag based on the target key;
[0053] (2) Find the linked list where the target value is located according to the label;
[0054] (3) Loop through the current linked list until the target key is found or the linked list is empty;
[0055] (4) After the target key value is successfully queried, the bound counter is incremented and it is determined whether the counter will overflow. If the counter overflows, another counter is randomly selected to be decremented.
[0056] The algorithm pseudo code is as follows:
[0057]
[0058]
[0059] 2. Data Query
[0060] 1) When the host initiates a query, the device can perform data query after successfully copying the data and creating the index.
[0061] Algorithm Description
[0062] Algorithm input: Btree, key, hashtable
[0063] Algorithm output: result
[0064] Description: Btree is an index built on GPU memory, its format is as follows Figure 2 As shown, hashtable is a hash table, and its format is as follows Figure 3 As shown in the example, key is the target key and result is the search result. If the key value is found, it returns the key value; if not, it returns NOT_FOUND.
[0065] Algorithm steps:
[0066] (1) Initiate a query, first access the cache hash table, and proceed to the second step if the cache misses;
[0067] (2) Traverse the B-Tree from the root node until it reaches the leaf node, using high concurrency to enable multi-threaded traversal;
[0068] (3) When reaching a leaf node, if the result exists, it is returned. If the result does not exist, information needs to be returned to the host, indicating that some data needs to be re-copied from the host.
[0069] The algorithm pseudo code is as follows:
[0070]
[0071]
Claims
1. An efficient cache optimization method based on GPU B-Tree, characterized by The implementation steps are: (1) Copy data from the CPU according to the size of the GPU memory and use B-Tree as the index data structure of the data; (2) Using GPU shared memory to establish a cache, including allocating cache space and setting it to a shared memory size, caching data items using a hash table, setting a replacement mechanism using a temporary counter, and implementing basic operations such as querying, inserting, and updating data within the cache; (3) Initiate a query from the CPU side, give priority to accessing cached data, and return directly if the cache hits. If the cache misses, traverse the index to query the GPU data. If the query is successful, return and replace the data. If the query fails, copy new data from the CPU side.
2. The efficient cache optimization method based on GPU B-Tree according to claim 1 is characterized in that This method is in the data copying stage: (1) Store data according to the size of GPU memory; (2) A B-Tree is built on the video memory as the index data structure of the data, and a query search is performed in a multi-threaded concurrent manner each time the data is accessed.
3. The efficient cache optimization method based on GPU B-Tree according to claim 1 is characterized in that This method has four steps in the cache establishment phase: (1) Allocate cache space. To properly utilize the high-speed access characteristics of GPU shared memory, use as much shared memory as possible to create cache space. Set the cache to the size of one shared memory in the current device. (2) Cache data items. In the cache, data is organized in a hash table format. A counter is bound to each key-value pair as a criterion for judging the access popularity of each cached data item. (3) Setting up a replacement mechanism, which stipulates that within a certain time range, a temporary counter is set for the data that is not in the cache, and the data item with the smallest counter value in the cache is replaced according to the threshold standard; (4) Implement basic operations within the cache, including data query, data insertion, and data update.
4. The efficient cache optimization method based on GPU B-Tree according to claim 1, characterized in that When this method initiates a query search: (1) Prioritize access to cached data and return the result directly if a hit occurs; (2) When the cache misses, the index is traversed to check whether the data already on the GPU contains the query content. If the query is successful, the result is returned and the data is replaced according to the cache replacement mechanism; (3) If the query fails, new data needs to be copied from the CPU.
Citation Information
Patent Citations
Open computing language (OpenCL)-based red-black tree acceleration algorithm
CN104036141A
Adaptive cardinal number tree dynamic indexing method based on GPU parallelism
CN112000847A