Reachability query method and system based on GPU acceleration

Through the GPU-accelerated accessibility query method, combined with the compression preprocessing of graph data and appropriate BFS algorithm, the problems of low efficiency of large-scale graph data query and unbalanced load in the existing technology are solved, and efficient graph data query and resource utilization are achieved.

CN120011601AActive Publication Date: 2025-05-16HUNAN UNIV

Patent Information

Application Number
CN202510088606.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-05-16
Estimated Expiration
2045-01-21

Smart Images

  • Figure CN120011601A_ABST
    Figure CN120011601A_ABST
Patent Text Reader

Abstract

The invention discloses a GPU (Graphics Processing Unit) acceleration-based reachability query method, which comprises the following steps of: receiving graph data from a user, carrying out compression preprocessing on the graph data to obtain compressed graph data, obtaining query points from the user, configuring an environment of a GPU, copying the compressed graph data and the query point into a memory of the GPU in a configured GPU environment, calculating an average out-degree of the graph data, judging whether the average out-degree is greater than a preset threshold value, if so, acquiring the graph data copied into the memory of the GPU, and if not, storing the graph data in the memory of the GPU; the two-stage BFS algorithm is used for querying whether the obtained query point can reach the query result or not in the graph data, and then the query result is copied to a CPU memory and output. The technical problem that an existing index-based reachability query algorithm cannot process large-scale graph data due to the fact that construction of the index for reachability query needs to consume a large number of storage resources can be solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of graph computing technology, and more specifically, relates to a reachability query method and system based on GPU acceleration. Background Art

[0002] The rapid development of the Internet and the advent of the era of big data have led to an increasing demand for rapid analysis of massive amounts of data. A graph is a general and effective data structure for representing many different types of objects. The vertices in a graph are used to describe entities, while the edges in a graph can be used to describe the associations between entities. For example, in a social networking platform, users are regarded as nodes, and the interactions between users are represented by edges, which effectively depicts the connection pattern of the social network. In addition, many non-graph-structured data and applications can also be converted into graph models for processing and analysis, including protein interaction networks, gene regulatory networks, knowledge graphs, and hardware structure design. Due to its wide range of application backgrounds, graph data processing research has become a key focus area in academia and industry.

[0003] Graph computing is a field that studies how to speed up the calculation, storage, and management of graph data under large-scale graph data. All analytical calculations based on graph data belong to graph computing, so the application fields involved in graph computing are very broad. As a classic problem in the field of graph computing, reachability query (RQ) is a research hotspot in computer science, geographic information science, and transportation. It can be used to solve optimization problems such as resource allocation, network connection analysis, and path planning; in addition, reachability query is also a basic technology for problems such as community query in graph theory.

[0004] There have been a lot of research on reachability query technology, including index-based reachability query and distributed reachability query. Among them, index-based reachability query can quickly query the reachability between nodes by pre-building the index structure of the reachability information of graph data. Distributed reachability query is to store large-scale graph data on multiple machines and process queries in parallel through a distributed computing framework.

[0005] However, both of the above two existing reachability query methods have some defects that cannot be ignored: First, the index-based reachability query algorithm will bring a lot of storage overhead when maintaining a large number of index data structures and cannot handle large-scale graph data; Second, the distributed reachability query algorithm will cause load imbalance problems when sharding large-scale graph data, and the frequent data exchange between nodes will bring a lot of communication overhead. Summary of the invention

[0006] In response to the above defects or improvement needs of the prior art, the present invention provides a reachability query method and system based on GPU acceleration, which aims to solve the technical problems that the existing index-based reachability query algorithm cannot process large-scale graph data because it consumes a large amount of storage resources to build an index for reachability query, and the technical problems that the existing distributed reachability query algorithm performs data sharding on large-scale graph data, which leads to load imbalance and high communication overhead.

[0007] To achieve the above object, according to one aspect of the present invention, a reachability query method based on GPU acceleration is provided, comprising the following steps:

[0008] (1) Receive graph data from a user and perform compression preprocessing on the graph data to obtain compressed graph data.

[0009] (2) Obtaining a query point from the user, which includes a source vertex and a target vertex, configuring a GPU environment, and copying the graph data compressed in step (1) and the query point to the GPU memory in the configured GPU environment;

[0010] (3) Calculate the average out-degree of the graph data copied to the GPU memory in step (2), and determine whether the average out-degree is greater than a preset threshold (the value range is 3 to 9, preferably 8). If so, proceed to step (4); otherwise, proceed to step (5);

[0011] (4) Obtain the graph data copied to the GPU memory in step (2), and use the two-stage BFS algorithm to query whether the query point obtained in step (2) is reachable in the graph data as the query result, and then proceed to step (6);

[0012] (5) Obtain the graph data copied to the GPU memory in step (2), and use the bidirectional BFS algorithm to query whether the query point obtained in step (2) is reachable in the graph data as the query result, and then proceed to step (6);

[0013] (6) Copy the query results to the CPU memory and output them.

[0014] Preferably, step (1) specifically includes the following sub-steps:

[0015] (1-1) Receive graph data from the user, set counter i = 0, obtain the total number of edges and vertices in the graph data, and initialize array C and array R to empty;

[0016] (1-2) Determine whether i is equal to the total number of vertices in the graph data. If so, use the obtained array C as the final adjacency list and array R as the final offset table. Array C and array R are the compressed graph data. The process ends. Otherwise, go to step (1-3).

[0017] (1-3) Determine whether the adjacent vertices of the i-th vertex in the graph data exist. If so, add the adjacent vertices to the array C in sequence, and record the position of the first adjacent vertex of the i-th vertex in the graph data in the array C in the array R, and then go to step (1-5), otherwise go to step (1-4);

[0018] (1-4) Set the i-th element R[i] in the array R to the i-1-th element R[i-1], and then go to step (1-5);

[0019] (1-5) Set i=i+1 and return to step (1-2).

[0020] Preferably, the average out-degree σ in step (3) is equal to:

[0021]

[0022] Where n represents the total number of vertices in the graph data, and out_deg(i) represents the out-degree of the i-th vertex in the graph data.

[0023] Preferably, this step specifically includes the following sub-steps:

[0024] (4-1) Create a vertex queue vque, an edge queue eque, and a distance array dis, initialize the vertex queue vque to the source vertex, initialize the edge queue eque to empty, initialize the distance corresponding to the source vertex in the distance array dis to 0, and initialize the distances corresponding to the remaining vertices in the distance array dis to infinity;

[0025] (4-2) Determine whether the vertex queue vque is empty. If so, take the source vertex and the target vertex as unreachable as the query result, and the process ends. Otherwise, determine whether the vertex queue vque contains the target vertex. If so, take the source vertex and the target vertex as reachable as the query result, and the process ends. Otherwise, go to step (4-3);

[0026] (4-3) Get the global ID of the current thread;

[0027] (4-4) Set the counter global_id1 = the global ID of the current thread obtained in step (4-3);

[0028] (4-5) Determine whether global_id1 is less than the length of the vertex queue vque. If so, obtain the vertex vque[global-id1] with global_id1 in the vertex queue vque, and then proceed to step (4-6). Otherwise, it means that all vertices in the vertex queue vque have been processed, and then proceed to step (4-9);

[0029] (4-6) Obtain the size of the adjacency list corresponding to the vertex vque[global-id1] with global_id1 in the vertex queue vque, and obtain the prefix and result of the thread corresponding to the vertex in the thread grid according to the size of the adjacency list, as the offset of the adjacent vertex in the corresponding adjacency list of the vertex vque[global-id1] with global_id1 in the edge queue eque, and store all the adjacent vertices in the edge queue eque;

[0030] (4-7) Set global_id1 = global_id1 + gridDim.x * blockDim.x, and return to step (4-5); where gridDim.x represents the number of thread blocks in the thread grid where the thread of dimension x is located, blockDim.x represents the number of threads contained in the thread block corresponding to the thread of dimension x, and gridDim.x * blockDim.x represents the total number of threads in the thread grid where the thread of dimension x is located;

[0031] (4-8) Copy the edge queue eque obtained in step (4-6) to the shared memory, and use the hash table in the shared memory to perform conflict detection processing on the edge queue eque (i.e., remove duplicate vertices) to obtain the edge queue eque after the conflict detection processing;

[0032] (4-9) Set the counter global_id2 = the global ID of the current thread obtained in step (4-3);

[0033] (4-10) Determine whether global_id2 is less than the length of the edge queue eque. If so, obtain the vertex eque[global_id2] with global_id2 in the edge queue eque, and then go to step (4-11). Otherwise, it means that all vertices in the edge queue eque have been processed, and then return to step (4-2);

[0034] (4-11) Determine whether the distance corresponding to the vertex eque[global_id2] with the global_id2th vertex in the distance array dis is infinite. If so, update the distance corresponding to the vertex eque[global_id2] with the global_id2th vertex in the distance array dis, and add the vertex eque[global_id2] with the global_id2th vertex to the vertex queue vque, otherwise proceed to step (4-12);

[0035] (4-12) Set global_id2=global_id2+gridDim.x*blockDim.x, and return to step (4-10).

[0036] Preferably, step (4-6) specifically includes the following sub-steps:

[0037] (4-6-1) Get the size of the adjacency list corresponding to the vertex vque[global-id1] with the global_id1 in the vertex queue vque, set the counter i=0, initialize the array prefix to be empty, which is used to record the prefix and result of the thread bundle corresponding to the thread corresponding to the vertex with the global_id1 in the thread grid, initialize the array nbr (whose size is 32) to be empty, which is used to record the size of the adjacency list corresponding to the vertex with the global_id1, and initialize the 0th element prefix[0] in the array prefix to 0;

[0038] (4-6-2) Determine whether i is equal to the size of the array prefix - 1. If so, the array prefix is ​​used as the final prefix and result table and the process ends. Otherwise, go to step (4-6-3);

[0039] (4-6-3) Set the i-th element in the array prefix to prefix[i-1]+nbr[i-1], where nbr[i-1] represents the size of the adjacency list corresponding to the vertex of global_id1, and then go to step (4-6-4);

[0040] (4-6-4) Set i=i+1 and return to step (4-6-2).

[0041] Preferably, the global ID of the thread in step (4-3) is equal to:

[0042] ID=blockIdx.x*blockDim.x+threadIdx.x

[0043] Among them, blockIdx.x represents the number of the thread block corresponding to the thread with dimension x in the entire thread grid, blockDim.x represents the number of threads contained in the thread block corresponding to the thread with dimension x, and threadIdx.x represents the number of the thread with dimension x in the thread bundle in the thread grid where it is located.

[0044] Preferably, step (5) specifically includes the following sub-steps:

[0045] (5-1) Create a forward queue fque, a backward queue bque, a forward distance array fdis, and a backward distance array bdis, initialize the forward queue fque as the source vertex, initialize the backward queue bque as the target vertex, initialize the distance corresponding to the source vertex in the forward distance array fdis to 0, and initialize the distances corresponding to the remaining vertices to infinity, initialize the distance corresponding to the target vertex in the backward distance array bdis to 0, and initialize the distances corresponding to the remaining vertices to infinity;

[0046] (5-2) Determine whether the forward distance array fdis and the backward distance array bdis have non-infinite distances at the same position. If so, the source vertex and the target vertex are reachable as the query result, and the process ends. Otherwise, determine whether the forward queue fque and the backward queue bque are both empty. If so, the source vertex and the target vertex are unreachable as the query result, and the process ends. Otherwise, go to step (5-3);

[0047] (5-3) Get the global ID of the current thread;

[0048] (5-4) Set the counter global_id3 = the global ID of the current thread obtained in step (5-3);

[0049] (5-5) Determine whether global_id3 is less than the length of the forward queue fque. If so, obtain the vertex fque[global_id3] with global_id3 in the forward queue fque, and then proceed to step (5-6). Otherwise, it means that all vertices in the forward queue fque have been processed, and then proceed to step (5-8);

[0050] (5-6) Determine whether the distance of the adjacent vertices in the adjacency list corresponding to the global_id3 vertex fque[global_id3] in the forward queue fque in the forward distance array fdis is infinite. If so, update the distance of the adjacent vertices in the adjacency list corresponding to the global_id3 vertex fque[global_id3] in the forward queue fque in the forward distance array fdis and add the adjacent vertices in the adjacency list corresponding to the global_id3 vertex fque[global_id3] in the forward queue fque to the forward queue fque. Otherwise, it means that the global_id3 vertex fque[global_id3] in the forward queue fque has been processed, remove the vertex from the forward queue fque, and then enter step (5-7);

[0051] (5-7) Set global_id3 = global_id3 + gridDim.x*blockDim.x, and return to step (5-5); where gridDim.x represents the number of thread blocks in the thread grid where the thread of dimension x is located, blockDim.x represents the number of threads contained in the thread block corresponding to the thread of dimension x, and gridDim.x*blockDim.x represents the total number of threads in the thread grid where the thread of dimension x is located (if the elements in the vertex queue are much larger than the total number of threads (gridDim.x*blockDim.x), this grid-striding loop mode can enable each thread to process multiple data elements and make rational use of computing resources).

[0052] (5-8) Set the counter global_id4 = the global ID of the current thread obtained in step (5-3);

[0053] (5-9) Determine whether global_id4 is less than the length of the backward queue bque. If so, obtain the vertex bque[global_id4] with global_id4 in the backward queue bque, and then proceed to step (5-10). Otherwise, it means that all vertices in the backward queue bque have been processed, and then return to step (5-2);

[0054] (5-10) Determine whether the distance between the adjacent vertices in the adjacency list corresponding to the global_id4 vertex bque[global_id4] in the backward queue bque and the corresponding distances in the backward distance array bdis are infinite. If so, update the distances between the adjacent vertices in the adjacency list corresponding to the global_id4 vertex bque[global_id4] in the backward queue bque and add the adjacent vertices in the adjacency list corresponding to the global_id4 vertex bque[global_id4] in the backward queue bque to the backward queue bque. Otherwise, it means that the global_id4 vertex bque[global_id4] in the backward queue bque has been processed. Remove the vertex from the backward queue bque and enter step (5-11).

[0055] (5-11) Set global_id4=global_id4+gridDim.x*blockDim.x, and return to step (5-9).

[0056] According to another aspect of the present invention, a reachability query system based on GPU acceleration is provided, comprising:

[0057] The first module is used to receive graph data from a user and perform compression preprocessing on the graph data to obtain compressed graph data.

[0058] The second module is used to obtain query points from the user, including source vertices and target vertices, configure the GPU environment, and copy the graph data compressed and processed by the first module and the query points to the GPU memory in the configured GPU environment;

[0059] The third module is used to calculate the average out-degree of the graph data copied to the GPU memory by the second module, and determine whether the average out-degree is greater than a preset threshold (the value range is 3 to 9, preferably 8). If yes, enter the fourth module, otherwise enter the fifth module;

[0060] The fourth module is used to obtain the graph data copied by the second module to the memory of the GPU, and use the two-stage BFS algorithm to query whether the query point obtained by the second module is reachable in the graph data as the query result, and then the sixth module;

[0061] The fifth module is used to obtain the graph data copied by the second module to the memory of the GPU, and use the bidirectional BFS algorithm to query whether the query point obtained by the second module is reachable in the graph data as the query result, and then transfer to the sixth module;

[0062] The sixth module is used to copy the query results to the CPU memory and output them.

[0063] In general, the above technical solutions conceived by the present invention can achieve the following beneficial effects compared with the prior art:

[0064] (1) The present invention adopts steps (1-1) to (1-5), which mainly realizes compressed storage of large-scale graph data by storing the adjacency list and offset information of the graph data. Therefore, it can solve the problem that the existing index-based reachability query algorithm cannot process large-scale graph data because the index occupies too much storage space;

[0065] (2) The present invention adopts steps (4-6), which mainly provides accurate workload statistics and a dynamic work allocation mechanism by calculating prefix sums, making parallel processing more efficient, thereby solving the problem of load imbalance caused by data fragmentation in existing distributed reachability query algorithms;

[0066] (3) The present invention adopts steps (4-9) to (4-13), which mainly reduces data access delay by performing information interaction in a shared memory, thereby making communication and synchronization faster. Therefore, the problem that the existing distributed reachability query algorithm brings a large communication overhead due to frequent data exchange between nodes can be solved.

[0067] (4) The present invention uses step (3) to adopt different algorithms for graphs with different average out-degrees, which can not only reasonably allocate tasks in parallel computing and improve resource utilization, but also select different data structures to reduce unnecessary calculations and speed up query speed. BRIEF DESCRIPTION OF THE DRAWINGS

[0068] Figure 1 This is a typical schematic diagram of the device memory model in CUDA;

[0069] Figure 2 This is a schematic diagram of the principle of CUDA parallel computing;

[0070] Figure 3 It is a schematic diagram of compression storage in the present invention;

[0071] Figure 4 It is a schematic diagram of prefix sum calculation in the present invention;

[0072] Figure 5 It is a flow chart of a reachability query method based on GPU acceleration of the present invention. DETAILED DESCRIPTION

[0073] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0074] The following first explains and illustrates the technical terms in the present invention:

[0075] Reachability Query: Given a graph G = (V, E), V is the set of vertices and E is the set of edges. Given any two vertices s and t in G, the reachability query is represented by q(s, t). If there is at least one path from s to t in G, then s is said to be reachable to t, represented by q(s, t) = true. If there is no path from s to t, then s is said to be unreachable to t, represented by q(s, t) = false.

[0076] Parallel computing: refers to solving a problem by using multiple processing units (such as CPU cores, GPU cores, computing nodes, etc.) at the same time, thereby speeding up the computing process. The core idea of ​​this method is to decompose the computing task into multiple subtasks that can be executed in parallel, and use multiple processing units to execute these tasks simultaneously, thereby improving efficiency and speed.

[0077] Thread: It is the basic unit for executing computing tasks. GPU can execute a large number of threads at the same time. These threads can execute the same instructions in parallel, thereby speeding up the computing process.

[0078] Warp: It is the smallest unit of NVIDIA GPU hardware scheduling, usually containing 32 threads. All threads in a warp share the same program counter, that is, they execute the same instructions at the same time.

[0079] Thread Block: It is a basic concept in the CUDA programming model in NVIDIA's GPU architecture. It is the organizational unit of threads and is used to perform parallel computing tasks on the GPU. Each thread block contains multiple thread warps, and a thread block is a basic scheduling unit that can be executed in parallel on the GPU.

[0080] Thread Grid: A thread grid is a group of multiple thread blocks. A thread grid provides a way to organize and manage a large number of thread blocks so that programs can be efficiently executed in parallel on the GPU.

[0081] Shared Memory: A storage area that is exclusive to each thread block. All threads in the thread block can access and operate this memory. A major advantage of shared memory is that its access latency is lower than that of global memory, which is suitable for collaboration and data sharing between threads within a thread block.

[0082] Figure 1 It is the framework of the device memory model in the typical Compute Unified Device Architecture (CUDA), which includes: first copying data from the host side (CPU) to the global memory of the device side (GPU), then performing thread-level parallel computing on the device side, and assigning a large number of computing tasks to different threads for computing in a single instruction multiple data (SIMD) manner. The grid is the highest-level thread organization unit, consisting of multiple thread blocks. Threads in the same thread block can choose to store intermediate results in shared memory with faster access speed, which can greatly improve the computing speed. Finally, the calculation results are copied from the device side (GPU) back to the host side (CPU).

[0083] Figure 2 Based on the principle of CUDA parallel computing, in the present invention, a thread block has 256 threads, and the thread global ID is indexed by the thread block number (blockIdx.x) and the thread number (threadIdx.x), and all threads process corresponding transactions according to their own IDs. The GPU includes a large-capacity global memory and a small-capacity shared memory. Threads in a thread block communicate quickly through the shared memory, and then write the results to the global memory for use by other thread blocks.

[0084] The basic idea of ​​the present invention is to achieve kernel-level task parallelism through an optimized GPU-accelerated reachability query method. The present invention first compresses the graph data to obtain graph data that occupies less storage space, and then copies the compressed graph data to the GPU memory for query. The graph data with a larger average out-degree is queried using a two-stage two-stage breadth first search (Breadth First Search, referred to as BFS) algorithm, and the graph data with a smaller average out-degree is queried using a bidirectional BFS algorithm. By using different algorithms and starting an appropriate number of threads according to the structure and resource status of the graph data, the query processing can be executed in parallel.

[0085] like Figure 5 As shown, the present invention provides a reachability query method based on GPU acceleration, comprising the following steps:

[0086] (1) Receive graph data from a user and perform compression preprocessing on the graph data to obtain compressed graph data.

[0087] The advantage of this step is that using compressed graph data can make large-scale graph data occupy less storage space, speed up the data transmission speed between the host and the device, enable the GPU to process larger-scale graph data, and the compressed data structure is more suitable for parallel processing.

[0088] like Figure 3 As shown, this step specifically includes the following sub-steps:

[0089] (1-1) Receive graph data from the user, set counter i = 0, obtain the total number of edges and the total number of vertices in the graph data, and initialize array C (whose size is equal to the total number of edges) and array R (whose size is equal to the total number of vertices) to empty;

[0090] Specifically, array C is used to record the adjacent vertices of all vertices in the graph data, and array R is used to record the offset of the starting positions of the adjacent vertices of all vertices in the graph data in array C.

[0091] More specifically, the array is initialized in this step by obtaining it from the computer using an input command, for example, using the statement "vector <int>C”.

[0092] (1-2) Determine whether i is equal to the total number of vertices in the graph data. If so, use the obtained array C as the final adjacency list and array R as the final offset table. Array C and array R are the compressed graph data. The process ends. Otherwise, go to step (1-3).

[0093] (1-3) Determine whether the adjacent vertices of the i-th vertex in the graph data exist. If so, add the adjacent vertices to the array C in sequence, and record the position of the first adjacent vertex of the i-th vertex in the graph data in the array C in the array R, and then go to step (1-5), otherwise go to step (1-4);

[0094] (1-4) Set the i-th element R[i] in the array R to the i-1-th element R[i-1], and then go to step (1-5);

[0095] (1-5) Set i=i+1 and return to step (1-2).

[0096] (2) Obtaining a query point from the user, which includes a source vertex and a target vertex, configuring a GPU environment, and copying the graph data compressed in step (1) and the query point to the GPU memory in the configured GPU environment;

[0097] Specifically, the GPU environment configured in this step is to install the NVIDIA driver and CUDA toolkit for the GPU, and install the nvcc compiler.

[0098] (3) Calculate the average out-degree of the graph data copied to the GPU memory in step (2), and determine whether the average out-degree is greater than a preset threshold (the value range is 3 to 9, preferably 8). If so, proceed to step (4); otherwise, proceed to step (5);

[0099] Specifically, in graph data, the degree is divided into in-degree and out-degree, which respectively represent the number of edges pointing to a node in the graph data and the number of edges departing from the node.

[0100] In this step, the average out-degree σ is equal to:

[0101]

[0102] Where n represents the total number of vertices in the graph data, and out_deg(i) represents the out-degree of the i-th vertex in the graph data.

[0103] If the average out-degree σ exceeds the preset threshold, it means that the structure of the graph data is relatively dense, and there are many intermediate results generated during the expansion process. It is likely that vertices will need to be deduplicated later. In this case, it is recommended to use the two-stage breadth-first search (BFS) algorithm. On the contrary, if the average out-degree σ does not exceed the threshold, it means that the structure of the graph data is relatively sparse, and there are few intermediate results generated during the expansion process. It is unlikely that vertices will need to be deduplicated later. In this case, it is recommended to use the bidirectional BFS algorithm.

[0104] (4) Obtain the graph data copied to the GPU memory in step (2), and use the two-stage BFS algorithm to query whether the query point obtained in step (2) is reachable in the graph data as the query result, and then proceed to step (6);

[0105] The advantage of this step is that the conflict detection feature of the hash table is used in the shared memory to deduplicate the intermediate results, and prefix sum and atomic operations are used to ensure accuracy and concurrency safety when multiple threads write data at the same time, ensuring full utilization of computing resources.

[0106] like Figure 5 As shown, this step specifically includes the following sub-steps:

[0107] (4-1) Create a vertex queue vque, an edge queue eque, and a distance array dis, initialize the vertex queue vque to the source vertex, initialize the edge queue eque to empty, initialize the distance corresponding to the source vertex in the distance array dis to 0, and initialize the distances corresponding to the remaining vertices in the distance array dis to infinity;

[0108] (4-2) Determine whether the vertex queue vque is empty. If so, take the source vertex and the target vertex as unreachable as the query result, and the process ends. Otherwise, determine whether the vertex queue vque contains the target vertex. If so, take the source vertex and the target vertex as reachable as the query result, and the process ends. Otherwise, go to step (4-3);

[0109] (4-3) Get the global ID of the current thread;

[0110] Specifically, each thread uses the global ID as the initial index to process the vertices in the queue;

[0111] More specifically, the global ID of a thread is equal to:

[0112] ID=blockIdx.x*blockDim.x+threadIdx.x

[0113] Among them, blockIdx.x represents the number of the thread block corresponding to the thread with dimension x in the entire thread grid, blockDim.x represents the number of threads contained in the thread block corresponding to the thread with dimension x, and threadIdx.x represents the number of the thread with dimension x in the thread bundle in the thread grid where it is located. The global ID of the thread in the entire thread grid can be calculated by multiplying the first two and adding the third.

[0114] (4-4) Set the counter global_id1 = the global ID of the current thread obtained in step (4-3);

[0115] (4-5) Determine whether global_id1 is less than the length of the vertex queue vque. If so, obtain the vertex vque[global-id1] with global_id1 in the vertex queue vque, and then proceed to step (4-6). Otherwise, it means that all vertices in the vertex queue vque have been processed, and then proceed to step (4-9);

[0116] (4-6) Obtain the size of the adjacency list corresponding to the vertex vque[global-id1] with global_id1 in the vertex queue vque, and obtain the prefix and result of the thread corresponding to the vertex in the thread grid according to the size of the adjacency list, as the offset of the adjacent vertex in the corresponding adjacency list of the vertex vque[global-id1] with global_id1 in the edge queue eque, and store all the adjacent vertices in the edge queue eque;

[0117] like Figure 4 As shown, this step specifically includes the following sub-steps:

[0118] (4-6-1) Get the size of the adjacency list corresponding to the vertex vque[global-id1] with the global_id1 in the vertex queue vque, set the counter i=0, initialize the array prefix (whose size is 32) to be empty, which is used to record the prefix and result of the thread bundle corresponding to the thread corresponding to the vertex with the global_id1 in the thread grid, initialize the array nbr (whose size is 32) to be empty, which is used to record the size of the adjacency list corresponding to the vertex with the global_id1, and initialize the 0th element prefix[0] in the array prefix to 0;

[0119] (4-6-2) Determine whether i is equal to the size of the array prefix - 1. If so, the array prefix is ​​used as the final prefix and result table and the process ends. Otherwise, go to step (4-6-3);

[0120] (4-6-3) Set the i-th element in the array prefix to prefix[i-1]+nbr[i-1], where nbr[i-1] represents the size of the adjacency list corresponding to the vertex of global_id1, and then go to step (4-6-4);

[0121] (4-6-4) Set i=i+1 and return to step (4-6-2).

[0122] The advantage of the above steps (4-6-1) to (4-6-4) is that the prefix and result are used to determine the offset of each thread in the writing process of the edge queue eque, so that multiple threads can write intermediate results to the edge queue in parallel, reasonably allocate the edge queue space, reduce thread competition, and achieve better load balancing.

[0123] (4-7) Set global_id1 = global_id1 + gridDim.x * blockDim.x, and return to step (4-5); where gridDim.x represents the number of thread blocks in the thread grid where the thread of dimension x is located, blockDim.x represents the number of threads contained in the thread block corresponding to the thread of dimension x, and gridDim.x * blockDim.x represents the total number of threads in the thread grid where the thread of dimension x is located;

[0124] If the number of elements in the vertex queue vque is much larger than the total number of threads (gridDim.x*blockDim.x), this grid-striding loop mode can enable each thread to process multiple data elements and make rational use of computing resources.

[0125] (4-8) Copy the edge queue eque obtained in step (4-6) to the shared memory, and use the hash table in the shared memory to perform conflict detection processing on the edge queue eque (i.e., remove duplicate vertices) to obtain the edge queue eque after the conflict detection processing;

[0126] (4-9) Set the counter global_id2 = the global ID of the current thread obtained in step (4-3);

[0127] (4-10) Determine whether global_id2 is less than the length of the edge queue eque. If so, obtain the vertex eque[global_id2] with global_id2 in the edge queue eque, and then go to step (4-11). Otherwise, it means that all vertices in the edge queue eque have been processed, and then return to step (4-2);

[0128] (4-11) Determine whether the distance corresponding to the vertex eque[global_id2] with the global_id2th vertex in the distance array dis is infinite. If so, update the distance corresponding to the vertex eque[global_id2] with the global_id2th vertex in the distance array dis, and add the vertex eque[global_id2] with the global_id2th vertex to the vertex queue vque, otherwise proceed to step (4-12);

[0129] (4-12) Set global_id2 = global_id2 + gridDim.x * blockDim.x, and return to step (4-10);

[0130] (5) Obtain the graph data copied to the GPU memory in step (2), and use the bidirectional BFS algorithm to query whether the query point obtained in step (2) is reachable in the graph data as the query result, and then proceed to step (6);

[0131] The advantage of this step is that graph data with lower average out-degree will generate fewer intermediate results during the expansion process. Therefore, the high concurrency characteristics of the GPU can be used to process nodes in two search directions at the same time, increase the number of intermediate results, make full use of GPU computing and storage resources, and accelerate the query process.

[0132] like Figure 5 As shown, this step specifically includes the following sub-steps:

[0133] (5-1) Create a forward queue fque, a backward queue bque, a forward distance array fdis, and a backward distance array bdis, initialize the forward queue fque as the source vertex, initialize the backward queue bque as the target vertex, initialize the distance corresponding to the source vertex in the forward distance array fdis to 0, and initialize the distances corresponding to the remaining vertices to infinity, initialize the distance corresponding to the target vertex in the backward distance array bdis to 0, and initialize the distances corresponding to the remaining vertices to infinity;

[0134] (5-2) Determine whether the forward distance array fdis and the backward distance array bdis have non-infinite distances at the same position. If so, the source vertex and the target vertex are reachable as the query result, and the process ends. Otherwise, determine whether the forward queue fque and the backward queue bque are both empty. If so, the source vertex and the target vertex are unreachable as the query result, and the process ends. Otherwise, go to step (5-3);

[0135] (5-3) Get the global ID of the current thread;

[0136] Specifically, each thread uses the global ID as the initial index to process the vertices in queue e;

[0137] More specifically, the thread global ID is equal to:

[0138] ID=blockIdx.x*blockDim.x+threadIdx.x

[0139] Among them, blockIdx.x represents the number of the thread block corresponding to the thread with dimension x in the entire thread grid, blockDim.x represents the number of threads contained in the thread block corresponding to the thread with dimension x, and threadIdx.x represents the number of the thread with dimension x in the thread bundle in the thread grid where it is located. The global ID of the thread in the entire thread grid can be calculated by multiplying the first two and adding the third.

[0140] (5-4) Set the counter global_id3 = the global ID of the current thread obtained in step (5-3);

[0141] (5-5) Determine whether global_id3 is less than the length of the forward queue fque. If so, obtain the vertex fque[global_id3] with global_id3 in the forward queue fque, and then proceed to step (5-6). Otherwise, it means that all vertices in the forward queue fque have been processed, and then proceed to step (5-8);

[0142] (5-6) Determine whether the distance of the adjacent vertices in the adjacency list corresponding to the global_id3 vertex fque[global_id3] in the forward queue fque in the forward distance array fdis is infinite. If so, update the distance of the adjacent vertices in the adjacency list corresponding to the global_id3 vertex fque[global_id3] in the forward queue fque in the forward distance array fdis and add the adjacent vertices in the adjacency list corresponding to the global_id3 vertex fque[global_id3] in the forward queue fque to the forward queue fque. Otherwise, it means that the global_id3 vertex fque[global_id3] in the forward queue fque has been processed, remove the vertex from the forward queue fque, and then enter step (5-7);

[0143] (5-7) Set global_id3 = global_id3 + gridDim.x*blockDim.x, and return to step (5-5); where gridDim.x represents the number of thread blocks in the thread grid where the thread of dimension x is located, blockDim.x represents the number of threads contained in the thread block corresponding to the thread of dimension x, and gridDim.x*blockDim.x represents the total number of threads in the thread grid where the thread of dimension x is located (if the elements in the vertex queue are much larger than the total number of threads (gridDim.x*blockDim.x), this grid-striding loop mode can enable each thread to process multiple data elements and make rational use of computing resources).

[0144] (5-8) Set the counter global_id4 = the global ID of the current thread obtained in step (5-3);

[0145] (5-9) Determine whether global_id4 is less than the length of the backward queue bque. If so, obtain the vertex bque[global_id4] with global_id4 in the backward queue bque, and then proceed to step (5-10). Otherwise, it means that all vertices in the backward queue bque have been processed, and then return to step (5-2);

[0146] (5-10) Determine whether the distance between the adjacent vertices in the adjacency list corresponding to the global_id4 vertex bque[global_id4] in the backward queue bque and the corresponding distances in the backward distance array bdis are infinite. If so, update the distances between the adjacent vertices in the adjacency list corresponding to the global_id4 vertex bque[global_id4] in the backward queue bque and add the adjacent vertices in the adjacency list corresponding to the global_id4 vertex bque[global_id4] in the backward queue bque to the backward queue bque. Otherwise, it means that the global_id4 vertex bque[global_id4] in the backward queue bque has been processed. Remove the vertex from the backward queue bque and enter step (5-11).

[0147] (5-11) Set global_id4 = global_id4 + gridDim.x * blockDim.x, and return to step (5-9);

[0148] (6) Copy the query results to the CPU memory and output them.

[0149] In one embodiment, a computer device is provided, which may be a gateway. The computer device includes a processor, a memory, a network interface, and a database connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store the recorded IP addresses and corresponding MAC address data of the terminals in the local area network. The network interface of the computer device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, a reachability query method based on GPU acceleration is implemented.

[0150] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, all steps of the reachability query method based on GPU acceleration of the present invention are implemented.

[0151] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, all steps of the reachability query method based on GPU acceleration of the present invention are implemented.

[0152] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), memory bus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).

[0153] It will be easily understood by those skilled in the art that the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the protection scope of the present invention.< / int>

Claims

1. A reachability query method based on GPU acceleration, characterized in that: The following steps are involved: (1) Receive graph data from a user and perform compression preprocessing on the graph data to obtain compressed graph data. (2) Obtaining a query point from the user, which includes a source vertex and a target vertex, configuring a GPU environment, and copying the graph data compressed in step (1) and the query point to the GPU memory in the configured GPU environment; (3) Calculate the average out-degree of the graph data copied to the GPU memory in step (2), and determine whether the average out-degree is greater than a preset threshold (the value range is 3 to 9, preferably 8). If so, proceed to step (4); otherwise, proceed to step (5); (4) Obtain the graph data copied to the GPU memory in step (2), and use the two-stage BFS algorithm to query whether the query point obtained in step (2) is reachable in the graph data as the query result, and then proceed to step (6); (5) Obtain the graph data copied to the GPU memory in step (2), and use the bidirectional BFS algorithm to query whether the query point obtained in step (2) is reachable in the graph data as the query result, and then proceed to step (6); (6) Copy the query results to the CPU memory and output them.

2. The GPU-accelerated reachability query method according to claim 1, characterized in that: Step (1) specifically includes the following sub-steps: (1-1) Receive graph data from the user, set counter i = 0, obtain the total number of edges and vertices in the graph data, and initialize array C and array R to empty; (1-2) Determine whether i is equal to the total number of vertices in the graph data. If so, use the obtained array C as the final adjacency list and array R as the final offset table. Array C and array R are the compressed graph data. The process ends. Otherwise, go to step (1-3). (1-3) Determine whether the adjacent vertices of the i-th vertex in the graph data exist. If so, add the adjacent vertices to the array C in sequence, and record the position of the first adjacent vertex of the i-th vertex in the graph data in the array C in the array R, and then go to step (1-5), otherwise go to step (1-4); (1-4) Set the i-th element R[i] in the array R to the i-1-th element R[i-1], and then go to step (1-5); (1-5) Set i=i+1 and return to step (1-2).

3. The GPU-accelerated reachability query method according to claim 1 or 2, characterized in that: The average out-degree σ in step (3) is equal to: Where n represents the total number of vertices in the graph data, and out_deg(i) represents the out-degree of the i-th vertex in the graph data.

4. The GPU-accelerated reachability query method according to any one of claims 1 to 3, characterized in that: Step (4) specifically includes the following sub-steps: (4-1) Create a vertex queue vque, an edge queue eque, and a distance array dis, initialize the vertex queue vque to the source vertex, initialize the edge queue eque to empty, initialize the distance corresponding to the source vertex in the distance array dis to 0, and initialize the distances corresponding to the remaining vertices in the distance array dis to infinity; (4-2) Determine whether the vertex queue vque is empty. If so, take the source vertex and the target vertex as unreachable as the query result, and the process ends. Otherwise, determine whether the vertex queue vque contains the target vertex. If so, take the source vertex and the target vertex as reachable as the query result, and the process ends. Otherwise, go to step (4-3); (4-3) Get the global ID of the current thread; (4-4) Set the counter global_id1 = the global ID of the current thread obtained in step (4-3); (4-5) Determine whether global_id1 is less than the length of the vertex queue vque. If so, obtain the vertex vque[global-id1] with global_id1 in the vertex queue vque, and then proceed to step (4-6). Otherwise, it means that all vertices in the vertex queue vque have been processed, and then proceed to step (4-9); (4-6) Obtain the size of the adjacency list corresponding to the vertex vque[global-id1] with global_id1 in the vertex queue vque, and obtain the prefix and result of the thread corresponding to the vertex in the thread grid according to the size of the adjacency list, as the offset of the adjacent vertex in the corresponding adjacency list of the vertex vque[global-id1] with global_id1 in the edge queue eque, and store all the adjacent vertices in the edge queue eque; (4-7) Set global_id1 = global_id1 + gridDim.x * blockDim.x, and return to step (4-5); where gridDim.x represents the number of thread blocks in the thread grid where the thread of dimension x is located, blockDim.x represents the number of threads contained in the thread block corresponding to the thread of dimension x, and gridDim.x * blockDim.x represents the total number of threads in the thread grid where the thread of dimension x is located; (4-8) Copy the edge queue eque obtained in step (4-6) to the shared memory, and use the hash table in the shared memory to perform conflict detection processing on the edge queue eque (i.e., remove duplicate vertices) to obtain the edge queue eque after the conflict detection processing; (4-9) Set the counter global_id2 = the global ID of the current thread obtained in step (4-3); (4-10) Determine whether global_id2 is less than the length of the edge queue eque. If so, obtain the vertex eque[global_id2] with global_id2 in the edge queue eque, and then go to step (4-11). Otherwise, it means that all vertices in the edge queue eque have been processed, and then return to step (4-2); (4-11) Determine whether the distance corresponding to the vertex eque[global_id2] with the global_id2th vertex in the distance array dis is infinite. If so, update the distance corresponding to the vertex eque[global_id2] with the global_id2th vertex in the distance array dis, and add the vertex eque[global_id2] with the global_id2th vertex to the vertex queue vque, otherwise proceed to step (4-12); (4-12) Set global_id2=global_id2+gridDim.x*blockDim.x, and return to step (4-10).

5. The GPU-accelerated reachability query method according to claim 4, characterized in that: Steps (4-6) specifically include the following sub-steps: (4-6-1) Get the size of the adjacency list corresponding to the vertex vque[global-id1] with the global_id1 in the vertex queue vque, set the counter i=0, initialize the array prefix to be empty, which is used to record the prefix and result of the thread bundle corresponding to the thread corresponding to the vertex with the global_id1 in the thread grid, initialize the array nbr (whose size is 32) to be empty, which is used to record the size of the adjacency list corresponding to the vertex with the global_id1, and initialize the 0th element prefix[0] in the array prefix to 0; (4-6-2) Determine whether i is equal to the size of the array prefix - 1. If so, the array prefix is ​​used as the final prefix and result table and the process ends. Otherwise, go to step (4-6-3); (4-6-3) Set the i-th element in the array prefix to prefix[i-1]+nbr[i-1], where nbr[i-1] represents the size of the adjacency list corresponding to the vertex of global_id1, and then go to step (4-6-4); (4-6-4) Set i=i+1 and return to step (4-6-2).

6. The GPU-accelerated reachability query method according to claim 5, characterized in that: The global ID of the thread in step (4-3) is equal to: ID=blockIdx.x*blockDim.x+threadIdx.x Among them, blockIdx.x represents the number of the thread block corresponding to the thread with dimension x in the entire thread grid, blockDim.x represents the number of threads contained in the thread block corresponding to the thread with dimension x, and threadIdx.x represents the number of the thread with dimension x in the thread bundle in the thread grid where it is located.

7. The GPU-accelerated reachability query method according to claim 6, characterized in that: Step (5) specifically includes the following sub-steps: (5-1) Create a forward queue fque, a backward queue bque, a forward distance array fdis, and a backward distance array bdis, initialize the forward queue fque as the source vertex, initialize the backward queue bque as the target vertex, initialize the distance corresponding to the source vertex in the forward distance array fdis to 0, and initialize the distances corresponding to the remaining vertices to infinity, initialize the distance corresponding to the target vertex in the backward distance array bdis to 0, and initialize the distances corresponding to the remaining vertices to infinity; (5-2) Determine whether the forward distance array fdis and the backward distance array bdis have non-infinite distances at the same position. If so, the source vertex and the target vertex are reachable as the query result, and the process ends. Otherwise, determine whether the forward queue fque and the backward queue bque are both empty. If so, the source vertex and the target vertex are unreachable as the query result, and the process ends. Otherwise, go to step (5-3); (5-3) Get the global ID of the current thread; (5-4) Set the counter global_id3 = the global ID of the current thread obtained in step (5-3); (5-5) Determine whether global_id3 is less than the length of the forward queue fque. If so, obtain the vertex fque[global_id3] with global_id3 in the forward queue fque, and then proceed to step (5-6). Otherwise, it means that all vertices in the forward queue fque have been processed, and then proceed to step (5-8); (5-6) Determine whether the distance of the adjacent vertices in the adjacency list corresponding to the global_id3 vertex fque[global_id3] in the forward queue fque in the forward distance array fdis is infinite. If so, update the distance of the adjacent vertices in the adjacency list corresponding to the global_id3 vertex fque[global_id3] in the forward queue fque in the forward distance array fdis and add the adjacent vertices in the adjacency list corresponding to the global_id3 vertex fque[global_id3] in the forward queue fque to the forward queue fque. Otherwise, it means that the global_id3 vertex fque[global_id3] in the forward queue fque has been processed, remove the vertex from the forward queue fque, and then enter step (5-7); (5-7) Set global_id3 = global_id3 + gridDim.x*blockDim.x, and return to step (5-5); where gridDim.x represents the number of thread blocks in the thread grid where the thread of dimension x is located, blockDim.x represents the number of threads contained in the thread block corresponding to the thread of dimension x, and gridDim.x*blockDim.x represents the total number of threads in the thread grid where the thread of dimension x is located (if the elements in the vertex queue are much larger than the total number of threads (gridDim.x*blockDim.x), this grid-striding loop mode can enable each thread to process multiple data elements and make rational use of computing resources). (5-8) Set the counter global_id4 = the global ID of the current thread obtained in step (5-3); (5-9) Determine whether global_id4 is less than the length of the backward queue bque. If so, obtain the vertex bque[global_id4] with global_id4 in the backward queue bque, and then proceed to step (5-10). Otherwise, it means that all vertices in the backward queue bque have been processed, and then return to step (5-2); (5-10) Determine whether the distance between the adjacent vertices in the adjacency list corresponding to the global_id4 vertex bque[global_id4] in the backward queue bque and the corresponding distances in the backward distance array bdis are infinite. If so, update the distances between the adjacent vertices in the adjacency list corresponding to the global_id4 vertex bque[global_id4] in the backward queue bque and add the adjacent vertices in the adjacency list corresponding to the global_id4 vertex bque[global_id4] in the backward queue bque to the backward queue bque. Otherwise, it means that the global_id4 vertex bque[global_id4] in the backward queue bque has been processed. Remove the vertex from the backward queue bque and enter step (5-11). (5-11) Set global_id4=global_id4+gridDim.x*blockDim.x, and return to step (5-9).

8. A GPU-accelerated reachability query system, characterized in that: include: The first module is used to receive graph data from a user and perform compression preprocessing on the graph data to obtain compressed graph data. The second module is used to obtain query points from the user, including source vertices and target vertices, configure the GPU environment, and copy the graph data compressed and processed by the first module and the query points to the GPU memory in the configured GPU environment; The third module is used to calculate the average out-degree of the graph data copied to the GPU memory by the second module, and determine whether the average out-degree is greater than a preset threshold (the value range is 3 to 9, preferably 8). If yes, enter the fourth module, otherwise enter the fifth module; The fourth module is used to obtain the graph data copied by the second module to the memory of the GPU, and use the two-stage BFS algorithm to query whether the query point obtained by the second module is reachable in the graph data as the query result, and then the sixth module; The fifth module is used to obtain the graph data copied by the second module to the memory of the GPU, and use the bidirectional BFS algorithm to query whether the query point obtained by the second module is reachable in the graph data as the query result, and then transfer to the sixth module; The sixth module is used to copy the query results to the CPU memory and output them.

Citation Information

Patent Citations

  • Semi-decompression data compression method for optimizing streaming data processing performance

    CN113992208A

  • Efficient parallel graph processing method and system based on direct calculation of compressed data

    CN115482147A

  • Compression index and query method of attribute graph

    CN115495617A

  • Dynamic and compressed trie for use in route lookup

    US20180212876A1

  • Graph database construction method and apparatus based on graph compression, and related component

    WO2022241813A1

Cited By

  • GPU (Graphics Processing Unit) instruction copying and spreading method and equipment, storage medium and program product

    CN121209967A

  • Copy propagation method, device, storage medium and program product of GPU instruction

    CN121209967B