Relationship connection method realized based on GPU (Graphics Processing Unit)
By designing efficient data organization, chunking and permutation methods, as well as task scheduling mechanisms in the database field, the problem that large-scale relational data sets cannot be fully loaded into GPU video memory is solved, and efficient relational connection operations are achieved.
Patent Information
- Application Number
- CN202510212805.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-05-30
AI Technical Summary
In traditional methods, large-scale relational data sets cannot be fully loaded into GPU video memory, which limits the execution efficiency of relational connection operations.
A GPU-based relational connection method is designed. Through efficient data organization, chunking and permutation methods, and task scheduling mechanism, the efficient exchange and processing of relational data between disk and memory, memory and GPU is realized.
It improves the accuracy and efficiency of GPU to implement relational connection operations and can effectively process large-scale relational data sets.
Smart Images

Figure CN120066727A_ABST
Abstract
Description
Technical Field
[0001] The implementation of the present invention relates to the field of databases. A relational join method implemented based on GPU is studied, which involves efficient data organization, permutation, and task scheduling methods to optimize the processing efficiency of large-scale relational datasets. The invention is applicable to the analysis and understanding of relational join operations in the database field. Background Art
[0002] In recent years, the rapid increase in the amount of data in various fields has caused the database scale to expand rapidly, which has brought major challenges to the efficiency of database query processing. As a parallel computing processor, GPU has a large number of computing cores and can process multiple data streams simultaneously, providing a feasible computing framework for improving the query efficiency of databases. Relational join operation is a relatively complex operation in database query applications. Improving its operation efficiency is extremely important for improving the query efficiency of database applications. Therefore, using GPU to improve the query performance of databases has become an active research topic in the database field. So far, the research on implementing relational join operations based on GPU has focused on the implementation mechanism of join operations on GPU. The data scale it targets can be loaded into the video memory at one time, so it has great limitations in applications. Some products have proposed the implementation mechanism of data query applications under the hybrid architecture of CPU and GPU. In these mechanisms, GPU is generally used as a cache supplement for CPU, mainly used to expand the data scale of operations, save time from data acquisition, and improve efficiency. For the join operation of large-scale data, no implementation method under the hybrid architecture of CPU and GPU has been found so far. The present invention studies the implementation mechanism of join operations based on GPU for large-scale relational tables and designs a relational join method implemented based on GPU. Summary of the Invention
[0003] The present invention proposes a relational join method implemented based on GPU, designs an efficient data organization and permutation method, and a task scheduling method, solves the problem that large-scale relational datasets in traditional methods cannot be completely loaded into the GPU video memory, thus restricting the execution of relational join operations, and improves the accuracy and efficiency of implementing relational join operations on GPU.
[0004] The present invention proposes a relational join method implemented based on GPU. The related technical solutions of this method include:
[0005] Design an organization method for relational data between disk and memory, and between memory and GPU;
[0006] Perform block processing on the relationship according to the GPU video memory capacity, and read the data blocks of the relationship into the memory;
[0007] Design a data permutation method to realize the exchange of relational data between disk and memory, and between memory and GPU;
[0008] A task scheduling mechanism is designed to coordinate the computing tasks of the CPU and GPU;
[0009] Based on the data organization, chunking and permutation methods and the task scheduling method, a complete relational join operation process is designed to achieve the join of large-scale relations.
[0010] Furthermore, the process of designing the organization method of relational data between disk and memory is as follows:
[0011] The complete data transfer is adopted between the disk and the memory for the relation, and all the attribute columns of the relation are completely transferred to the memory and organized into a one-dimensional array form; the purpose of adopting the above organization method in the present invention is that when the CPU assembles the result set, the complete attribute columns of the relation can be obtained from the memory, thereby reducing the data exchange frequency between the disk and the memory.
[0012] Furthermore, the process of designing the organization method of relational data between memory and GPU is as follows:
[0013] (1) Data transfer from memory to GPU video memory: Extract the tuple numbers and the participating attribute columns of the relation, organize them into a one-dimensional array form, i.e., <tuple number, attribute column>, and then transfer them to the GPU; adopting the above organization method helps to improve the continuity of memory access, and the threads in the GPU can access adjacent memory addresses in a coalesced manner, reducing the latency caused by random access;
[0014] (2) Data transfer from GPU video memory to memory: The GPU adopts the sorted merge join method to generate the join result set, and the join result set is organized in the form of <tuple number of relation R, tuple number of relation S>. When the GPU finishes processing the tuples, the join result set is transferred to the memory.
[0015] Furthermore, the process of chunking the relation according to the GPU video memory capacity is as follows:
[0016] Relations R and S are divided into B R 、B S blocks, and B R 、B S are calculated by the following formula:
[0017]
[0018] B R = N R / T
[0019] B S = N S / T
[0020] Where: a is the spatial expansion coefficient, and its value range is 0.5 < a < 1; M GPU is the video memory capacity of the GPU; C R , C S are the spaces occupied by the attributes and tuple numbers participating in the join in relations R and S respectively; T is the average number of tuples passed into the GPU; N R , N S are the total numbers of tuples in R and S respectively.
[0021] Furthermore, a data replacement method is designed, and the process of realizing the exchange of relational data between disk and memory is as follows:
[0022] Read all data blocks of relation R and the first k data blocks S j of S into memory for calculation, where 1 ≤ j ≤ k; k is calculated by the following formula:
[0023]
[0024] Where: M CPU represents the capacity of the memory; M R represents the capacity of relation R, represents rounding down;
[0025] During the relational join process, if the data block S j has completed the join, then transfer the unjoined data block S x in the disk to the host memory and overwrite the space occupied by the data block S j ; The present invention only needs to read data from the disk once and write data to the disk once, reducing the number of blocks for long-term data transfer.
[0026] Furthermore, a data replacement method is designed, and the process of realizing the exchange of relational data between disk and memory is as follows:
[0027] The join of two relations is realized by using an ordered loop method: number the data blocks of R and S read into memory in order, and then the S data blocks are sequentially joined with all blocks of R in the positive or reverse order of the numbers, that is, first use the first block of S to join with all blocks of R in the order of R 1 , R 2 ,..., R BR for the join operation. After completion, then use the second block of S to join with all blocks of R in the order of R BR , R BR -1,..., R 1 for the join operation. Then, the third block of S is joined with all blocks of R in the order of R 1 , R 2 ,..., R BRPerform the join operation in sequence, and so on until all blocks of S in memory are completed with the join operation, then start the next round of loop; after each block of S completes the join operation with R, it is deleted from memory and a new block of S is read from disk to fill its position until the last data block of S is read into memory and the join operation is completed;
[0028] The sequential loop join method enables each block of S to exchange data with the GPU only once when performing the join operation with R: upload once and write back once, that is, transfer to the GPU when first joining with the R block, write back the sorted result after joining, and then remain in the GPU to join with other blocks of R; for the R data block, when the first block of S finishes the join operation with all data blocks of R, the last block of R remains in the GPU, and then R joins with the second block of S in the reverse order, so that one data exchange can be reduced during the loop process;
[0029] Based on the above method, the number of exchanged blocks L between memory and GPU is calculated by the following formula, including the number of blocks from memory to GPU and from GPU to memory:
[0030] L = B R × (B S + 1) + 1
[0031] The number of exchanged blocks Q between disk and memory is calculated by the following formula, including the number of blocks read in and written out:
[0032] Q = 2(B R + B S )
[0033] Furthermore, the process of designing a task scheduling mechanism to coordinate the CPU and GPU computing tasks is as follows:
[0034] The GPU mainly executes the sorting within the block and the join task of the two relations, and uses the sort-merge join method to implement this task, and the sorting uses bitonic sort;
[0035] The CPU has two tasks to execute: assembling the join result set and performing the join operation based on the sorted data blocks; these two tasks can be executed simultaneously with the GPU, but if the CPU faces two tasks at the same time, the task execution needs to be performed according to the following priority: result set assembly task > join task of sorted data blocks;
[0036] At the beginning of the operation, the CPU first reads the data blocks of the two relations participating in the join operation from the disk and organizes them to be transferred to the GPU. Then the GPU starts to perform the calculation. At this time, the CPU is in an idle state. After the GPU finishes the calculation and transfers the sorted data blocks and the join result to the memory, the CPU transfers data blocks to the GPU again. At this time, when the GPU performs the operation again, the CPU can assemble tuples based on the join result set output by the GPU to generate the result of the join operation. If the GPU has not finished executing after the CPU assembles the result, the CPU can participate in the tuple matching work. If the GPU has finished the operation and is about to output the result when the CPU performs the tuple matching operation, the CPU should immediately stop the tuple matching operation and perform the data transfer to the GPU. After the data transfer is completed, continue to complete the current matching operation. To implement this mechanism, record the current matching positions posR and posS of the CPU executing the tuple blocks R m and S n so that when the CPU executes the tuple block join task, it will immediately start the tuple matching operation from the positions of posR and posS;
[0037] To ensure that the CPU and GPU execute tasks in correct logic, the present invention adopts a CUDA event-driven mechanism to achieve dynamic scheduling and control. The specific process is as follows: First, the CPU transfers the unjoined data blocks in the memory to the GPU and starts the GPU to execute the join operation task. At the same time, it calls cudaEventCreate to create an event flag and records the event status through cudaEventRecord. During the execution of the GPU task, the CPU queries the event status in real time through cudaEventQuery: If the status is cudaSuccess, it means that the GPU task has been completed, and the CPU immediately stops the current task and jumps to the subsequent steps; if the status is cudaErrorNotReady, it means that the GPU is still running, and the CPU continues to execute the result set assembly task; if the GPU is still not ready after the result set assembly task is completed, the CPU performs the tuple matching task based on the sorted data blocks and continues the operation from the recorded positions of posR and posS until the GPU task is completed; if there are unjoined data blocks in the memory, the CPU restarts the GPU task; otherwise, continue to execute the remaining tasks and then terminate the process. The present invention uses CUDA events to sense the GPU status in real time, avoids the CPU waiting blindly, and dynamically allocates CPU tasks according to the priority to maximize the resource utilization rate.
[0038] Furthermore, based on the data organization, chunking and permutation method and the task scheduling method, a complete relational join operation process is designed. The process of realizing large-scale relational join is as follows:
[0039] (1) Organize the relations R and S in the disk according to the data organization method, partition R and S according to the video memory capacity of the GPU, and then read the first k blocks of R and S into the memory;
[0040] (2) According to the data replacement method, transfer the data of R and S in the memory to the GPU according to the attribute values for connection in the blocks. If the data blocks have been sorted, the GPU only performs the matching operation and transfers the matching results to the memory; if the data blocks have not been sorted, the GPU performs the sorting and matching operations and transfers the sorting results and matching results to the memory;
[0041] (3) While the GPU is performing operations, after the CPU receives the results returned by the GPU, on the one hand, it assembles the tuples of the matching connection attributes and the corresponding data blocks and outputs them, and on the other hand, it can also implement the tuple matching operation based on the data blocks sorted by the GPU;
[0042] (4) While the GPU is performing operations, if a certain block of S has completed the matching operation, delete the data of this block in the memory, read in a new block of S, and continue to repeat the data replacement method and the task scheduling method until all blocks of S have been calculated.
[0043] The creativity of the present invention is mainly reflected in:
[0044] (1) By designing the data organization and replacement methods, the present invention efficiently manages the transmission of relational data between the memory and the GPU video memory, reducing the impact of data transmission on the implementation of relational joins by the GPU.
[0045] (2) By designing the task scheduling method, the present invention realizes the collaborative computing between the CPU and the GPU in the relational join operation, maximizing the computing power of heterogeneous devices.
[0046] (3) Based on the data organization and replacement methods and the task scheduling method, the present invention designs a complete relational join operation process, improving the efficiency of performing join operations on large-scale relational tables on the GPU. Brief Description of the Drawings
[0047] Figure 1 It is a schematic diagram of the organization of relational data in the memory and the GPU.
[0048] Figure 2 It is a schematic diagram of the replacement of tuple blocks between the disk and the memory.
[0049] Figure 3 It is a schematic diagram of the sequential conversion of relational joins.
[0050] Figure 4 It is a schematic diagram of the task scheduling between the CPU and the GPU. Detailed Embodiments
[0051] The following is a detailed description of a relationship connection method based on GPU implementation of the present invention in conjunction with the accompanying drawings:
[0052] S1 designed an organization method for relational data between disk and memory, and between memory and GPU;
[0053] S2 performs block processing on the relationship according to the GPU video memory capacity, and reads the data blocks of the relationship into memory;
[0054] S3 designed a data replacement method to realize the exchange of relational data between disk and memory, and between memory and GPU;
[0055] S4 designed a task scheduling mechanism to coordinate the computing tasks of CPU and GPU;
[0056] S5 designed a complete relational join operation process based on data organization, chunking and replacement methods, and task scheduling methods, and realized the join of large-scale relationships.
[0057] In S1, the process of designing the organization method for relational data between disk and memory is as follows:
[0058] The relationship uses complete data transfer between disk and memory, transfers all attribute columns of the relationship to memory in a complete manner, and organizes them into a one-dimensional array form; the present invention adopts the above organization method with the aim of enabling the CPU to obtain the complete attribute columns of the relationship from memory when assembling the result set, thereby reducing the data exchange frequency between disk and memory.
[0059] In S1, as Figure 1 shown, the process of designing the organization method for relational data between memory and GPU is as follows:
[0060] (1) Data transfer from memory to GPU video memory: Extract the tuple numbers of the relationship and the attribute columns participating in the join, organize them into a one-dimensional array form, i.e., <tuple number, attribute column>, and then transfer them to the GPU; adopting the above organization method helps to improve the continuity of storage access, and the threads in the GPU can access adjacent storage addresses in a merged manner, reducing the latency caused by random access;
[0061] (2) Data transfer from GPU video memory to memory: The GPU uses the sorted merge join method to generate the join result set, and the join result set is organized in the form of <tuple number of relationship R, tuple number of relationship S>. When the GPU finishes processing the tuples, it transfers the join result set to memory.
[0062] In S2, the process of performing block processing on the relationship according to the GPU video memory capacity is as follows:
[0063] Divide relationships R and S into B R 、BS Block, B R , B S is calculated by the following formula:
[0064]
[0065] B R = N R / T
[0066] B S = N S / T
[0067] Where: a is the spatial expansion coefficient, and the value range is 0.5 < a < 1; M GPU is the video memory capacity of the GPU; C R , C S are the spaces occupied by the attributes and tuple numbers participating in the join in relations R and S respectively; T is the average number of tuples passed into the GPU; N R , N S are the total numbers of tuples of R and S respectively.
[0068] In S3, as Figure 2 shown, a data replacement method is designed, and the process of realizing the exchange of data blocks between disk and memory is as follows:
[0069] Read all data blocks of relation R and the first k data blocks S j of S into memory for calculation, where 1 ≤ j ≤ k; k is calculated by the following formula:
[0070]
[0071] Where: M CPU represents the capacity of the memory; M R represents the capacity of relation R, represents rounding down;
[0072] During the relation join process, if the data block S j has completed the join, then transfer the unjoined data block S x in the disk to the host memory and overwrite the space occupied by the data block S j ; The present invention only needs to read data from the disk once and write data to the disk once, reducing the number of long-term data transfer blocks;
[0073] Furthermore, a data replacement method is designed, and the process of realizing the exchange of data blocks between memory and GPU is as follows:
[0074] As Figure 3As shown in the figure, the connection of two relations is implemented in an ordered loop manner: the data blocks of R and S read into the memory are numbered in order, and then the S data blocks are sequentially connected to all the blocks of R in the positive or reverse order of the sequence numbers, that is, first use the first block of S to connect to all the blocks of R in the order of R 1 ,R 2 ,...,R BR 's order for the connection operation. After completion, then use the second block of S to connect to all the blocks of R in the order of R BR ,R BR -1,…,R 1 's order for the connection operation. Then, the third block of S is connected to all the blocks of R in the order of R 1 ,R 2 ,…,R BR 's order for the connection operation, and so on until all the blocks of S in the memory complete the connection operation, and then start the next round of loop; after each block of S completes the connection operation with R, it is deleted from the memory, and a new S block is read from the disk to fill its position until the last data block of S is read into the memory from the disk and the connection operation is completed;
[0075] The ordered loop connection method enables each S block to have only one data exchange with the GPU when connecting to R: uploading once and writing back once, that is, it is transferred to the GPU when connecting to the R block for the first time, and the sorted result is written back after the connection, and then retained in the GPU to connect with other blocks of R; for the R data block, when the first block of data of S finishes the connection operation with all the data blocks of R, the last block of data of R is retained in the GPU, and then R connects with the second block of data of S in the reverse order, so that one data exchange can be reduced during the loop process;
[0076] Based on the above method, the number of exchanged blocks L of data between the memory and the GPU is calculated by the following formula, including the number of blocks from the memory to the GPU and from the GPU to the memory:
[0077] L = B R ×(B S +1)+1
[0078] The number of exchanged blocks Q of data between the disk and the memory is calculated by the following formula, including the number of blocks read in and written out:
[0079] Q = 2(B R +B S )
[0080] In S4, as Figure 4 shown, the process of designing a task scheduling mechanism to coordinate the CPU and GPU computing tasks is as follows:
[0081] The GPU mainly performs the tasks of sorting data within a block and joining two relations. The sorting-merge join method is used to implement these tasks, and bitonic sorting is adopted for sorting.
[0082] The CPU has two tasks to execute: assembling the joined result set and performing the join operation based on the sorted data blocks. These two tasks can be executed simultaneously with the GPU. However, if the CPU faces both tasks at the same time, the task execution needs to follow the following priority: the result set assembly task > the join task on the sorted data blocks.
[0083] At the beginning of the operation, the CPU first reads the data blocks of the two relations involved in the join operation from the disk and organizes them to be transmitted to the GPU. Then the GPU starts to calculate, and at this time the CPU is in an idle state. After the GPU finishes the calculation and transmits the sorted data blocks and the join result to the memory, the CPU transmits data blocks to the GPU again. At this time, when the GPU performs the operation again, the CPU can assemble tuples based on the join result set output by the GPU to generate the result of the join operation. If the GPU has not finished the operation after the CPU assembles the result, the CPU can participate in the tuple matching work. If the GPU has finished the operation and is about to output the result when the CPU is performing the tuple matching operation, the CPU should immediately stop the tuple matching operation, perform the data transmission to the GPU, and then continue to complete the current matching operation after the data transmission is completed. To implement this mechanism, record the current matching positions posR and posS of the CPU when executing the tuple blocks R m and S n so that when the CPU executes the tuple block join task, it will immediately start the tuple matching operation from the positions of posR and posS.
[0084] To ensure that the CPU and GPU execute tasks in correct logic, the present invention adopts a CUDA event-driven mechanism to achieve dynamic scheduling and control. The specific process is as follows: First, the CPU transfers the unconnected data blocks in the memory to the GPU and starts the GPU to execute the connection operation task. At the same time, it calls cudaEventCreate to create an event flag and records the event status through cudaEventRecord. During the execution of the GPU task, the CPU queries the event status in real time through cudaEventQuery: If the status is cudaSuccess, it means that the GPU task has been completed, and the CPU immediately stops the current task and jumps to the subsequent steps; if the status is cudaErrorNotReady, it means that the GPU is still running, and the CPU continues to execute the result set assembly task; if the GPU is still not ready after the result set assembly task is completed, the CPU performs a tuple matching task based on the sorted data blocks and continues the operation from the recorded posR and posS positions until the GPU task is completed; if there are unconnected data blocks in the memory, the CPU restarts the GPU task; otherwise, it continues to execute the remaining tasks and then terminates the process. The present invention uses CUDA events to perceive the GPU status in real time, avoids the CPU waiting blindly, and dynamically allocates CPU tasks according to the priority to maximize the resource utilization rate.
[0085] In S5, based on the data organization, chunking and permutation method and the task scheduling method, a complete relational join operation process is designed. The process of realizing large-scale relational join is as follows:
[0086] (1) Organize the relations R and S in the disk according to the data organization method, chunk R and S according to the video memory capacity of the GPU, and then read the first k chunks of R and S into the memory.
[0087] (2) According to the data permutation method, transfer the data of R and S in the memory to the GPU according to the attribute values for connection in the chunks. If the data chunks have been sorted, the GPU only performs the matching operation and generates the matching results and transfers them to the memory; if the data chunks have not been sorted, the GPU performs the sorting and matching operations and generates the sorting results and matching results and transfers them to the memory.
[0088] (3) While the GPU is performing the operation, after the CPU receives the results returned by the GPU, on the one hand, it assembles the tuples of the matching connection attributes and the corresponding data chunks and outputs them, and on the other hand, it can also perform the tuple matching operation based on the data chunks sorted by the GPU.
[0089] (4) While the GPU is performing the operation, if a certain chunk of S has completed the matching operation, delete the data of this chunk in the memory, read a new chunk of S, and continue to repeat the data permutation method and the task scheduling method until all chunks of S have been calculated.
Claims
1. A GPU-based relationship connection method, characterized in that: The steps include: S1 designs the organization method of relational data between disk and memory, memory and GPU; S2 processes the relationship in blocks according to the GPU memory capacity and reads the data blocks of the relationship into the memory; S3 has designed a data replacement method to realize the exchange of relational data between disk and memory, and between memory and GPU; S4 designed a task scheduling mechanism to coordinate the computing tasks of the CPU and GPU; Based on data organization, block and replacement methods and task scheduling methods, S5 has designed a complete set of relational connection operation processes to achieve the connection of large-scale relations.
2. The GPU-based relationship connection method according to claim 1, characterized in that: The process of designing the organization of relational data between disk and memory is as follows: Relationships use full data transfer between disk and memory, transferring all attribute columns of the relationship to memory and organizing them into a one-dimensional array.
3. The GPU-based relationship connection method according to claim 1, characterized in that: The process of designing the method of organizing relational data between memory and GPU is as follows: (1) Data transfer from memory to GPU memory: Extract the tuple number of the relation and the attribute columns involved in the connection, organize them into a one-dimensional array, i.e., <tuple number, attribute column>, and then transfer it to the GPU; (2) Data transfer from GPU video memory to main memory: GPU uses sort-merge join method to generate the join result set. The join result set is organized in the form of <tuple number of relation R, tuple number of relation S>. After GPU processes the tuple, it transfers the join result set to main memory.
4. The GPU-based relationship connection method according to claim 1, characterized in that: The process of dividing the relationship into blocks according to the GPU memory capacity is as follows: Partition relations R and S into B R , B S Block, B R , B S Calculated by the following formula: B R =N R / T B S =N S / T Where: a is the spatial expansion coefficient, which is 0.5 <a<1;M GPU is the memory capacity of GPU; C R , C S are the spaces occupied by the attributes and tuple numbers involved in the connection in relations R and S respectively; T is the average number of tuples passed into the GPU; N R 、N S are the total number of tuples of R and S respectively.
5. The method for connecting relations based on GPU according to claim 4, characterized in that: A data replacement method is designed to realize the process of exchanging relational data between disk and memory: Take all the data blocks of relation R and the first k data blocks of S j Read into memory for calculation, where 1≤j≤k; k is calculated by the following formula: Where: M CPU Indicates the capacity of memory; M R represents the capacity of relation R; Indicates rounding down; In the process of relation connection, if the data block S j The connection is completed, then the unconnected data block S in the disk x Transfer to host memory and overwrite data block S j The space occupied.
6. The method for connecting relations based on GPU according to claim 4, characterized in that: A data replacement method is designed to realize the process of exchanging relational data between disk and memory: The connection of two relations is realized by sequential looping: the data blocks of R and S read into the memory are numbered in sequence, and then the S data blocks are connected with all the blocks of R in the order of the sequence numbers in the positive or reverse order of the sequence numbers. That is, the first block of S is connected with all the blocks of R in the order of the sequence numbers. After the connection operation is completed, the second block of S is connected with all blocks of R in the order of Then the third block of S and all blocks of R are connected in the order of The connection operation is performed in the order of S, and so on until all blocks of S in the memory have completed the connection operation, and then the next cycle begins; after each block of S is connected with R, it is deleted from the memory, and a new block of S is read from the disk to fill its place, until the last data block of S is read from the disk into the memory and the connection operation is completed.
7. The method for connecting relations based on GPU according to claim 4, characterized in that: The task scheduling mechanism is designed to coordinate the CPU and GPU computing tasks as follows: The GPU performs the task of sorting data within the block and connecting two relations. The sort merge join method is used to implement this task. The sorting adopts bitonic sorting. The CPU has two tasks to perform: assembling the connection result set and performing connection operations based on the sorted data blocks. These two tasks can be executed simultaneously with the GPU. However, if the CPU faces two tasks at the same time, the task execution must be performed according to the following priority: result set assembly task > sorted data block connection task. The CUDA event-driven mechanism is used to realize the dynamic scheduling and control of CPU and GPU tasks. The specific process is as follows: First, the CPU transfers the unconnected data blocks in the memory to the GPU, and starts the GPU to perform the connection operation task. At the same time, it calls cudaEventCreate to create an event marker and records the event status through cudaEventRecord. During the GPU execution of the task, the CPU queries the event status in real time through cudaEventQuery: if the status is cudaSuccess, it means that the GPU task has been completed, and the CPU immediately stops the current task and jumps to the subsequent steps; if the status is cudaErrorNotReady, it means that the GPU is still running, and the CPU continues to execute the result set assembly task; if the GPU is still not ready after the result set assembly task is completed, the CPU performs the tuple matching task based on the sorted data blocks, and continues the operation from the recorded posR and posS positions until the GPU task is completed; if there are unconnected data blocks in the memory, the CPU restarts the GPU task; otherwise, it continues to execute the remaining tasks and terminates the process.
8. The GPU-based relationship connection method according to claim 1, characterized in that: Based on data organization, block and replacement methods and task scheduling methods, a complete set of relational connection operation processes is designed to realize the process of large-scale relational connection, including: (1) Organize the relations R and S in the disk according to the data organization method, divide R and S into blocks according to the GPU memory capacity, and then read the first k blocks of R and S into the memory; (2) Based on the data replacement method, the data of R and S in the memory are transferred to the GPU according to the attribute values used for connection in the block. If the data block has been sorted, the GPU only performs the matching operation and generates the matching result and transfers it to the memory; if the data block has not been sorted, the GPU performs the sorting and matching operations and generates the sorting result and the matching result and transfers them to the memory; (3) While the GPU is executing the calculation, the CPU receives the result sent back by the GPU and, on the one hand, assembles and outputs the matched connection attributes and the corresponding data blocks into tuples. On the other hand, it implements the tuple matching operation based on the data blocks sorted by the GPU. (4) While the GPU is executing the calculation, if a block of S has completed the matching calculation, the block data is deleted from the memory and a new block of S is read in. The data replacement method and task scheduling method are repeatedly executed until all blocks of S are calculated.