OLAP query processing method for GPU platform
By adopting the OLAP query processing method of thread block vectorization and thread intra-line processing on the GPU platform, the problems of materialization of intermediate results and communication overhead of kernel functions in vectorization processing mode are solved, and the GPU computing efficiency and resource utilization are improved.
Patent Information
- Application Number
- CN202510034766.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-09
- Publication Date
- 2025-05-16
AI Technical Summary
In the vectorization processing mode, the OLAP query processing on the GPU side leads to an increase in materialization overhead of intermediate results, and the load of a single thread is too heavy, which affects GPU performance. The query processing method of one-time operators faces the problems of materialization of intermediate results and communication overhead between different operator kernel functions.
The method of thread block vectorization and thread intra-line processing is adopted. The kernel kernel function in the first stage realizes the selection, projection and grouping of dimension tables. The kernel kernel function in the second stage slices the virtual fact tables and batches are processed by each thread block. Each thread adopts row processing method to realize the filtering, projection, grouping and aggregation calculation of the fact tables.
It improves the computing efficiency of the GPU, solves the problem of excessive load of a single thread caused by vectorization processing, reduces the materialization of intermediate results and kernel function communication overhead, and records intermediate processing results through faster register cache lines, eliminating memory materialization access overhead.
Smart Images

Figure CN120011405A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data analysis and processing, and in particular to an OLAP query processing method for a GPU platform. Background Art
[0002] OLAP (On-Line Analytical Processing) technology provides decision support for enterprises through analysis and calculation on massive data and multi-dimensional analysis results. It is the most important application technology in data warehouse systems. OLAP technology can be implemented on CPU platforms or GPU platforms.
[0003] Technically, OLAP query processing on the GPU has evolved from column processing to vectorized processing.
[0004] In the column processing mode, GPU programming adopts the kernel mode, with the kernel function as a computing carrier, and an operator is often defined as a kernel. Kernel-based query processing aims to optimize the resource utilization of a single kernel and execute the GPU kernels involved in the query plan one by one. In this optimization mode, the kernels are logically and data independent from each other. When a single kernel operator is optimized, the overall query collaborative processing performance is optimized.
[0005] The query processing engine based on kernel functions will lead to serious underutilization of GPU resources, generate cross-core communication delays, and require the materialization of intermediate results to global memory. This query processing model leads to high GPU global memory occupancy, increased global memory access times and delays, and ultimately reduces GPU computing performance. Therefore, vectorized processing has emerged. The key feature of GPU vectorized query processing is that each thread in a GPU thread block processes 4-8 rows of data. It is a thread block vectorized processing and thread vectorized processing model. This model increases GPU resource utilization and reduces global memory access. However, in the vectorized processing mode, it is inevitable to cause the problem of increased overhead of intermediate result materialization. The amount of resources required for each thread is increasing, and the complexity of thread computing is also relatively high. Although GPU threads have strong concurrency capabilities and large numbers, they are not good at processing complex logical operations. Summary of the invention
[0006] The present invention provides an OLAP query processing method for a GPU platform, which is a thread block vectorization and intra-thread "row processing" method; thread block vectorization processing improves the computing efficiency of the GPU, and intra-thread row processing solves the problem of a single thread overload caused by vectorization affecting GPU performance and the problem of a large number of intermediate result materialization and communication overhead between different operator kernel functions faced by a one-time operator query processing method, and eliminates memory materialization access overhead by recording intermediate processing results through faster register cache lines.
[0007] To achieve the above object, the present invention adopts the following technical solutions:
[0008] In a first aspect, the present application provides an OLAP query processing method for a GPU platform, the method comprising:
[0009] Step 1, in response to the SQL statement instruction of the OLAP query processing, the dimension table and the fact table in the database are loaded into the global memory of the GPU platform;
[0010] Step 2, parsing the SQL statement instructions and constructing a virtual fact table;
[0011] Step 3: Implement the selection, projection and grouping operations of the dimension table related to the SQL statement instruction through the kernel function of the first stage;
[0012] Step 4: Based on the results of dimension table selection, projection, and grouping operations, the virtual fact table is partitioned according to the thread block structure of the GPU platform through the kernel function of the second stage, and then batch processed by each thread block. Each thread uses row processing to implement filtering, projection, grouping, and aggregation calculations of the fact table.
[0013] In one implementation, in step 1, the data storage model in the global memory includes:
[0014] Map the dimension table primary key to a continuous surrogate key and update the corresponding fact table foreign key. Map the dimension table column to be accessed to a virtual fact table column based on the fact table foreign key offset address through the fact table foreign key.
[0015] In one implementation, in step 2, query-related dimension table columns are mapped to virtual fact table columns based on array nested access through fact table foreign keys, and query-corresponding multi-table record connection access is mapped to virtual fact table wide table record access.
[0016] In one implementation, in step 3, the dimension table-side operator pushdown is implemented through the table-level kernel function, the selection, projection, and grouping operations on the dimension table are pushed down to the dimension table preprocessing kernel function, and a bitmap or vector is generated as a table-level calculation proxy vector. The virtual fact table column accesses a single dimension table proxy vector through foreign key mapping, reducing the access to multiple columns on the dimension table to one column.
[0017] In one implementation, it specifically includes:
[0018] For each dimension table, the group value corresponding to the GROUP BY clause is projected according to its WHERE condition, and a dynamic dictionary table is created for data compression. A continuously growing unique ID is assigned to each group, and the group ID is stored in the corresponding unit of the proxy vector with the same length as the dimension table. Records that do not meet the WHERE condition are stored as null values in the corresponding proxy vector unit. When only the WHERE clause exists on the dimension table, the proxy vector is simplified to a proxy bitmap, with 1 and 0 representing records that meet and do not meet the WHERE condition respectively.
[0019] In one implementation, in S4, the operation mode of filtering, projecting, grouping and aggregation calculation of the virtual fact table is determined according to the dimension table proxy vector.
[0020] In one implementation, it specifically includes:
[0021] When the dimension table proxy vector corresponding to the virtual fact table column is a bitmap, only the record filtering operation is performed. When the corresponding dimension table proxy vector is a vector, the null value is used to filter the fact table records, and the non-null value is mapped to the multi-dimensional address of the grouping cube composed of the dictionary table encoding dynamically generated by each dimension table, and the fact table metric expression result is mapped to the grouping cube Corresponding unit Perform aggregate calculations.
[0022] In one implementation, during aggregation calculation, each thread block uses the grouping cube stored in the shared memory to perform aggregation calculation between threads, and finally merges the aggregation calculation results of each grouping cube in the global memory as a query grouping aggregation cube, and maps it to the original grouping attributes through a dynamic dictionary table, and outputs a grouping aggregation result set.
[0023] The technical solution of the present invention and the proposed method have the following advantages:
[0024] (1) By converting complex relational operations into efficient array access and calculation, the complexity of the query implementation algorithm code and data structure is reduced, and an optimization model based on kernel fusion is implemented. The core query functions are all run in a kernel function, solving the problems of intermediate result materialization and kernel function communication overhead.
[0025] (2) Through the dimension table preprocessing kernel function, the query clause is pushed down to the dimension table processing to generate a streamlined proxy bitmap or proxy vector. The multi-column access corresponding to the selection, projection, and grouping operations on the dimension table is mapped to a single bitmap or vector data structure, reducing the number of accesses to the virtual fact table columns. The size of the proxy vector is reduced through dynamic dictionary table compression, the size of the grouping cube is minimized, and the high-performance shared memory and cache are fully utilized to cache the dimension table proxy bitmap, proxy vector, and grouping cube, thereby reducing the latency of accessing the global memory.
[0026] (3) A query processing method for strip-based calculations based on the characteristics of GPU hardware architecture is proposed. The complex pattern is converted into a virtual fact table access based on array nested address access. The number of records in a batch processing data strip matches the number of parallel execution threads in the thread block. The intermediate calculation results of each thread processing record are cached through the register file. On the one hand, GPU core switching has almost zero overhead, and the number of thread blocks increases as the amount of thread processing data decreases, which improves the utilization of GPU computing resources. On the other hand, GPU threads are difficult to adapt to complex data processing logic and reduce computing efficiency. The method of each thread processing only one row of data is more suitable for the characteristics of GPU ultra-large-scale simple computing core thread processing, and can give full play to the parallel computing performance of GPU. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] Figure 1 is a schematic diagram of a query execution plan in an embodiment of the present invention;
[0028] Figure 2 It is a schematic diagram of GPU table storage surrogate key conversion and virtual fact table wide table access in one embodiment of the present invention;
[0029] Figure 3 Schematic diagram of wide table access of a virtual fact table based on a dimension table preprocessing mechanism in one embodiment of the present invention;
[0030] Figure 4 It is a schematic diagram of a query processing process based on stripe calculation in one embodiment of the present invention. DETAILED DESCRIPTION
[0031] In order to make the purpose, technical solution and advantages of the embodiment of the present invention clearer, the technical solution of the embodiment of the present invention will be clearly and completely described below in conjunction with the drawings of the embodiment of the present invention. Obviously, the described embodiment is a part of the embodiment of the present invention, not all of the embodiments. Based on the described embodiment of the present invention, all other embodiments obtained by ordinary technicians in this field belong to the scope of protection of the present invention.
[0032] In view of the defects and problems of the prior art, the present application provides an OLAP query processing method for a GPU platform. Law, The method comprises:
[0033] Step 1, in response to the SQL statement instruction of the OLAP query processing, the dimension table and the fact table in the database are loaded into the global memory of the GPU platform;
[0034] Step 2, parsing the SQL statement instructions and constructing a virtual fact table;
[0035] Step 3: Implement the selection, projection and grouping operations of the dimension table related to the SQL statement instruction through the kernel function of the first stage;
[0036] Step 4: Based on the results of dimension table selection, projection, and grouping operations, the virtual fact table is partitioned according to the thread block structure of the GPU platform through the kernel function of the second stage, and then batch processed by each thread block. Each thread uses row processing to implement filtering, projection, grouping, and aggregation calculations of the fact table.
[0037] In the prior art, GPU-side OLAP query and vectorized query methods use thread block vectorization and intra-thread vectorization processing, while this method uses a thread block vectorization and intra-thread "row processing" model. Thread block vectorization processing improves the computing efficiency of the GPU, and intra-thread row processing solves the problem of single thread overload caused by vectorization affecting GPU performance, as well as the problem of a large number of intermediate result materialization and communication overhead between different operator kernel functions faced by the one-operator query processing method at a time. The intermediate processing results are recorded by faster register cache lines to eliminate memory materialization access overhead.
[0038] In one embodiment, the process of the method of the present application is described in detail below.
[0039] The present application provides an OLAP query processing method for a GPU platform, the method comprising:
[0040] S11, in the OLAP data set, the dimension table primary key is mapped to a single continuous surrogate key (for example, continuous integer data 1, 2, 3...) and the corresponding fact table foreign key is updated, the fact table and dimension table data in the database are loaded into the GPU global memory, and the dimension table column to be accessed is mapped to a virtual fact table column based on the fact table foreign key offset address through the fact table foreign key.
[0041] S12, when parsing the query SQL statement, the query-related dimension table columns are mapped to virtual fact table columns based on array nested access through the fact table foreign key, and the query corresponding multi-table record connection access is mapped to access based on the virtual fact table wide table record. In the kernel function, only the selection and grouping aggregation calculation operations on the corresponding fields of the virtual fact table wide table record need to be performed. Through the dictionary table compression method for the grouping attribute, only the corresponding dictionary table code is stored on the GPU side. During the grouping operation, multiple grouping attributes are mapped to a grouping cube, and the grouping attribute dictionary table code is used as a multi-dimensional address to be mapped to the grouping cube unit for aggregation calculation.
[0042] S13, in order to further improve GPU query performance and possible group cube sparsity problem in S1, the dimension table side operator pushdown is implemented through the table-level kernel function, and the selection, projection, and grouping operations on the dimension table are pushed down to the dimension table preprocessing kernel function. A bitmap or vector is generated as a table-level calculation proxy vector. The virtual fact table column accesses a single dimension table proxy vector through foreign key mapping, reducing the access of multiple columns on the dimension table to one column, improving the data access locality of the proxy vector in the GPU cache, and improving data access performance.
[0043] S14, the query virtual wide table is composed of the columns related to the fact table in the query and the columns of the virtual fact table based on the foreign key of the fact table. In the multi-dimensional computing kernel function, data is sliced according to the number of threads in the GPU thread block, as a number The data is processed by the GPU computing core, and each thread processes one record.
[0044] Specifically, when the dimension table proxy vector corresponding to the virtual fact table column is a bitmap, only the record filtering operation is performed. When the corresponding dimension table proxy vector is a vector, the null value is used to filter the fact table records, and the non-null value is mapped to the multi-dimensional address of the grouping cube composed of the dictionary table encoding dynamically generated by each dimension table, and the fact table metric expression result is mapped to the corresponding unit of the grouping cube for aggregation calculation. Each thread block uses the grouping cube stored in the shared memory to perform aggregation calculation between threads, and finally merges the aggregation calculation results of the grouping cubes in each SM (Streaming Multiprocessor, the streaming multiprocessor on the GPU) in the global memory as the query grouping aggregation cube, and maps it to the original grouping attribute through the dynamic dictionary table, and outputs the grouping aggregation result set.
[0045] From the above embodiments, it can be seen that: the optimized query processing process is completed using a two-stage kernel function. In the first stage, the dimension table filtering, projection, and grouping construction operators are merged into a dimension table preprocessing kernel function. Each dimension table calls the kernel function. There is no data transfer and dependency between the dimension tables, and they are executed in sequence. In the second stage, the fact table filtering, projection, grouping and aggregation calculation operations are merged into a multi-dimensional calculation kernel function. In this stage, only one kernel function is needed to complete the calculation tasks of multiple operators. Each operator (fact table virtual column access, selection, grouping, aggregation calculation) is converted into efficient array access and calculation, and the kernel function algorithm is simplified. Each thread processes one record, and the intermediate results are stored in the register file, avoiding the overhead of multiple kernel function calls, the materialization of intermediate results, and the global memory communication overhead.
[0046] The process and technical advantages of the above method are explained below in a more detailed embodiment in conjunction with the drawings in the specification of this application.
[0047] In a more detailed embodiment, an OLAP query processing method for a GPU platform is provided.
[0048] Figure 1 The query statement in this embodiment is illustrated. In this regard, the method flow in this embodiment includes:
[0049] S21, loading the fact table and dimension table data in the database into the GPU global video memory.
[0050] The data in the database needs to be loaded into the GPU global memory. In the pre-loading phase, the primary keys of the dimension tables customer, part, supplier and orders are updated to surrogate keys. The GPU memory table uses a column storage structure, and the sequentially increasing ID is mapped to the memory offset address of the primary key table record. The foreign key values in the corresponding foreign key table are also updated synchronously, so that the foreign key values can be mapped to the offset address of the corresponding primary key table record, such as Figure 2 shown.
[0051] S22, when parsing the query SQL statement, a virtual fact table wide table is constructed, and the columns accessed in the dimension table are constructed into virtual fact table columns by nesting arrays based on foreign key values, mapping the column records on the corresponding dimension table, such as Figure 2The first record in the fact table lineitem is mapped to the first record in the orders table through the l_orderkey foreign key value 1, and the o_custkey foreign key value 4 is obtained. Then the foreign key value is mapped to the fourth record in the customer table to obtain the c_acctbal record value 12000. Each column in the table is stored as a GPU array structure. The array nested access expression of the fact table virtual column c_acctbal is: c_acctbal[o_custkey[l_orderkey[i]]], where i is the offset address of the lineitem table record. In the same way, based on the foreign key columns l_orderkey, l_partkey, l_suppkey constructs the fact table virtual columns c_nationkey, c_ acctbal, p_size, s_nationkey, for GPU Provides a unified view for accessing virtual wide tables.
[0052] When each virtual fact table wide table record is accessed, the corresponding dimension table record value is accessed through the nested array address, and the corresponding where selection operation is performed. The grouping attribute value is mapped to the group cube constructed by the original dictionary table of the corresponding attribute of group by according to the stored dictionary table encoding, such as Figure 1 In the group by attribute c_nation, s_nation, and p_size, the original dictionary table sizes are 25, 25, and 20 respectively. The query grouping cube is Agg
[25]
[25]
[20] . The first fact table record meets the where selection condition, and the l_quantity field value is mapped to the Agg[7]
[16]
[12] cell for cumulative calculation.
[0053] S23, dimension table selection, projection and grouping operations.
[0054] Through the dimension table preprocessing kernel function, the corresponding selection, projection, and grouping operations of the dimension table in the SQL query are applied to the dimension table, and a proxy vector is generated for each dimension table related to the query as a compressed vector structure of the selection, projection, and grouping operation results on the dimension table.
[0055] The specific process of this step can be as follows Figure 3As shown in the figure, the clauses where c_acctbal>8500and c_nationkey<10 and group by c_nation on the customer table are applied to the customer table through the dimension table preprocessing kernel function. A dynamic dictionary table is created for the c_nationkey of the records that meet the conditions, and a proxy vector is created. The null value NULL is used to represent the records that do not meet the selection conditions. The corresponding dynamic dictionary table code is filled in the corresponding proxy vector unit of the records that meet the conditions. The dimension table preprocessing kernel function is called in turn to execute the query clauses on the part table and the supplier table, and the corresponding proxy vector is created. The length of the corresponding dictionary table is 2, 2, 2, and the grouping cube Agg[2][2][2] is created for grouping aggregation calculation.
[0056] S24, filtering, projection, grouping and aggregation calculation of fact tables.
[0057] In this detailed embodiment, Figure 4 As shown, it is assumed that there are two SMs in the GPU, corresponding to two BLOCKs, and each BLOCK has four threads in the thread block. Four records in the virtual fact table width table are batch processed as a data strip by a thread block of a BLOCK of the GPU. Each thread in the thread block processes one record. The GPU can execute parallel query processing tasks on the thread blocks in two BLOCKs at the same time.
[0058] By obtaining the compressed vector structure through step S23, each SM creates a grouping cube Agg[2][2][2]. When the records accessed by the threads in the thread block meet the non-empty condition of the virtual fact table column value corresponding to the dimension table proxy vector, such as Figure 3 In fact table record 1, the l_quantity value is mapped to the Agg[1][1][0] grouping cube unit corresponding to the dynamic dictionary compression code in the virtual column of the fact table for cumulative calculation. In this thread block, the shared memory grouping cube is used for aggregation calculation, and the shared memory in the thread block completes the data calculation in the thread block in the form of atomicAdd operation. Each SM generates a local grouping cube. After all fact table records are processed, the grouping cubes of each SM are merged and calculated in the global memory of the GPU to generate a global grouping cube, and then mapped to the dynamic dictionary table according to each dimension address to restore it to the original groupby attribute output.
[0059] The aggregation calculation in this detailed embodiment adopts a calculation method based on strip calculation.
[0060] like Figure 4 As shown in the figure, the processing model of strip computing is characterized by vectorized processing at the thread block level. It is row-wise processing at the thread level.
[0061] The size of the stripe is kept consistent with the number of threads in the thread block, so as to achieve the effect of vectorized batch processing at the thread block level and one-to-one row processing at the thread level. When computing on the GPU, the corresponding number of data stripes are read from the global memory at the granularity of data stripes according to the number of SMs or the BLOCK parameter configuration value, and processed in parallel on the GPU.
[0062] Dynamic dictionary table compression uses the actual number of groups to build a group cube based on the selection conditions. Usually, the group cube is much smaller than the group cube built by the original group dictionary table, so it can be stored in a smaller but higher-performance shared memory. GPU atomic operations are used to implement shared group aggregation calculations within thread blocks, greatly reducing the latency of frequent writing of group aggregation results to global memory. The query processing kernel function on the fact table virtual wide table only uses simple array data structures and efficient array access, reducing the complexity of the GPU implementation code, so that complex multi-operator query tasks can be integrated into one kernel function, eliminating the communication latency between different operators in the kernel function and the cost of materializing intermediate results, and improving GPU query processing performance.
[0063] In the several embodiments provided by the present invention, it should be understood that the disclosed method can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the above units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0064] The above-mentioned integrated unit implemented in the form of a software functional unit can be stored in a computer-readable storage medium. The above-mentioned software functional unit is stored in a storage medium, including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor (Processor) to perform some steps of the above-mentioned method of each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (Read-Only Memory, ROM), random access memory (Random Access Memory, RAM), disk or optical disk and other media that can store program codes.
[0065] The above are only preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. An OLAP query processing method for a GPU platform, characterized in that: The method comprises: Step 1, in response to the SQL statement instruction of the OLAP query processing, the dimension table and the fact table in the database are loaded into the global memory of the GPU platform; Step 2, parsing the SQL statement instructions and constructing a virtual fact table; Step 3: Implement the selection, projection and grouping operations of the dimension table related to the SQL statement instruction through the kernel function of the first stage; Step 4: Based on the results of dimension table selection, projection, and grouping operations, the virtual fact table is partitioned according to the thread block structure of the GPU platform through the kernel function of the second stage, and then batch processed by each thread block. Each thread uses row processing to implement filtering, projection, grouping, and aggregation calculations of the fact table.
2. The OLAP query processing method for GPU platform according to claim 1, characterized in that: In step 1, the data storage model in the global memory includes: Map the dimension table primary key to a continuous surrogate key and update the corresponding fact table foreign key. Map the dimension table column to be accessed to a virtual fact table column based on the fact table foreign key offset address through the fact table foreign key.
3. The OLAP query processing method for GPU platform according to claim 2, characterized in that: In step 2, query-related dimension table columns are mapped to virtual fact table columns based on array nested access through fact table foreign keys, and query-corresponding multi-table record connection access is mapped to virtual fact table wide table record access.
4. The OLAP query processing method for GPU platform according to claim 3, characterized in that: In step 3, the dimension table-side operator pushdown is implemented through the table-level kernel function, the selection, projection, and grouping operations on the dimension table are pushed down to the dimension table preprocessing kernel function, and a bitmap or vector is generated as a table-level calculation proxy vector. The virtual fact table column accesses a single dimension table proxy vector through foreign key mapping, reducing the access of multiple columns on the dimension table to one column.
5. The OLAP query processing method for GPU platform according to claim 4, characterized in that: Specifically include: For each dimension table, the group value corresponding to the GROUP BY clause is projected according to its WHERE condition, and a dynamic dictionary table is created for data compression. A continuously growing unique ID is assigned to each group, and the group ID is stored in the corresponding unit of the proxy vector with the same length as the dimension table. Records that do not meet the WHERE condition are stored as null values in the corresponding proxy vector unit. When only the WHERE clause exists on the dimension table, the proxy vector is simplified to a proxy bitmap, with 1 and 0 representing records that meet and do not meet the WHERE condition respectively.
6. The OLAP query processing method for GPU platform according to claim 4, characterized in that: In S4, the operation mode of filtering, projecting, grouping and aggregation calculation of the virtual fact table is determined according to the dimension table proxy vector.
7. The OLAP query processing method for GPU platform according to claim 6, characterized in that: Specifically include: When the dimension table proxy vector corresponding to the virtual fact table column is a bitmap, only record filtering operations are performed. When the corresponding dimension table proxy vector is a vector, null values are used to filter fact table records, and non-null values are mapped to the multi-dimensional addresses of the grouping cube composed of the dictionary table encoding dynamically generated for each dimension table. The fact table metric expression results are mapped to the corresponding units of the grouping cube for aggregation calculation.
8. The OLAP query processing method for GPU platform according to claim 7, characterized in that: During the aggregation calculation, each thread block uses the grouping cube stored in the shared memory to perform aggregation calculations between threads, and finally merges the aggregation calculation results of each grouping cube in the global memory as the query grouping aggregation cube, and maps it to the original grouping attributes through the dynamic dictionary table to output the grouping aggregation result set.