SQL (Structured Query Language) request data acceleration processing method and device and DPU (Distributed Processing Unit) based on KPU architecture

By determining the target column on the DPU chip and merging the Filter and Project operators, the problem of data transmission without computation in traditional Spark SQL is solved, achieving efficient and reliable SQL request data processing.

CN120849435APending Publication Date: 2025-10-28YUSUR TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510679044.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-26
Publication Date
2025-10-28

AI Technical Summary

Technical Problem

When traditional Spark SQL offloads SQL request data processing to devices other than the CPU, data that does not need to participate in the calculation is also transmitted, resulting in reduced processing efficiency and increased storage resource utilization.

Method used

The DPU chip, based on the KPU architecture, determines the target column corresponding to the SQL request statement before executing the Scan operator, and uses the merged Filter and Project operators to perform data scanning and filtering with the Scan operator, transmitting only the necessary data to the DPU memory unit.

Benefits of technology

Improves the efficiency and reliability of SQL request data processing, reduces CPU resource usage and storage resource usage of offload devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120849435A_ABST
    Figure CN120849435A_ABST
Patent Text Reader

Abstract

The invention provides an SQL request data acceleration processing method and device and a DPU based on a KPU architecture, the method is executed by a data processor DPU based on a kernel processor KPU architecture, and the method comprises the steps that a target column corresponding to an SQL request statement is determined from a data table stored outside the DPU; and scanning the target column in the data table on the basis of a Scan operator preset in the DPU, so as to take a field corresponding to the target column obtained by scanning as request result data corresponding to the SQL request statement, and transferring the request result data into a memory unit of the DPU. On the basis of reducing the resource occupancy rate of the CPU and improving the calculation scheduling efficiency and reliability of the CPU, the SQL request data processing efficiency and reliability can be effectively improved, and the storage resource occupancy rate of the unloading equipment used for processing the SQL request data outside the CPU can be effectively reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing technology, and in particular to a method, apparatus, and DPU based on KPU architecture for accelerating SQL request data processing. Background Technology

[0002] Spark SQL is a computation module of Apache Spark, a fast and general-purpose computing engine designed for large-scale data processing, specifically for handling structured data. Traditional Spark SQL breaks down a single SQL statement into a physical execution plan tree, the specific of which varies depending on the SQL statement. Regardless of the SQL statement, there is a basic Scan operator. The Scan operator locates one or more corresponding tables based on the table names mentioned in the SQL statement, scans these tables, reads data from a database on devices such as hard drives into memory, and then performs subsequent calculations on the CPU. However, this increases CPU resource consumption and affects the efficiency and reliability of CPU computation scheduling. Therefore, it is necessary to offload the SQL request data processing operations to devices other than the CPU.

[0003] Currently, while offloading the SQL request data processing to devices outside the CPU can effectively reduce CPU resource utilization and improve the efficiency and reliability of CPU computation scheduling, it requires using a Scan operator located outside the CPU to offload all data in the database specified by the SQL statement to the device outside the CPU before performing the SQL request data processing. This results in data that does not need to participate in the calculation being transferred to the device outside the CPU as well. This not only reduces the efficiency of SQL request data processing but also increases the storage resource utilization of the device outside the CPU. Summary of the Invention

[0004] In view of this, embodiments of this application provide a method, apparatus and DPU based on KPU architecture for accelerating SQL request data processing, in order to eliminate or improve one or more defects existing in the prior art.

[0005] One aspect of this application provides a method for accelerating SQL request data processing, the method being executed by a Data Processing Unit (DPU) based on a Kernel Processor (KPU) architecture, the method comprising:

[0006] The target column corresponding to the SQL request statement is determined from a data table stored outside the DPU;

[0007] Based on the Scan operator preset in the DPU, the target column in the data table is scanned, and the field corresponding to the target column obtained by scanning is used as the request result data corresponding to the SQL request statement, and the request result data is transferred to the memory unit of the DPU.

[0008] In some embodiments of this application, before determining the target column corresponding to the SQL request statement in a data table stored outside the DPU, the method further includes:

[0009] It receives SQL request statements, reads the specified data table from the SQL request statement, determines the attribute type of the target, the judgment conditions, and the attribute type of the returned object.

[0010] In some embodiments of this application, determining the target column corresponding to the SQL request statement from a data table stored outside the DPU includes:

[0011] According to the SQL request statement, select the columns corresponding to the judgment target and the returned object respectively in the data table stored outside the DPU, and perform field filtering on the column corresponding to the judgment target based on the SQL request statement. Each column contains all fields of an attribute type in the data table.

[0012] Based on the column corresponding to the target after field filtering, the column corresponding to the returned object is further filtered to obtain the target column corresponding to the SQL request statement.

[0013] In some embodiments of this application, the step of selecting the columns corresponding to the judgment target and the returned object from a data table stored outside the DPU according to the SQL request statement, and performing field filtering on the column corresponding to the judgment target based on the SQL request statement, includes:

[0014] The filter operator preset in the DPU is invoked. Based on the attribute types of the judgment target and the return object specified in the SQL request statement, the columns corresponding to the judgment target and the return object are selected in the data table stored outside the DPU. Based on the judgment conditions specified in the SQL request statement, the column corresponding to the judgment target is filtered.

[0015] In some embodiments of this application, before performing field filtering on the column corresponding to the returned object based on the column corresponding to the target after field filtering, the method further includes:

[0016] The Filter operator is invoked to mask fields that meet the judgment conditions in the column corresponding to the judgment target using a first identifier, and fields that do not meet the judgment conditions using a second identifier, wherein the first identifier and the second identifier are different.

[0017] In some embodiments of this application, the step of filtering the column corresponding to the returned object based on the column corresponding to the target after field filtering to obtain the target column corresponding to the SQL request statement includes:

[0018] The Project operator, which is preset in the DPU, is invoked. In the data table, the first identifier in the column corresponding to the judgment target is logically summed with each field in the column corresponding to the returned object to obtain the filtered column corresponding to the returned object. The filtered column corresponding to the returned object is then used as the target column corresponding to the SQL request statement.

[0019] Another aspect of this application provides an SQL request data acceleration processing apparatus, the apparatus being disposed in a data processor DPU based on a kernel processor KPU architecture, the apparatus comprising:

[0020] The pre-operation module is used to determine the target column corresponding to the SQL request statement from a data table stored outside the DPU;

[0021] The Scan operator module is used to scan the target column in the data table based on the Scan operator preset in the DPU, so as to use the field corresponding to the scanned target column as the request result data corresponding to the SQL request statement, and to transfer the request result data to the memory unit of the DPU.

[0022] In some embodiments of this application, the pre-operation module includes: a Filter operator unit;

[0023] The Filter operator unit is used to call the Filter operator preset in the DPU, select the columns corresponding to the judgment target and the return object respectively in the data table stored outside the DPU according to the attribute types of the judgment target and the return object specified in the SQL request statement, and perform field filtering on the column corresponding to the judgment target based on the judgment conditions specified in the SQL request statement.

[0024] In some embodiments of this application, the pre-operation module includes: a Project operator unit;

[0025] The Project operator unit is used to call the Project operator preset in the DPU, and in the data table, perform a logical summation operation on the first identifier in the column corresponding to the judgment target and each field in the column corresponding to the returned object to obtain the filtered column corresponding to the returned object, and use the filtered column corresponding to the returned object as the target column corresponding to the SQL request statement.

[0026] The third aspect of this application provides a DPU based on a KPU architecture, wherein the DPU includes an SQL request data acceleration processing device for executing the SQL request data acceleration processing method.

[0027] The SQL request data acceleration processing device communicates with a database outside the DPU to search for the data table in the database;

[0028] The SQL request data acceleration processing device communicates with the interactive query module in the distributed computing platform to receive SQL request statements from the interactive query module.

[0029] The SQL request data acceleration processing method provided in this application is executed by a data processor (DPU) based on a kernel processor (KPU) architecture. The method determines the target column corresponding to the SQL request statement by storing it in a data table outside the DPU; based on a Scan operator preset in the DPU, the target column in the data table is scanned, and the field corresponding to the scanned target column is used as the request result data corresponding to the SQL request statement. The request result data is then transferred to the memory unit of the DPU. This method effectively improves the efficiency and reliability of SQL request data processing while reducing CPU resource utilization and improving the efficiency and reliability of CPU computation scheduling. It also effectively reduces the storage resource utilization of the offloading device used to process SQL request data outside the CPU.

[0030] Additional advantages, objectives, and features of this application will be set forth in part in the description which follows, and will in part become apparent to those skilled in the art upon review of the following description, or may be learned by practice of the application. The objectives and other advantages of this application can be realized and obtained by means of the structures specifically pointed out in the specification and drawings.

[0031] Those skilled in the art will understand that the purposes and advantages that can be achieved with this application are not limited to those specifically described above, and that the above and other purposes that this application can achieve will be more clearly understood from the following detailed description. Attached Figure Description

[0032] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, do not constitute a limitation thereof. The components in the drawings are not drawn to scale but are merely for illustrating the principles of this application. For ease of illustration and description of certain parts of this application, corresponding portions in the drawings may be enlarged, i.e., may appear larger relative to other components in an exemplary device actually manufactured according to this application. In the drawings:

[0033] Figure 1 This is a schematic diagram of the first type of SQL request data acceleration processing method in one embodiment of this application.

[0034] Figure 2 This is a schematic diagram of a second process of the SQL request data acceleration processing method in one embodiment of this application.

[0035] Figure 3 This is a schematic diagram of the plan tree corresponding to the execution of SQL statements using DPU, provided as an example in this application.

[0036] Figure 4 This is a schematic diagram of the execution plan tree for an example of this application, which uses DPU and combines the operations of the FILTER operator, Project operator and Scan operator to execute SQL statements.

[0037] Figure 5 This is a schematic diagram of the third process of the SQL request data acceleration processing method in one embodiment of this application.

[0038] Figure 6 This is a schematic diagram of a first structure of an SQL request data acceleration processing device in one embodiment of this application.

[0039] Figure 7 This is a schematic diagram of a second structure of the SQL request data acceleration processing device in one embodiment of this application.

[0040] Figure 8 This is a schematic diagram of the structure of a DPU based on the KPU architecture in one embodiment of this application. Detailed Implementation

[0041] To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the embodiments and accompanying drawings. Here, the illustrative embodiments and their descriptions are used to explain this application, but are not intended to limit it.

[0042] It should also be noted that, in order to avoid obscuring this application with unnecessary details, only the structures and / or processing steps closely related to the solution according to this application are shown in the accompanying drawings, while other details that are not closely related to this application are omitted.

[0043] It should be emphasized that the term "include / comprises" when used herein refers to the existence of features, elements, steps or components, but does not exclude the existence or addition of one or more other features, elements, steps or components.

[0044] It should also be noted that, unless otherwise specified, the term "connection" herein may refer not only to a direct connection but also to an indirect connection involving an intermediate.

[0045] In the following description, embodiments of the present application will be illustrated with reference to the accompanying drawings. In the drawings, the same reference numerals represent the same or similar parts, or the same or similar steps.

[0046] In one or more embodiments of this application, Spark is a distributed computing platform, a computing framework written in Scala, and a fast, general-purpose, and scalable big data analysis engine based on memory. Spark SQL is a module based on Spark for processing structured data. Spark SQL supports querying and loading from various data sources, is compatible with Hive, and can execute SQL statements using JDBC / ODBC connections. It provides important technical support for the Spark framework in structured data analysis.

[0047] Spark SQL allows users to execute SQL queries and read / write operations using standard SQL statements, and also allows them to execute queries and read / write operations on Hive repositories using Hive SQL.

[0048] In order to improve the efficiency of SQL request data processing while reducing CPU resource utilization and improving the efficiency and reliability of CPU computing scheduling, this application provides an SQL request data acceleration processing method, an SQL request data acceleration processing device for executing the SQL request data acceleration processing method, and a DPU based on KPU architecture. These methods can improve the efficiency of SQL request data processing while reducing CPU resource utilization and improving the efficiency and reliability of CPU computing scheduling.

[0049] The following examples will provide a detailed description.

[0050] Based on this, embodiments of this application provide an SQL request data acceleration processing method that can be implemented by an SQL request data acceleration processing device, executed by a data processor DPU based on a kernel processor KPU architecture, see [link to relevant documentation]. Figure 1 The SQL request data acceleration processing method specifically includes the following:

[0051] Step 100: Determine the target column corresponding to the SQL request statement from the data table stored outside the DPU.

[0052] In step 100, the target column refers to the column in the data table containing the query result specified by the SQL request statement.

[0053] Since CPUs, as general-purpose processing chips, do not offer significant advantages in high-intensity big data computation, computational power becomes the main performance bottleneck when Spark SQL is run on CPU. Therefore, this application employs a DPU chip based on the KPU architecture to accelerate SQL request data processing.

[0054] Understandably, DPU chips based on the KPU architecture, as dedicated data processing chips, offer significantly higher performance than CPUs when handling complex data computations. Therefore, offloading Spark SQL data computations from the CPU to the DPU can greatly improve Spark SQL performance, accelerate Spark SQL computations in big data scenarios, and allow the CPU to focus on Spark computation scheduling while the DPU focuses on data computations within Spark SQL.

[0055] However, traditional Spark SQL decomposes a single SQL statement into a physical execution plan tree, the specific of which varies depending on the SQL statement. Regardless of the SQL statement, there is a basic Scan operator. The Scan operator finds one or more corresponding tables based on the table names mentioned in the SQL statement, scans these tables, reads the data from disk into memory, and then performs subsequent calculations using the CPU or DPU. However, most SQL statements, in addition to the Scan operator, also have a Filter operator, which corresponds to the WHERE clause in the SQL expression. This means that because the Scan operator, located on a device outside the CPU, is used to offload all data from the database specified by the SQL statement to a device outside the CPU for SQL request data processing, data that doesn't need to participate in the calculation is also transferred to the device outside the CPU. This not only reduces the efficiency of SQL request data processing but also increases the storage resource usage of the device outside the CPU.

[0056] Based on this, in one or more embodiments of this application, a method for accelerating SQL request data processing based on DPU is proposed. The basis for executing the columnar data is the above-mentioned step 100. Before executing the Scan operator, the target column corresponding to the SQL request statement needs to be determined from the data table stored outside the DPU. Then, in step 200 below, the DPU's Scan operator scans the target column in the data table to use the field corresponding to the scanned target column as the request result data corresponding to the SQL request statement, and transfers the request result data to the memory unit of the DPU.

[0057] It is understood that the SQL request statement refers to an SQL statement containing an operation type, which may include SELECT, etc.; and each SQL request statement contains: operation type, data table in row-based database, attribute type corresponding to the judgment target, judgment condition, and attribute type corresponding to the return object. The judgment target refers to the attribute used as the judgment object in the judgment condition, and the return object specifies the attribute to be returned. In a special case, the attribute corresponding to the target and the attribute corresponding to the return object may be the same, but usually they are different.

[0058] Step 200: Based on the Scan operator preset in the DPU, scan the target column in the data table, so as to use the field corresponding to the target column obtained by scanning as the request result data corresponding to the SQL request statement, and transfer the request result data to the memory unit of the DPU.

[0059] In other words, when dealing with large amounts of data, the Scan operator scans and transmits data that does not need to be used in the calculation, resulting in data transmission losses and a decrease in memory utilization. To address this, this embodiment of the application first locks the target column specified by the SQL request statement in the data table before executing the Scan operator operation, and then executes step 200 to perform the Scan operator operation and transfers the operation result to the memory unit of the DPU, thereby reducing the transmission of data that does not participate in the calculation between the database and the DPU, and effectively saving the memory space of the DPU.

[0060] As can be seen from the above description, the SQL request data acceleration processing method provided in this application embodiment can effectively improve the efficiency and reliability of SQL request data processing by reducing the CPU resource utilization rate and improving the efficiency and reliability of CPU computing scheduling, and can effectively reduce the storage resource utilization rate of the offloading device used to process SQL request data outside the CPU.

[0061] To further improve the effectiveness and reliability of SQL request data acceleration processing, an SQL request data acceleration processing method is provided in an embodiment of this application, see [link to relevant documentation]. Figure 2 The SQL request data acceleration processing method includes the following content before step 100:

[0062] Step 010: Receive the SQL request statement, and read the specified data table, determine the attribute type of the target, the judgment condition, and the attribute type of the returned object from the SQL request statement.

[0063] To further improve the effectiveness and reliability of determining the target column corresponding to the SQL request statement in a data table stored outside the DPU, an SQL request data acceleration processing method is provided in this application embodiment, see [link to relevant documentation]. Figure 2 Step 100 in the SQL request data acceleration processing method specifically includes the following:

[0064] Step 110: Select the columns corresponding to the judgment target and the returned object from the data table stored outside the DPU according to the SQL request statement, and perform field filtering on the column corresponding to the judgment target based on the SQL request statement, wherein each column contains all fields of an attribute type in the data table.

[0065] It is understood that the data table stored outside the DPU is a row-based data table, in which each row represents an independent record, and each record contains a field corresponding to each attribute type. Each column is used to store all fields corresponding to each attribute type, and each column in the data table corresponds one-to-one with each of the data types.

[0066] In one example of this application, the rows of the data table refer to all rows except the first row, which is used only to represent the various data types. As shown in Table 1, the row-based data table can contain three attribute types: "name", "age", and "score".

[0067] Table 1 is a row-based data table.

[0068] Name Age Score Zhang San 17 95 Li Si 19 59 Wang Wu 21 64

[0069] Step 120: Based on the column corresponding to the target after field filtering, perform field filtering on the column corresponding to the returned object to obtain the target column corresponding to the SQL request statement.

[0070] Based on Table 1 above, in one example, the SQL request statement can be:

[0071] The query "SELECT name FROM table1 where age>18" means: find the names in table1 whose age is greater than 18.

[0072] In the SQL request statement above, the operation type is "SELECT"; the data table in the row-oriented database is "TABLE1"; the attribute type corresponding to the target is "name"; the condition is "age>18"; and the attribute type corresponding to the returned object is "age". In the example of the SQL request statement above, the target column shown in Table 2 will be returned.

[0073] Table 2 returns the target columns.

[0074] Name Li Si Wang Wu

[0075] The results show that this SQL statement doesn't involve the `score` column at all. Therefore, when the Scan operator executes, there's no need to read the `score` column data and load it into memory. Because the DPU reads a columnar data structure, the `score` column can be completely skipped when using the DPU. However, when reading the `age` column, the original approach was to read the entire `score` column before performing calculations on the DPU. But if the WHERE condition is checked during the Scan process, and fields that don't meet the condition are skipped, data transfer can be reduced, memory space can be saved, and this offers a significant advantage for calculations involving large amounts of data.

[0076] Based on this, see Figure 3 For large datasets, the Scan operator scans and transmits data that does not need to be used in the calculation, resulting in data transmission overhead and reduced memory utilization. To address this, this application uses an operator pull-up operation to combine the operations of the FILTER and Project operators with the Scan operator, thereby reducing the transmission of data that does not participate in the calculation and saving memory space.

[0077] In other words, taking Table 1 above as an example, the existing method first executes the Scan operator to scan all columns involved in the SQL statement in the entire table and stores the data in memory. That is, all data in the `name` and `age` columns is read into memory (column-by-column). Then, a FILTER operation is executed to select those with an age greater than 18, generating a mask where 0 represents less than 18 and 1 represents greater than 18. The mask is then applied to the `name` column for a Project operation, where 0 represents not meeting the criteria and 1 represents meeting the criteria, selecting Li Si as a compliant candidate. However, this embodiment merges the Filter, Project, and Scan operators. The merged operator can be called the DPU-integrated Scan operator, and the related plan tree becomes... Figure 4 As shown, a DPU-synthesized Scan operation is first performed on the age column, which returns the mask mentioned above. Then, a synthesized Scan operation is performed on the target column, which is the name column. The Filter condition in the synthesis operator is the mask obtained from the first operator run. Thus, after the second operator operation is completed, the final result is obtained. This will be explained in detail through the following examples.

[0078] Among them, Figure 3 In SQL, the Scan operator reads specified raw data from disk into memory; the Filter operator filters the dataset, selecting data that meets certain conditions; the Project operator, for columnar data, projects data from a dataset to specific columns, essentially projecting high-dimensional data to low-dimensional data. `DPUProjectExec` refers to executing the Project operator via the DPU; `DPU FilterExec` refers to executing the Filter operator via the DPU; and `DPU ScanExec` refers to executing the Scan operator via the DPU. The SQL result is the target column corresponding to the SQL request statement.

[0079] Based on this, in order to further improve the efficiency, effectiveness, and reliability of field filtering on the column corresponding to the judgment target, in an embodiment of this application, a method for accelerating SQL request data processing is provided, see... Figure 5 Step 110 of the SQL request data acceleration processing method specifically includes the following:

[0080] Step 111: Invoke the Filter operator preset in the DPU, select the columns corresponding to the judgment target and the return object respectively in the data table stored outside the DPU according to the attribute types of the judgment target and the return object specified in the SQL request statement, and perform field filtering on the column corresponding to the judgment target based on the judgment conditions specified in the SQL request statement.

[0081] To further improve the efficiency, effectiveness, and reliability of field filtering on the columns corresponding to the returned object, an SQL request data acceleration processing method is provided in this application embodiment, see [link to relevant documentation]. Figure 5 The SQL request data acceleration processing method further includes the following content between steps 111 and 120:

[0082] Step 112: Call the Filter operator to mask the fields that meet the judgment conditions in the column corresponding to the judgment target with a first identifier, and mask the fields that do not meet the judgment conditions with a second identifier, wherein the first identifier and the second identifier are different.

[0083] In one example, the mask can be 1, 0, where 1 is the first identifier and 0 is the second identifier. 1 represents a field that meets the judgment condition, and 0 represents a field that does not meet the judgment condition.

[0084] To further improve the efficiency, effectiveness, and reliability of field filtering on the columns corresponding to the returned object, an SQL request data acceleration processing method is provided in this application embodiment, see [link to relevant documentation]. Figure 5 Step 120 in the SQL request data acceleration processing method specifically includes the following:

[0085] Step 121: Call the Project operator preset in the DPU, and in the data table, perform a logical summation operation on the first identifier in the column corresponding to the judgment target and each field in the column corresponding to the returned object to obtain the filtered column corresponding to the returned object, and use the filtered column corresponding to the returned object as the target column corresponding to the SQL request statement.

[0086] Specifically, a logical summation (AND) operation is performed between the row containing the first identifier of the field that meets the judgment condition and the attribute information row corresponding to the returned object. If the mask corresponding to "Zhang San" is 0, it means that the field "Zhang San" does not meet the judgment condition and is not returned; if the mask corresponding to "Li Si" is 1, it means that "Li Si" meets the judgment condition and is returned.

[0087] Understandably, this application uses operator pull-up operations to combine the operations of the FILTER and Project operators with the Scan operator, thereby reducing the transfer of data that does not participate in the operation and saving memory space.

[0088] In other words, this application embodiment uses a mask and the target result row to perform a logical summation operation to filter out the results that meet the conditions.

[0089] This application also provides an apparatus for executing all or part of the SQL request data acceleration processing method, wherein the apparatus is disposed in a data processor DPU based on a kernel processor KPU architecture. See [link to relevant documentation] Figure 6 The SQL request data acceleration processing device specifically includes the following components:

[0090] The pre-operation module 10 is used to determine the target column corresponding to the SQL request statement from a data table stored outside the DPU;

[0091] Scan operator module 20 is used to scan the target column in the data table based on the Scan operator preset in the DPU, so as to use the field corresponding to the scanned target column as the request result data corresponding to the SQL request statement, and to transfer the request result data to the memory unit of the DPU.

[0092] The embodiments of the SQL request data acceleration processing device provided in this application can be used to execute the processing flow of the SQL request data acceleration processing method embodiments described above. Its functions will not be repeated here, but can be referred to the detailed description of the SQL request data acceleration processing method embodiments described above.

[0093] As can be seen from the above description, the SQL request data acceleration processing device provided in this application embodiment can effectively improve the efficiency and reliability of SQL request data processing while reducing the CPU resource utilization rate and improving the efficiency and reliability of CPU computing scheduling. It can also effectively reduce the storage resource utilization rate of the offloading device used to process SQL request data outside the CPU.

[0094] To further improve the efficiency, effectiveness, and reliability of field filtering on the columns corresponding to the judgment target, an SQL request data acceleration processing device is provided in this application embodiment, see [link to relevant documentation]. Figure 7 The pre-operation module 10 in the SQL request data acceleration processing device includes a Filter operator unit 11;

[0095] The Filter operator unit 11 is used to call the Filter operator preset in the DPU, select the columns corresponding to the judgment target and the return object respectively in the data table stored outside the DPU according to the attribute types of the judgment target and the return object specified in the SQL request statement, and perform field filtering on the column corresponding to the judgment target based on the judgment conditions specified in the SQL request statement.

[0096] To further improve the efficiency, effectiveness, and reliability of field filtering on the columns corresponding to the judgment target, an SQL request data acceleration processing device is provided in this application embodiment, see [link to relevant documentation]. Figure 7 The pre-operation module 10 in the SQL request data acceleration processing device further includes a Project operator unit 12;

[0097] The Project operator unit 12 is used to call the Project operator preset in the DPU, and in the data table, perform a logical summation operation on the first identifier in the column corresponding to the judgment target and each field in the column corresponding to the returned object to obtain the filtered column corresponding to the returned object, and use the filtered column corresponding to the returned object as the target column corresponding to the SQL request statement.

[0098] This application also provides a DPU based on a KPU architecture, wherein the DPU includes an SQL request data acceleration processing device, which is used in the SQL request data acceleration processing method described in the foregoing embodiments; see also Figure 8 The SQL request data acceleration processing device communicates with a database outside the DPU to search for the data table in the database;

[0099] The SQL request data acceleration processing device communicates with the interactive query module in the distributed computing platform to receive SQL request statements from the interactive query module. The distributed computing platform can be Spark, and the interactive query module can be Spark SQL.

[0100] This application also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of the aforementioned SQL request data acceleration processing method. The computer-readable storage medium can be a tangible storage medium, such as random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, floppy disks, hard disks, removable storage disks, CD-ROMs, or any other form of storage medium known in the art.

[0101] Those skilled in the art will understand that the exemplary components, systems, and methods described in conjunction with the embodiments disclosed herein can be implemented in hardware, software, or a combination of both. Whether implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application. When implemented in hardware, it can be, for example, electronic circuits, application-specific integrated circuits (ASICs), appropriate firmware, plug-ins, function cards, etc. When implemented in software, the elements of this application are programs or code segments used to perform the required tasks. The programs or code segments can be stored in a machine-readable medium or transmitted over a transmission medium or communication link via data signals carried on a carrier wave.

[0102] It should be clarified that this application is not limited to the specific configurations and processes described above and shown in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of this application is not limited to the specific steps described and shown. Those skilled in the art can make various changes, modifications, and additions, or change the order of steps, after understanding the spirit of this application.

[0103] In this application, features described and / or illustrated for one embodiment may be used in the same or similar manner in one or more other embodiments, and / or combined with or in place of features of other embodiments.

[0104] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Various modifications and variations can be made to the embodiments of this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

Claims

1. A method for accelerating SQL request data processing, characterized in that, The method is executed by a data processor (DPU) based on a kernel processor (KPU) architecture, and the method includes: The target column corresponding to the SQL request statement is determined from a data table stored outside the DPU; Based on the Scan operator preset in the DPU, the target column in the data table is scanned, and the field corresponding to the target column obtained by scanning is used as the request result data corresponding to the SQL request statement, and the request result data is transferred to the memory unit of the DPU.

2. The SQL request data acceleration processing method according to claim 1, characterized in that, Before determining the target column corresponding to the SQL request statement in the data table stored outside the DPU, the method further includes: It receives SQL request statements, reads the specified data table from the SQL request statement, determines the attribute type of the target, the judgment conditions, and the attribute type of the returned object.

3. The SQL request data acceleration processing method according to claim 2, characterized in that, The process of determining the target column corresponding to the SQL request statement from a data table stored outside the DPU includes: According to the SQL request statement, select the columns corresponding to the judgment target and the returned object respectively in the data table stored outside the DPU, and perform field filtering on the column corresponding to the judgment target based on the SQL request statement. Each column contains all fields of an attribute type in the data table. Based on the column corresponding to the target after field filtering, the column corresponding to the returned object is further filtered to obtain the target column corresponding to the SQL request statement.

4. The SQL request data acceleration processing method according to claim 3, characterized in that, The step of selecting the columns corresponding to the judgment target and the returned object from a data table stored outside the DPU according to the SQL request statement, and performing field filtering on the column corresponding to the judgment target based on the SQL request statement, includes: The filter operator preset in the DPU is invoked. Based on the attribute types of the judgment target and the return object specified in the SQL request statement, the columns corresponding to the judgment target and the return object are selected in the data table stored outside the DPU. Based on the judgment conditions specified in the SQL request statement, the column corresponding to the judgment target is filtered.

5. The SQL request data acceleration processing method according to claim 4, characterized in that, Before performing field filtering on the column corresponding to the returned object based on the column corresponding to the target after field filtering, the method further includes: The Filter operator is invoked to mask fields that meet the judgment conditions in the column corresponding to the judgment target using a first identifier, and fields that do not meet the judgment conditions using a second identifier, wherein the first identifier and the second identifier are different.

6. The SQL request data acceleration processing method according to claim 5, characterized in that, The step of filtering the columns corresponding to the returned object based on the columns corresponding to the target after field filtering to obtain the target column corresponding to the SQL request statement includes: The Project operator, which is preset in the DPU, is invoked. In the data table, the first identifier in the column corresponding to the judgment target is logically summed with each field in the column corresponding to the returned object to obtain the filtered column corresponding to the returned object. The filtered column corresponding to the returned object is then used as the target column corresponding to the SQL request statement.

7. A device for accelerating SQL request data processing, characterized in that, The device is installed in a data processor (DPU) based on a kernel processor (KPU) architecture, and the device includes: The pre-operation module is used to determine the target column corresponding to the SQL request statement from a data table stored outside the DPU; The Scan operator module is used to scan the target column in the data table based on the Scan operator preset in the DPU, so as to use the field corresponding to the scanned target column as the request result data corresponding to the SQL request statement, and to transfer the request result data to the memory unit of the DPU.

8. The SQL request data acceleration processing device according to claim 7, characterized in that, The pre-operation module includes: a Filter operator unit; The Filter operator unit is used to call the Filter operator preset in the DPU, select the columns corresponding to the judgment target and the return object respectively in the data table stored outside the DPU according to the attribute types of the judgment target and the return object specified in the SQL request statement, and perform field filtering on the column corresponding to the judgment target based on the judgment conditions specified in the SQL request statement.

9. The SQL request data acceleration processing device according to claim 8, characterized in that, The pre-operation module includes: a Project operator unit; The Project operator unit is used to call the Project operator preset in the DPU, and in the data table, perform a logical summation operation on the first identifier in the column corresponding to the judgment target and each field in the column corresponding to the returned object to obtain the filtered column corresponding to the returned object, and use the filtered column corresponding to the returned object as the target column corresponding to the SQL request statement.

10. A DPU based on a KPU architecture, characterized in that, The DPU is equipped with an SQL request data acceleration processing device, which is used to execute the SQL request data acceleration processing method according to any one of claims 1 to 6. The SQL request data acceleration processing device communicates with a database outside the DPU to search for the data table in the database; The SQL request data acceleration processing device communicates with the interactive query module in the distributed computing platform to receive SQL request statements from the interactive query module.

Citation Information

Patent Citations

  • Spark SQL acceleration method based on GPU

    CN116303550A

  • Data operation method and device, electronic equipment, storage medium and program product

    CN117851413A

  • Method and system for executing structured query statement aggregation calculation based on DPU

    CN118861097A