An Execution Management Method and Device for Spark SQL Query Plan Trees Based on DPU

By deploying the rule plug-in and adaptive query execution mode in Spark SQL, we can determine whether the operators in the query plan tree can be offloaded to the DPU, and by configuring the execution method of the single-manage query plan tree, the data format conversion and sub-plan tree offload logic synchronization problems when mixed calls to the CPU and DPU are solved, and efficient computing performance and resource utilization are achieved.

CN118861096BActive Publication Date: 2025-06-27YUSUR TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410953084.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-16
Publication Date
2025-06-27
Estimated Expiration
2044-07-16

AI Technical Summary

Technical Problem

In the prior art, when mixing the CPU and DPU, it is necessary to repeatedly convert the data format, resulting in large calculation overhead and inability to effectively synchronize the unloading logic between sub-plan trees, resulting in large row-column conversion overhead and data replication overhead.

Method used

It provides an execution management method of Spark SQL query plan tree based on DPU. By obtaining and deploying rule plug-ins, self-testing of adaptive query execution mode, traversing the query plan tree, determining whether the operator can be offloaded to the DPU, and by configuring the execution method of the query plan tree, avoiding the introduction of additional row-column conversion operators and data replication.

Benefits of technology

It effectively avoids the introduction of additional row-column conversion operators and data replication overhead, improves computing performance and resource utilization, and ensures the overall deployment and optimization of query tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118861096B_ABST
    Figure CN118861096B_ABST
Patent Text Reader

Abstract

The present invention provides an execution management method and device for a Spark SQL query plan tree based on DPU, which allows hybrid computing of DPU and CPU for query tasks, improves computing performance and resource utilization rate. When the adaptive query execution mode is not running, the query plan tree is deployed as a whole. When all operators conform to the types supported by DPU, the operators are preferentially offloaded to DPU for operation, otherwise the whole is handed over to CPU for processing. When the adaptive query execution mode is running, first judge whether the currently intercepted sub-plan tree has an adaptive query execution operator. If so, first judge whether the entire query plan tree can be offloaded and marked in the configuration sheet. If not, query the mark in the configuration sheet and hand over the current sub-plan tree to DPU or CPU for processing according to the mark. This execution management method deploys the whole for a single query task, avoiding the introduction of additional row-column conversion operators and the data replication overhead brought by the row-column conversion operators, saving resources and improving computing performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and in particular, to an execution management method and device for a Spark SQL query plan tree based on a DPU. Background Art

[0002] Apache Spark is a general big data processing and computing framework. Compared with the MapReduce computing model proposed by Hadoop, Spark introduces the abstraction of resilient distributed datasets (RDDs). It provides richer operations, such as map, filter, reduce, union, etc. These rich operations enable the rapid development of very complex data processing programs based on the Spark framework. On the basis of Spark core, components such as Spark Streaming, Spark SQL, Spark Mllib, SparkR, and GraphX are provided to support multiple application scenarios. In terms of computing performance, the caching mechanism and lazy execution mechanism of Spark enable Spark to merge multiple operations and thus make full use of memory, resulting in a significant improvement in computing performance.

[0003] Spark SQL is one of the computing modules of Spark, dedicated to processing structured data. Spark SQL allows users to use standard SQL statements to perform queries and read / write operations on a Hive warehouse, so it has become the most widely used module in Spark applications. With the rapid development of Spark SQL, although its performance has been greatly improved, all its computations are still based on the central processing unit (CPU). In addition to maintaining the scheduling of the entire computation, the CPU also requires additional computing power for data-intensive computations. The computing power during CPU operation has gradually become the main bottleneck of performance.

[0004] As a dedicated data processing chip, the data center processor (DPU) chip based on the KPU architecture has extremely high performance improvement compared with the CPU when processing complex data computations. However, there are a very large number of built-in operators and functions in Spark SQL, not to mention various user-defined functions. At present, it is impossible for the DPU to fully cover them, that is, at the current stage, it is impossible to unload all computing tasks to the DPU in all computing scenarios.

[0005] Existing hybrid operation acceleration solutions offload some operators to the DPU and adapt to the differences in input data formats required by the CPU and DPU by inserting row-column conversion operators and data replication. Since the calculations on the CPU are row-based data calculations and the calculations on the DPU are column-based data calculations, if an operator executed on the CPU finds that the data is column-based data, the column-based data needs to be converted to row-based data, which requires copying the data from the DPU to the CPU. Conversely, if an operator executed on the DPU finds that the data is row-based data, the row-based data needs to be converted to column-based data, which requires copying the data from the CPU to the DPU. This method introduces row-column conversion calculations and a relatively large computational overhead. At the same time, row-column conversion requires copying data back and forth between the CPU and the DPU, bringing additional data replication overhead and resulting in poor overall performance of hybrid operation calculations.

[0006] Since Spark 3.0 introduced AQE (Adaptive Query Execution), that is, the adaptive query execution mode, the original query plan tree is split into multiple sub-plan trees. The existing hybrid operation framework cannot synchronize the offloading logic between sub-plan trees, that is, each sub-plan tree is optimized independently, and additional row-column conversion overhead may be introduced between sub-plan trees. The existing hybrid operation framework cannot well evaluate this part of the cost, so it cannot achieve global optimization and may even result in negative optimization.

[0007] Therefore, there is an urgent need for a new query plan tree execution management solution. Summary of the Invention

[0008] In view of this, the embodiments of the present invention provide a method and device for executing and managing a Spark SQL query plan tree based on a DPU to eliminate or improve one or more defects existing in the prior art and solve the problem of large computational overhead caused by repeated data format conversion when the CPU and DPU are mixedly called in the prior art.

[0009] One aspect of the present invention provides a method for executing and managing a Spark SQL query plan tree based on a DPU. The method runs on a host device and includes the following steps:

[0010] Obtain and deploy a rule plugin, and the rule plugin is controlled to start and stop by a preset control signal;

[0011] Perform a self-check on the adaptive query execution mode based on the rule plugin:

[0012] If the adaptive query execution mode is in an unrun state, traverse the query plan tree and verify all operators that make up the query plan tree based on a set standard: If there is a target operator that cannot be offloaded to the data center processor, hand over the entire query plan tree to the central processor for execution; if there is no such target operator, offload the query plan tree to the data center processor for execution;

[0013] If the adaptive query execution mode is in a running state, traverse the query plan tree and verify whether there is an adaptive query execution operator for marking a new query task in the current sub-plan tree: If there is such an adaptive query execution operator in the current sub-plan tree, search for and verify the entire query plan tree, recursively determine whether there is a target operator in the current sub-plan tree that cannot be offloaded to the data center processor, and write a mark indicating whether the query plan tree can be offloaded into the configuration sheet through a preset interface; if there is no such adaptive query execution operator in the current sub-plan tree, query the configuration sheet through the preset interface to obtain the mark. If the mark indicates non-offloadable, hand over the current sub-plan tree to the central processor for execution; if the mark indicates offloadable, offload the current sub-plan tree to the data center processor for execution.

[0014] In some embodiments, before obtaining and deploying the rule plugin, the method further includes:

[0015] Configure the rule plugin for query plan tree execution management, expand Spark SQL for processing structured data based on the Spark Plugin mechanism, and allow dynamic loading or unloading of the rule plugin according to the specific requirements of the running environment.

[0016] In some embodiments, verifying all operators that make up the query plan tree based on a set standard includes:

[0017] Obtain the operator types supported by the data center processor recorded in the rule plugin. If the operator in the query plan tree belongs to the supported operator types, it is determined that the operator can be offloaded to the data center processor; if the operator in the query plan tree does not belong to the supported operator types, it is determined that the operator cannot be offloaded to the data center processor; the supported operator types include: scan, filter, projection, merge, hash aggregation, sorting, join, and data exchange;

[0018] Alternatively, select operators to be adapted to the data center processor based on the computational complexity, data transmission volume, parallel adaptability, and resource utilization rate of the operators. If the operator in the query plan tree belongs to the adapted operator types, it is determined that the operator can be offloaded to the data center processor; otherwise, it is determined that the operator cannot be offloaded to the data center processor.

[0019] In some embodiments, the method further includes:

[0020] According to the amount of data processed by the operator unloaded to the data center processor, perform dynamic data partitioning on the data center processor and adopt a streaming processing form to balance the data volume of each partition and avoid data skew.

[0021] In some embodiments, the method further includes:

[0022] Save the configuration list as a file in JSON format or YAML format in local memory, configure the expiration time according to a preset duration and update it;

[0023] The method also establishes a log file for saving the configuration lists generated at each moment in history, and establishes an index for querying and backtracking.

[0024] In some embodiments, the method further includes:

[0025] Monitor the load of the data center processor. When the load is higher than the set value, the operators in the query plan tree or sub-plan tree corresponding to the new query task are preferentially executed by the central processor.

[0026] In some embodiments, the method further includes:

[0027] When the data center processor fails, re-submit the current sub-plan tree to the central processor for operation; when the central processor fails, re-unload the current sub-plan tree to the data center processor for operation.

[0028] On the other hand, the present invention also provides an execution management device for a Spark SQL query plan tree based on a DPU, including a processor, a memory, and a computer program / instruction stored on the memory. The processor is used to execute the computer program / instruction, and when the computer program / instruction is executed, the device implements the steps of the above method.

[0029] On the other hand, the present invention also provides a computer-readable storage medium, on which a computer program / instruction is stored. When the computer program / instruction is executed by a processor, the steps of the above method are implemented.

[0030] On the other hand, the present invention also provides a computer program product, including a computer program / instruction. When the computer program / instruction is executed by a processor, the steps of the above method are implemented.

[0031] The beneficial effects of the present invention are at least:

[0032] The execution management method and device for the Spark SQL query plan tree based on DPU according to the present invention allow hybrid computing of DPU and CPU for query tasks, improving computing performance and resource utilization. When the adaptive query execution mode is not running, the query plan tree is deployed as a whole. When all operators conform to the types supported by the DPU, the operators are preferentially unloaded to the DPU for operation; otherwise, the whole is handed over to the CPU for processing. When the adaptive query execution mode is running, first determine whether the currently intercepted sub-plan tree has an adaptive query execution operator. If so, first determine whether the entire query plan tree can be unloaded and mark it in the configuration sheet. If not, query the mark in the configuration sheet and hand over the current sub-plan tree to the DPU or CPU for processing according to the mark. This execution management method deploys the whole for a single query task, avoiding the introduction of additional row-column conversion operators and the data replication overhead brought by the row-column conversion operators, saving resources and improving computing performance.

[0033] Additional advantages, objects, and features of the present invention will be partially described below and will become partially apparent to those of ordinary skill in the art after studying the following text, or may be learned from the practice of the present invention. The objects and other advantages of the present invention can be achieved and obtained through the structures specifically pointed out in the specification and the drawings.

[0034] Those skilled in the art will understand that the objects and advantages that can be achieved by the present invention are not limited to the above specifically described, and the above and other objects that the present invention can achieve will be more clearly understood according to the following detailed description. Brief Description of the Drawings

[0035] The drawings described herein are used to provide a further understanding of the present invention, form a part of this application, and do not limit the present invention. In the drawings:

[0036] Figure 1 It is a logical schematic diagram of the execution management method for the Spark SQL query plan tree based on DPU according to an embodiment of the present invention.

[0037] Figure 2 It is a schematic diagram of the query plan tree structure and allocation method in the execution management method for the Spark SQL query plan tree based on DPU according to another embodiment of the present invention. Detailed Embodiments

[0038] To make the objects, technical solutions, and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below in combination with the embodiments and the drawings. Here, the illustrative embodiments of the present invention and their descriptions are used to explain the present invention, but do not limit the present invention.

[0039] Here, it should also be noted that in order to avoid obscuring the present invention with unnecessary details, only the structures and / or processing steps closely related to the solution according to the present invention are shown in the drawings, while other details less relevant to the present invention are omitted.

[0040] It should be emphasized that the term "comprising / including" when used herein refers to the presence of features, elements, steps or components, but does not exclude the presence or addition of one or more other features, elements, steps or components.

[0041] Here, it should also be noted that if not otherwise specified, the term "connection" in this article can not only refer to a direct connection, but also represent an indirect connection with an intermediate.

[0042] It should be pre-stated that the DPU (Data Processing Unit) and the CPU (Central Processing Unit) are different in architecture and function: the CPU is a general-purpose processor suitable for complex control and computing tasks, with high flexibility and single-thread performance, while the DPU is designed specifically for data-intensive tasks, improving data processing efficiency through hardware acceleration and parallel processing, and is applicable to scenarios such as network processing, storage management, and big data analysis. The CPU is responsible for the execution of a wide range of applications, while the DPU optimizes specific data operations. The combination of the two can improve the overall system performance. The calculations on the CPU are row-based data calculations, while the calculations on the DPU are column-based data calculations.

[0043] The present invention proposes a solution for eliminating the Spark SQL row-column conversion operator in the DPU heterogeneous computing scenario. Automatically detect the offloadability of all operators in the query task. If there are non-offloadable operators, all sub-plan trees directly execute all plans on the CPU; otherwise, all sub-plan trees are offloaded to the DPU for execution. This completely avoids introducing additional row-column conversion operators and avoids data replication between the host and DPU hardware.

[0044] Specifically, the present invention provides an execution management method for a Spark SQL query plan tree based on DPU, which runs on a host device, referring to Figure 1 , and this method includes the following steps S101 to S104:

[0045] Step S101: Obtain and deploy a rule plugin, and the rule plugin is controlled to start and stop by a preset control signal.

[0046] Step S102: Perform a self-check on the adaptive query execution mode based on the rule plugin.

[0047] Step S103: If the adaptive query execution mode is in an unrun state, traverse the query plan tree and verify all operators that make up the query plan tree based on the set criteria: If there is a target operator that cannot be offloaded to the data center processor, hand over the entire query plan tree to the central processor for execution; if there is no target operator, offload the query plan tree to the data center processor for execution.

[0048] Step S104: If the adaptive query execution mode is in a running state, traverse the query plan tree and verify whether there is an adaptive query execution operator for marking a new query task in the current sub-plan tree: If there is an adaptive query execution operator in the current sub-plan tree, search for and verify the complete query plan tree, recursively determine whether there is a target operator that cannot be offloaded to the data center processor, and write the mark of whether the query plan tree can be offloaded into the configuration sheet through a preset interface; if there is no adaptive query execution operator in the current sub-plan tree, query the configuration sheet through the preset interface to obtain the mark. If the mark is non-offloadable, hand over the current sub-plan tree to the central processor for execution; if the mark is offloadable, offload the current sub-plan tree to the data center processor for execution.

[0049] In step S101, deploy the rule plugin based on the SparkPlugin mechanism. SparkPlugin allows users to insert custom logic during the life cycle of a Spark application through the plugin system. These plugins can intercept and modify the execution plan of Spark at different stages (such as task submission, task execution, plan optimization, etc.) to meet specific requirements. The plugins can be used for data source extension, plan optimization, computing acceleration, etc.

[0050] In some embodiments, before obtaining and deploying the rule plugin, the method further includes: configuring the rule plugin for query plan tree execution management, expanding Spark SQL for processing structured data based on the SparkPlugin mechanism, and allowing dynamic loading or unloading of the rule plugin according to the specific requirements of the running environment.

[0051] Specifically, deploy the rule plugin in Spark SQL through Spark Plugin as follows: Create a custom rule class that implements the Rule[SparkLogicalPlan] interface, and then register the plugin in the SparkSession. The registration can be completed by setting the spark.sql.extensions parameter in the SparkSession.builder to the fully qualified name of the plugin class, or directly adding the plugin instance to the spark.experimental.extraOptimizations list. In this way, the rules defined by the plugin will be automatically applied during the Spark SQL execution plan optimization phase.

[0052] In step S102, the Adaptive Query Execution (AQE) is a dynamic optimization mechanism in Apache Spark that can split the original query plan tree into multiple sub-plan trees for optimizing the query execution plan during query runtime. This application first needs to determine whether the Adaptive Query Execution mode is running. If not, the entire query plan tree will not be split into multiple sub-plan trees. To prevent the introduction of row-column conversion operators, the query plan tree is directly deployed to the CPU or DPU. If it is, it is necessary to uniformly coordinate the unloading of all sub-plan trees of the current query task to avoid introducing row-column conversion calculations and data replication overhead between the CPU and DPU.

[0053] In step S103, when the Adaptive Query Execution mode is not enabled, the query plan tree is directly deployed as a whole.

[0054] In some embodiments, all operators that make up the query plan tree are verified based on a set standard, including:

[0055] Obtain the operator types supported by the data center processor recorded in the rule plugin. If the operator in the query plan tree belongs to the supported operator types, it is determined that the operator can be unloaded to the data center processor; otherwise, it is a target operator that cannot be unloaded to the data center processor. The supported operator types include: scan, filter, projection, merge, hash aggregation, sorting, join, and data exchange.

[0056] Specifically, the data center processor DPU supports multiple operators through pre-configuration and records them in the rule plugin, and can support unloading the corresponding operators to the data center processor for running during actual application. Among them, those that have been supported can be unloaded to the DPU, otherwise not. The principle of the entire solution is to deploy the query plan tree as a whole and give priority to deploying it to the DPU. For the DPU, if there are operators that are not configured, they cannot be directly unloaded. Therefore, as long as there are operators that cannot be unloaded, the whole is handed over to the central processing unit CPU for processing, otherwise it is completely unloaded to the DPU for processing.

[0057] Alternatively, select an operator based on its computational complexity, data transmission volume, parallel adaptability, and resource utilization rate and adapt it to the data center processor. If the operator in the query plan tree belongs to the type of operator for which adaptation has been completed, it is determined that the operator can be offloaded to the data center processor; otherwise, it belongs to the target operator that cannot be offloaded to the data center processor. In some cases, an operator can be converted and called by a calling program loaded on the DPU, but this real-time conversion consumes computing power and also requires consideration of data transmission volume, parallel adaptability, and resource utilization rate. For operators with high computational complexity, large data transmission volume, which can be adapted to be split and parallelized, and high resource utilization rate, it is more suitable to be converted and offloaded to the DPU. If the operator itself has low computational complexity, small data transmission volume, is not easily split and parallelized, and has low resource utilization rate, then the conversion is not cost-effective, and CPU computing can be considered as a priority.

[0058] Specifically, the division of Spark SQL query plan tree operators that can be offloaded to the DPU is mainly based on criteria such as computational complexity, data transmission volume, parallel processing adaptability, resource utilization efficiency, and I / O performance requirements. By offloading computationally intensive, data-intensive, or high-I / O demand operators to the DPU, the overall execution efficiency of Spark jobs can be significantly improved.

[0059] In step S104, when the adaptive query execution mode is enabled, the query plan tree is split into multiple sub-plan trees, and at this time, more fine-grained allocation can be performed. At this time, in the AQE mode, the Spark SQL optimizer inserts an AdaptiveSparkPlanExec operator at the root node position of the original plan tree to mark a new query task, and the plan tree after this operator is divided into multiple sub-plan trees for execution. When intercepting a sub-plan tree, it is determined whether it is a new query task. If so, first make an overall judgment on whether it can be offloaded and mark it in the configuration sheet. If it is not a new query task, then determine whether it can be offloaded by querying the mark in the configuration sheet of the current query task, and determine whether the current sub-plan tree is allocated to the DPU or the CPU.

[0060] In some embodiments, the method further includes: dynamically partitioning the data center processor according to the amount of data processed by the operator offloaded to the data center processor and adopting a streaming processing form to balance the data volume of each partition and avoid data skew.

[0061] Specifically, AQE dynamically monitors the data distribution during query execution and identifies partitions with data skew. The specific practices include:

[0062] Identifying data skew. When performing operations such as shuffle or join, AQE identifies abnormally large partitions (i.e., hot partitions) by counting the sizes of each partition.

[0063] Repartitioning. For hot partitions, AQE will repartition them, reallocating data to create more and more uniform partitions, thus avoiding overloading of a single partition. For example, if a partition contains far more records than other partitions, AQE will further subdivide this partition into multiple small partitions.

[0064] Partition size adjustment. Beyond the initial partitioning scheme, AQE dynamically adjusts the partition size to balance the load. For example, by automatically adjusting the number of partitions based on the actual data size, the load of each partition is made more balanced.

[0065] In some embodiments, the method further includes: saving the configuration sheet in a file in JSON format or YAML format in local memory, configuring the expiration time according to a preset duration and updating it.

[0066] The method also establishes a log file for saving the configuration sheets generated at each moment in history, and establishes an index for querying and backtracking.

[0067] In some embodiments, the method further includes: monitoring the load of the data center processor, and when the load is higher than the set value, the operators in the query plan tree or sub - plan tree corresponding to the new query task are preferentially executed by the central processor.

[0068] In some embodiments, the method further includes: configuring a supervision and optimization strategy for the operating states of the data center processor and the central processor. When the data center processor fails, the current sub - plan tree is recomputed by the central processor; when the central processor fails, the current sub - plan tree is unloaded to the data center processor for recomputation.

[0069] On the other hand, the present invention also provides an execution management device for a Spark SQL query plan tree based on a DPU, including a processor, a memory, and a computer program / instructions stored on the memory. The processor is used to execute the computer program / instructions, and when the computer program / instructions are executed, the device implements the steps of the above - mentioned method.

[0070] On the other hand, the present invention also provides a computer - readable storage medium, on which computer program / instructions are stored, and when the computer program / instructions are executed by a processor, the steps of the above - mentioned method are implemented.

[0071] On the other hand, the present invention also provides a computer program product, including computer program / instructions, and when the computer program / instructions are executed by a processor, the steps of the above - mentioned method are implemented.

[0072] The present invention is described below with a specific embodiment:

[0073] This embodiment proposes an execution management method for the Spark SQL query plan tree based on DPU. Referring to Figure 1 , the steps are as follows:

[0074] 1. When Spark SQL optimizes the physical plan tree, it applies the plug-in rules and the plug-in starts to work.

[0075] 2. When the AQE mode is turned off, the plug-in first traverses the query plan tree. If there are non-unloadable operators, the entire plan tree is handed over to the CPU for execution; otherwise, it is all handed over to the DPU for execution.

[0076] 3. In the AQE mode, the Spark SQL optimizer inserts an AdaptiveSparkPlanExec operator at the root node position of the original plan tree. After the Spark SQL optimizer intercepts this operator, it splits the original plan tree into multiple sub-plan trees, and then optimizes and executes the sub-plan trees one by one. Each sub-plan tree applies the plug-in rules during optimization, that is, the plug-in will intercept them all.

[0077] 4. Isolation of the offloading logic between query tasks. Only one root node of the sub-plan trees in the same query task is the AdaptiveSparkPlanExec operator. Therefore, the plug-in rules determine whether it is a new query task by judging whether the intercepted plan tree root node is the AdaptiveSparkPlanExec operator.

[0078] 5. Synchronize the offloading process of each sub-plan tree of the same query task through configuration:

[0079] a) If the root node of the plan tree intercepted by the plug-in is the AdaptiveSparkPlanExec operator, it means this is a new query task. Obtain the original query plan tree through the AdaptiveSparkPlanExec operator. Then, recursively judge whether there are operators that cannot be offloaded to the DPU for the plan tree, and write the judgment result into the configuration of the SparkSession.

[0080] b) If the root node of the intercepted plan tree is not the AdaptiveSparkPlanExec operator, it means that the plan tree is a sub-plan tree of the query task being executed. At this time, it is necessary to query the configuration information in the SparkSession to judge whether to perform the offloading logic. If it cannot be unloaded, the sub-plan tree is directly handed over to the CPU for execution; otherwise, it means that the entire sub-plan tree can be unloaded to the DPU, and it is directly unloaded to the DPU in its entirety, and the DPU executes this plan tree.

[0081] Further, as Figure 2 shown, there are operators in query task 1 that cannot be offloaded to the DPU. Through this solution, all operators belonging to the current task can be automatically handed over to the CPU for execution, avoiding row-column conversion and data copying between the host and DPU hardware. All operators in query task 2 can be offloaded and will be all handed over to the DPU for execution through this solution. Through this solution, the purpose of hybrid operation of the CPU and DPU between query tasks is achieved, and the purpose of not introducing additional overhead within a single query task is also achieved.

[0082] In this application, introducing row-column conversion operators in the query plan tree of a single query task is avoided, additional calculations and data copying are avoided, and the overall execution performance is improved. The query tasks still operate in a hybrid manner. The query performance of some query tasks is accelerated through the DPU, and at the same time, the CPU resources are fully utilized. The introduction of additional row-column conversion operators and the data copying overhead brought by the row-column conversion operators are completely avoided, saving resources and improving the computing performance.

[0083] Correspondingly, the present invention also provides an apparatus / system. The apparatus / system includes a computer device, the computer device includes a processor and a memory, computer instructions are stored in the memory, and the processor is used to execute the computer instructions stored in the memory. When the computer instructions are executed by the processor, the apparatus / system implements the steps of the method described above.

[0084] The embodiment of the present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the foregoing edge computing server deployment method are implemented. The computer-readable storage medium may be a tangible storage medium, such as a random access memory (RAM), memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, register, floppy disk, hard disk, removable storage disk, CD-ROM, or any other form of storage medium known in the technical field.

[0085] In summary, for the execution management method and device of the Spark SQL query plan tree based on DPU according to the present invention, for a query task, it allows hybrid operation of DPU and CPU, improves computing performance and resource utilization rate. When the adaptive query execution mode is not running, the query plan tree is deployed as a whole. When all operators conform to the types supported by DPU, the operators are preferentially offloaded to DPU for operation, otherwise the whole is handed over to CPU for processing. When the adaptive query execution mode is running, first determine whether the currently intercepted sub-plan tree has an adaptive query execution operator. If so, first determine whether the entire query plan tree can be offloaded and marked in the configuration sheet. If not, query the mark in the configuration sheet and hand over the current sub-plan tree to DPU or CPU for processing according to the mark. This execution management method deploys the whole for a single query task, avoiding the introduction of additional row-column conversion operators and the data replication overhead brought by the row-column conversion operators, saving resources and improving computing performance.

[0086] Those of ordinary skill in the art should understand that the various exemplary components, systems, and methods described in conjunction with the embodiments disclosed herein can be implemented in hardware, software, or a combination of both. Specifically, whether to implement in hardware or software depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention. When implemented in hardware, it can be, for example, an electronic circuit, an application-specific integrated circuit (ASIC), appropriate firmware, a plug-in, a functional card, etc. When implemented in software, the elements of the present invention are programs or code segments used to perform the required tasks. The program or code segment can be stored in a machine-readable medium or transmitted via a data signal carried in a carrier wave on a transmission medium or a communication link.

[0087] It should be clear that the present invention is not limited to the specific configurations and processes described above and shown in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of the present invention is not limited to the specific steps described and shown. Those skilled in the art can make various changes, modifications, and additions, or change the order between steps after understanding the spirit of the present invention.

[0088] In the present invention, the features described and / or illustrated for one embodiment can be used in the same way or in a similar way in one or more other embodiments, and / or combined with the features of other embodiments or replace the features of other embodiments.

[0089] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, various modifications and variations can be made to the embodiments of the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for managing the execution of a Spark SQL query plan tree based on a DPU, characterized in that: The method is executed on a host device and comprises the following steps: Obtain and deploy a rule plug-in, wherein the rule plug-in is started and stopped by a preset control signal; Based on the rule plugin, perform a self-check on the adaptive query execution mode: If the adaptive query execution mode is not in operation, the query plan tree is traversed, and all operators constituting the query plan tree are checked based on the set criteria: if there is a target operator that cannot be offloaded to the data center processor, the query plan tree is completely handed over to the central processor for execution; if there is no target operator, the query plan tree is offloaded to the data center processor for execution; If the adaptive query execution mode is in operation, the query plan tree is traversed to verify whether there is an adaptive query execution operator for marking new query tasks in the current sub-plan tree: if the adaptive query execution operator exists in the current sub-plan tree, the complete query plan tree is searched and verified, and a recursive judgment is performed to determine whether the current sub-plan tree has the target operator that cannot be unloaded to the data center processor, and a mark indicating whether the query plan tree is unloadable is written into the configuration sheet through a preset interface; if the adaptive query execution operator does not exist in the current sub-plan tree, the configuration sheet is queried through the preset interface to obtain the mark, and if the mark is not unloadable, the current sub-plan tree is handed over to the central processing unit for execution; if the mark is unloadable, the current sub-plan tree is unloaded to the data center processor for execution.

2. The method for managing the execution of a Spark SQL query plan tree based on a DPU according to claim 1, characterized in that: Before acquiring and deploying the rule plug-in, the method further includes: The rule plug-in for query plan tree execution management is configured, Spark SQL for processing structured data is expanded based on the Spark Plugin mechanism, and the rule plug-in is allowed to be dynamically loaded or unloaded according to the specific requirements of the operating environment.

3. The execution management method of Spark SQL query plan tree based on DPU according to claim 1, characterized in that: All operators constituting the query plan tree are checked based on the set criteria, including: Obtain the operator types that the data center processor supports and that are recorded in the rule plug-in. If the operator in the query plan tree belongs to the operator type that has been supported for running, it is determined that the operator can be unloaded to the data center processor; if the operator in the query plan tree does not belong to the operator type that has been supported for running, it is determined that the operator cannot be unloaded to the data center processor; the operator types that have been supported for running include: scanning, filtering, projection, merging, hash aggregation, sorting, joining, and data exchange; Alternatively, an operator is selected to adapt to the data center processor based on the operator's computational complexity, data transmission volume, parallel adaptability, and resource utilization; if the operator in the query plan tree belongs to an operator type that has completed adaptation, it is determined that the operator can be offloaded to the data center processor; otherwise, it is determined that the operator cannot be offloaded to the data center processor.

4. The method for managing the execution of a Spark SQL query plan tree based on DPU according to claim 1, characterized in that: The method further comprises: According to the amount of data processed by the operator unloaded to the data center processor, the data center processor is dynamically partitioned and stream processing is adopted to balance the data amount of each partition to avoid data skew.

5. The method for managing the execution of a Spark SQL query plan tree based on DPU according to claim 1, characterized in that: The method further comprises: The configuration list is saved in a local memory in a JSON format or a YAML format file, and the expiration time is configured and updated according to a preset time length; The method also creates a log file for storing the configuration orders generated at various moments in history, and creates an index for query and backtracking.

6. The method for managing the execution of a Spark SQL query plan tree based on DPU according to claim 1, characterized in that: The method further comprises: The load of the data center processor is monitored. When the load is higher than a set value, the operators in the query plan tree or sub-plan tree corresponding to the new query task are preferentially handed over to the central processor for execution.

7. The execution management method of Spark SQL query plan tree based on DPU according to claim 1, characterized in that: The method further comprises: When the data center processor fails, the current sub-plan tree is handed over to the central processor for calculation again; when the central processor fails, the current sub-plan tree is unloaded to the data center processor for calculation again.

8. A DPU-based Spark SQL query plan tree execution management device, comprising a processor, a memory, and a computer program / instruction stored in the memory, characterized in that: The processor is used to execute the computer program / instructions. When the computer program / instructions are executed, the device implements the steps of the method according to any one of claims 1 to 7.

9. A computer-readable storage medium having a computer program / instruction stored thereon, characterized in that: When the computer program / instructions are executed by a processor, the steps of the method as claimed in any one of claims 1 to 7 are implemented.

10. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Data processing method and system under storage and calculation separation architecture based on DPU, and storage medium

    CN117370020A

  • Multi-operator fusion data processing method and system based on DPU

    CN117874053A