A data processing method and system based on DPU multi-operator fusion

By adopting the DPU-based multi-operator fusion method in Spark SQL, the problem of CPU performance bottleneck in the existing technology is solved, more efficient data processing is achieved, and the execution speed and throughput of Spark SQL are improved.

CN117874053BActive Publication Date: 2025-06-06YUSUR TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311649240.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-12-04
Publication Date
2025-06-06
Estimated Expiration
2043-12-04

AI Technical Summary

Technical Problem

When existing Spark SQL processes large-scale data and intensive computing, it is unable to take full advantage of the hardware accelerator due to its CPU reliance on performance bottlenecks.

Method used

The multi-operator fusion method based on DPU is adopted to generate a physical plan tree through a semantic analyzer, replace the operator node as a DPU operator node, and use the marking algorithm to determine the fusion node, generate fusion code, and perform multi-operator fusion to reduce the read and write data and the storage of intermediate results.

Benefits of technology

Offload data-intensive computing from the CPU to the DPU, reduce data read and write and storage of intermediate results through operator fusion, reduce calculation delay, improve throughput, improve Spark SQL execution speed, and reduce CPU resource consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117874053B_ABST
    Figure CN117874053B_ABST
Patent Text Reader

Abstract

The present invention provides a data processing method and system based on DPU multi-operator fusion, the method comprising: receiving a data processing task and generating a physical plan tree using a semantic analyzer. The operator nodes in the physical plan tree are replaced with DPU operator nodes. The DPU operator nodes are traversed, and identifiers are set for these nodes using a marking algorithm to determine the fused nodes. Each fused operator node generates a corresponding fusion code, and the fusion execution module uses a marking algorithm to perform multi-operator fusion. The data set to be processed is read and converted into an appropriate format. The data will be passed to the fusion execution module for batch cyclic processing to perform multi-operator fusion operations. After all the data sets to be processed are processed, the results are converted into an appropriate format, and the results of the data processing are output. The present invention unloads data-intensive calculations from the CPU to the DPU, merges multiple operators, releases CPU resources, and improves data processing speed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of distributed computing technology, and in particular to a data processing method and system based on DPU multi-operator fusion. Background Art

[0002] Apache Spark is a fast general-purpose computing engine designed for large-scale data processing. It provides an open source cluster computing environment similar to Hadoop. Spark SQL is one of Spark's computing modules, which is specifically used to process structured data. Spark SQL allows users to use standard SQL statements to perform SQL queries and read and write operations. It also supports the use of Hive SQL to query and read and write Hive warehouses. This allows users to easily use the familiar SQL language for data processing without having to write complex code in depth.

[0003] However, the current Spark SQL still relies on CPU for computing. Although the Spark framework can schedule computing, the CPU's computing power becomes the main bottleneck of performance when processing large-scale data and intensive computing. As the CPU is a general-purpose processing chip, it has no obvious advantage in high-density computing of big data. Summary of the invention

[0004] In view of this, an embodiment of the present invention provides a data processing method and system based on DPU multi-operator fusion to eliminate or improve one or more defects existing in the prior art.

[0005] One aspect of the present invention provides a data processing method based on DPU multi-operator fusion, the method comprising the following steps:

[0006] Receive a data processing task, generate a physical plan tree through a semantic analyzer, and replace operator nodes in the physical plan tree with DPU operator nodes;

[0007] Traversing the DPU operator nodes, and setting identifiers for the DPU operator nodes using a marking algorithm to mark whether the DPU operator nodes can be fused;

[0008] A corresponding fusion code is generated for each fusionable operator node, and the fusion execution module uses a marking algorithm to perform multi-operator fusion.

[0009] In some embodiments of the present invention, the semantic analyzer is a Spark SQL semantic analyzer.

[0010] In some embodiments of the present invention, traversing the DPU operator nodes and setting identifiers for the DPU operator nodes using a marking algorithm to mark whether the DPU operator nodes can be fused specifically include:

[0011] The DPU operator nodes are traversed, and the input dependency of the current node is checked using the marking algorithm. If the input of the current node only comes from the output of the previous node and does not undergo a data repartitioning operation, the node is marked as the fusible operator.

[0012] In some embodiments of the present invention, the method is performed under a Producer-Consumer framework.

[0013] In some embodiments of the present invention, it further includes:

[0014] Reading a data set to be processed, converting the data set to be processed into data in a first form, and then transferring the data set to the fusion execution module for batch loop processing to execute multi-operator fusion;

[0015] After the entire data set to be processed is processed and converted into the second form of data, the data processing result is output.

[0016] In some embodiments of the present invention, the form of reading the data set to be processed includes: table scanning, repartition reading or broadcasting.

[0017] In some embodiments of the present invention, the first form of data is a columnar data structure.

[0018] In some embodiments of the present invention, the second form of data is a data structure suitable for Spark.

[0019] Another aspect of the present invention provides a data processing system based on DPU multi-operator fusion, including a DPU, a processor and a memory, wherein the memory stores computer instructions, and the processor is used to execute the computer instructions stored in the memory. When the computer instructions are executed by the processor, the system implements the steps of the above method.

[0020] Another aspect of the present invention provides a computer-readable storage medium having a computer program stored thereon, which implements the steps of the above method when executed by a processor.

[0021] The beneficial effects of the present invention are at least:

[0022] The present invention provides a data processing method and system based on DPU multi-operator fusion, the method comprising: receiving a data processing task and generating a physical plan tree using a semantic analyzer. Replacing the operator nodes in the physical plan tree with DPU operator nodes. Traversing the DPU operator nodes, and using a marking algorithm to set identifiers for these nodes to determine the fusionable nodes. Each fusionable operator node generates a corresponding fusion code, and the fusion execution module uses the marking algorithm to perform multi-operator fusion. Read the data set to be processed and convert it into an appropriate format. The data will be passed to the fusion execution module for batch cyclic processing to perform multi-operator fusion operations. After all the data sets to be processed are processed, the results are converted into an appropriate format, and the results of the data processing are output. The present invention unloads data intensive calculations from the CPU to the DPU, uses a marking algorithm at the DPU layer, merges multiple operators, avoids the storage and transmission of intermediate results, reduces calculation delays, and improves throughput. Operator fusion reduces the reading and writing of data and the storage of intermediate results, improves the execution speed of Spark SQL, can reduce CPU resource consumption, and improves data processing speed.

[0023] Additional advantages, purposes, and features of the present invention will be described in part in the following description, and will become apparent to those skilled in the art after studying the following, or may be learned from the practice of the present invention. The purposes and other advantages of the present invention may be achieved and obtained by the structures specifically indicated in the specification and the accompanying drawings.

[0024] Those skilled in the art will appreciate that the objectives and advantages that can be achieved with the present invention are not limited to the above specific description, and the above and other objectives that can be achieved by the present invention will be more clearly understood from the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] The drawings described herein are used to provide a further understanding of the present invention, constitute a part of the present application, and do not constitute a limitation of the present invention. In the drawings:

[0026] Figure 1 The present invention is a flowchart of a data processing method based on DPU multi-operator fusion according to an embodiment of the present invention.

[0027] Figure 2 This is a processing flow chart of the fusible operators according to another embodiment of the present invention.

[0028] Figure 3 This is a Spark fusion input-output data flow diagram according to another embodiment of the present invention.

[0029] Figure 4 This is a flowchart of the execution of the fusion operator according to another embodiment of the present invention. DETAILED DESCRIPTION

[0030] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the embodiments and the accompanying drawings. Here, the illustrative embodiments of the present invention and their descriptions are used to explain the present invention, but are not intended to limit the present invention.

[0031] It should also be noted that, in order to avoid obscuring the present invention due to unnecessary details, only structures and / or processing steps closely related to the solutions according to the present invention are shown in the accompanying drawings, while other details that are not closely related to the present invention are omitted.

[0032] It should be emphasized that the term “include / comprises” when used herein refers to the presence of features, elements, steps or components, but does not exclude the presence or addition of one or more other features, elements, steps or components.

[0033] Hereinafter, embodiments of the present invention will be described with reference to the accompanying drawings. In the accompanying drawings, the same reference numerals represent the same or similar components, or the same or similar steps.

[0034] DPU is the abbreviation of Data Processing Unit, which is a hardware accelerator specially designed for processing data, designed to speed up data processing in data centers and edge computing devices. DPU usually provides efficient computing on specific types of data processing tasks, can efficiently process large-scale data streams, and accelerate data transmission, processing and storage. DPU integrates network functions, security functions and data processing functions, can process and manage data streams, and provide higher performance and efficiency.

[0035] Spark is an open source distributed computing framework that provides efficient and easy-to-use data processing and analysis tools, supporting fast data processing, machine learning, and graph computing on large-scale clusters. SQL (Structured Query Language) is a standardized query language for managing and operating relational database management systems (RDBMS). It provides a simple and powerful way for users to define, manipulate, and query data in a database. The functions of SQL mainly include data query, data manipulation, data definition, and data control.

[0036] An aspect of an embodiment of the present invention provides a data processing method based on DPU multi-operator fusion, the method comprising the following steps S101 to S103:

[0037] Step S101: receiving a data processing task, generating a physical plan tree through a semantic analyzer, and replacing the operator nodes in the physical plan tree with DPU operator nodes.

[0038] Among them, the physical plan tree is a data structure used in the process of data processing or query optimization. It represents the actual physical operation process and describes how to generate the required output results from the input data. The physical plan tree usually contains nodes, connections, execution order, and optimization strategies.

[0039] In some embodiments, the physical plan tree may include multiple subtrees, where a subtree refers to a substructure in the physical plan tree, which can be an independent execution plan unit or a combination of other subtrees. When analyzing the physical plan tree, the subtrees in the physical plan tree are traversed to determine whether there are operator nodes in each subtree that cannot be replaced. If so, all operator nodes in the subtree will not be replaced; if not, all operator nodes in the subtree can be replaced. Therefore, the operator nodes in the physical plan tree will be partially or completely replaced with DPU operator nodes.

[0040] Step S102: traverse the DPU operator nodes, and use a marking algorithm to set identifiers for the DPU operator nodes to mark whether the DPU operator nodes can be merged.

[0041] Step S103: Generate corresponding fusion code for each fusionable operator node, and the fusion execution module uses a marking algorithm to perform multi-operator fusion.

[0042] Operator fusion is the process of combining multiple consecutive operator operations into a single operation to improve computing efficiency. Operator fusion can reduce the overhead of reading, writing, and communicating data, and reduce the memory required to calculate intermediate results. In Spark, when multiple operator operations are applied continuously to RDD, Spark will fuse them into a more efficient sequence of operations to reduce the storage of intermediate results and improve computing performance.

[0043] In some embodiments of the present invention, the semantic analyzer is a Spark SQL semantic analyzer.

[0044] In some embodiments of the present invention, traversing the DPU operator nodes and using a marking algorithm to set an identifier for the DPU operator nodes to mark whether the DPU operator nodes can be merged specifically includes:

[0045] Traverse the DPU operator nodes and use the marking algorithm to check the input dependency of the current node. If the input of the current node only comes from the output of the previous node and does not undergo data repartitioning, the node is marked as a fusionable operator.

[0046] In some embodiments of the present invention, the method is performed under a Producer-Consumer framework.

[0047] Among them, the Producer-Consumer framework is a concurrent programming model used to solve the data exchange and synchronization problems between producers and consumers. In this framework, producers are responsible for generating data and putting it into a shared buffer, while consumers are responsible for taking data out of the buffer and processing it. By using this framework, the work between producers and consumers can be effectively coordinated to avoid common concurrent programming problems such as data competition and deadlock.

[0048] In some embodiments of the present invention, it further includes:

[0049] The data set to be processed is read, converted into data in the first form, and then transmitted to the fusion execution module for batch loop processing to execute multi-operator fusion.

[0050] After the entire data set to be processed is processed and converted into the second form of data, the data processing result is output.

[0051] In some embodiments of the present invention, the form of reading the data set to be processed includes: table scanning, repartition reading or broadcasting.

[0052] In some embodiments of the present invention, the first form of data is a columnar data structure.

[0053] In some embodiments of the present invention, the second form of data is a data structure suitable for Spark.

[0054] Another aspect of an embodiment of the present invention provides a data processing system based on DPU multi-operator fusion, including a DPU, a processor and a memory, wherein computer instructions are stored in the memory, and the processor is used to execute the computer instructions stored in the memory. When the computer instructions are executed by the processor, the system implements the steps of the above method.

[0055] Another aspect of an embodiment of the present invention provides a computer-readable storage medium having a computer program stored thereon, which implements the steps of the above method when executed by a processor.

[0056] Another embodiment of the present invention provides a data processing method and system based on DPU multi-operator fusion, which uses a DPU based on KPU architecture to accelerate multi-operator fusion in Spark SQL. The specific implementation scheme of the method includes:

[0057] 1. Whether the marking operator can be fused:

[0058] like Figure 2As shown, first, a statement is processed by the Spark SQL semantic analyzer to generate an executable physical plan tree. After being processed by this system, the operator nodes in the CPU physical plan tree are replaced by the operator nodes for DPU operations. Then, the plan is traversed from top to bottom to set marks for the fusible operators.

[0059] 2. Generate fusion code:

[0060] 1) The Producer-Consumer framework is used at the software layer. Each operator is responsible for generating its own fusion code, which is finally merged by TaskStageOperatorExec. In the final implementation, a large loop is used at the outer layer. Each operator implements part of the fusion code, and then batches the data while maintaining the original operator pipeline calculation model.

[0061] 2) Use a marking algorithm at the hardware layer to fuse multiple operators. Multiple operators notify the board whether to fuse and the hardware storage location of the data based on the identifier and address pointer.

[0062] Specifically, the embodiments of the present invention optimize from the perspective of reducing IO and CPU memory. The fused operator copies data from the CPU memory to the DPU board. The DPU board caches the data in the DPU board register according to the operator identifier for use by the next operator, which can reduce the number of data copies and reduce CPU consumption here. A marking algorithm is used at the DPU layer to clarify the frequency of data flow use, avoid the storage and transmission of intermediate results, and thus enable the DPU board to reduce calculation delays and improve calculation throughput.

[0063] Specifically, TaskStageOperatorExec represents a single operation in an execution stage, such as an operator or a specific transformation operation. It contains the logical execution plan and physical execution plan of the operation, as well as some context information and data structures required to execute the operation. TaskStageOperatorExec encapsulates the specific operations in the execution stage in Spark, and provides the Spark execution engine with the relevant information and methods required to execute these operations. Through TaskStageOperatorExec, Spark can manage, schedule, and execute operations in each execution stage, thereby realizing distributed parallel processing of the entire computing process.

[0064] 3. Spark SQL integrated use:

[0065] Based on the ColumnarBatch method, data is transmitted to the downstream in batches by storing each column in a vector format. The ColumnarBatch method is an internal data structure for efficient processing of columnar data. Spark fusion input and output data flow is as follows Figure 3 As shown, specifically including:

[0066] For input, according to the Spark operating mechanism, the data input provided by the upstream is Table Scan, Shuffle Read, and Broad Cast. There are actually only two types of operator fusion algorithms: Columnar Batch or Columnar Iterator.

[0067] For output, after operator fusion, the fixed value is Columnar Iterator. The corresponding output data of Spark is Shuffle write and Input Adapter.

[0068] Furthermore, the fusion operator execution process is as follows Figure 4 As shown, first determine whether the operator is a fusion operator or a non-fusion operator based on the current tag. In the fusion operator branch, first execute TaskInit to perform task-level initialization work, including initializing the operator fusion algorithm, loading the compiled fusion algorithm module, reading Table Scan data, and other operations. Then execute a loop, each loop reads a ColumnarBatch (column batch) for the fusion operator to execute, and finally converts the output to Shuffle Writer or Input Adapter.

[0069] The input of ColumnarBatch is Scan Table, Shuffle Reader, or BroadCast.

[0070] Among them, the embodiment of the present invention adopts the Producer-Consumer framework at the Jvm layer, processes data in column batches, processes one column at a time, breaks the boundaries between internal operators, generates fused operators that are consistent with the source logic, reduces the overhead of virtual function calls, eliminates unnecessary object creation, and improves the execution speed of Spark SQL.

[0071] The Producer-Consumer framework is a concurrent programming model that solves the problem of data delivery and coordination between producers and consumers. In this model, the producer is responsible for generating data or tasks and passing them to the consumer for processing. The basic idea is that the producer and consumer communicate through a shared data buffer. The producer puts the data into the buffer, and the consumer gets the data from the buffer and processes it. This model can effectively decouple producers and consumers, allowing them to work at different speeds without blocking each other. In software development, the Producer-Consumer framework is often used to solve data sharing and communication problems between multiple threads or processes. It can help implement scenarios such as task distribution and data processing in concurrent programming, and improve the efficiency and performance of the system.

[0072] In summary, the present invention provides a data processing method and system based on DPU multi-operator fusion, the method comprising: receiving a data processing task and generating a physical plan tree using a semantic analyzer. Replace the operator nodes in the physical plan tree with DPU operator nodes. Traverse the DPU operator nodes, and use a marking algorithm to set identifiers for these nodes to determine the fused nodes. Each fused operator node generates a corresponding fusion code, and the fusion execution module uses a marking algorithm to perform multi-operator fusion. Read the data set to be processed and convert it into an appropriate format. The data will be passed to the fusion execution module for batch loop processing to perform multi-operator fusion operations. After all the data sets to be processed are processed, the results are converted into an appropriate format, and the results of the data processing are output. The present invention offloads data-intensive calculations from the CPU to the DPU, merges multiple operators, releases CPU resources, and improves data processing speed.

[0073] Corresponding to the above method, the present invention also provides a system, which includes a computer device, the computer device includes a processor and a memory, the memory stores computer instructions, the processor is used to execute the computer instructions stored in the memory, and when the computer instructions are executed by the processor, the system implements the steps of the method described above.

[0074] The embodiment of the present invention also provides a computer-readable storage medium on which a computer program is stored, and when the computer program is executed by a processor, the steps of the aforementioned edge computing server deployment method are implemented. The computer-readable storage medium can be a tangible storage medium, such as a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a floppy disk, a hard disk, a removable storage disk, a CD-ROM, or any other form of storage medium known in the technical field.

[0075] It should be understood by those skilled in the art that the exemplary components, systems and methods described in conjunction with the embodiments disclosed herein can be implemented in hardware, software or a combination of the two. Whether it is performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention. When implemented in hardware, it can be, for example, an electronic circuit, an application specific integrated circuit (ASIC), appropriate firmware, a plug-in, a function card, etc. When implemented in software, the elements of the present invention are programs or code segments used to perform the required tasks. The program or code segment can be stored in a machine-readable medium, or transmitted on a transmission medium or a communication link via a data signal carried in a carrier.

[0076] It should be clear that the present invention is not limited to the specific configuration and processing described above and shown in the figures. For the sake of simplicity, a detailed description of the known method is omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of the present invention is not limited to the specific steps described and shown, and those skilled in the art can make various changes, modifications and additions, or change the order between the steps after understanding the spirit of the present invention.

[0077] In the present invention, features described and / or illustrated for one embodiment may be used in the same or similar manner in one or more other embodiments, and / or combined with features of other embodiments or replace features of other embodiments.

[0078] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the embodiments of the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A data processing method based on multi-operator fusion of DPU, It is characterized in that The method comprises the following steps: Receive a data processing task, generate a physical plan tree through a semantic analyzer, and replace operator nodes in the physical plan tree with DPU operator nodes; Traverse the DPU operator nodes and use the marking algorithm to check the input dependency of the current node. If the input of the current node only comes from the output of the previous node and does not undergo data repartitioning, mark the node as a fusionable operator. Generate corresponding fusion code for each fusionable operator node, and the fusion execution module uses the marking algorithm to perform multi-operator fusion; Reading a data set to be processed, converting the data set to be processed into data in a first form, and then transmitting the data set to the fusion execution module for batch loop processing to execute multi-operator fusion; Until the entire data set to be processed is completely processed and converted into the second form of data, the data processing result is output; The first form of data is a columnar data structure; the second form of data is a data structure suitable for Spark.

2. The data processing method based on DPU multi-operator fusion according to claim 1, It is characterized in that The semantic analyzer is a Spark SQL semantic analyzer.

3. The data processing method based on DPU multi-operator fusion according to claim 1, It is characterized in that, The method is implemented in a Producer-Consumer framework.

4. The data processing method based on DPU multi-operator fusion according to claim 1, It is characterized in that The form of reading the data set to be processed includes: table scanning, repartition reading or broadcasting.

5. A data processing system based on DPU multi-operator fusion, comprising a DPU, a processor and a memory, It is characterized in that The memory stores computer instructions, and the processor is used to execute the computer instructions stored in the memory. When the computer instructions are executed by the processor, the system implements the steps of the method as claimed in any one of claims 1 to 4.

6. A computer-readable storage medium having a computer program stored thereon, It is characterized in that When the program is executed by a processor, the steps of the method according to any one of claims 1 to 4 are implemented.

Citation Information

Patent Citations

  • Operator calculation method and device, equipment and medium

    CN116149856A

  • Deep learning inference task compiler-oriented operator fusion method and system

    CN116861359A