A Compilation-Based Automatic Multi-Stream Scheduling Method for Kernel Functions

The kernel function dependencies are automatically analyzed and DAG is constructed through the LLVM compiler framework. The topological sorting and stream scheduling algorithms are used to allocate kernel functions to different streams for parallel execution, solving the problems of complexity and low performance of parallel development in the existing technology, and achieving efficient kernel function parallelization and GPU resource utilization.

CN114549277BActive Publication Date: 2025-06-24SUN YAT SEN UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210172808.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-02-24
Publication Date
2025-06-24
Estimated Expiration
2042-02-24

AI Technical Summary

Technical Problem

When the prior art parallelizes serial kernel functions, developers need to learn in-depth CUDA framework and task-oriented programming framework, which increases development costs and learning burdens, and introduces additional runtime overhead and reduces program performance.

Method used

The data dependencies between kernel functions are automatically analyzed through the LLVM compiler framework, directed acyclic graph (DAG), and use topological sorting and stream scheduling algorithms to allocate kernel functions to different streams for parallel execution, avoiding additional programming framework and runtime overhead.

Benefits of technology

It realizes parallelizing serial kernel functions without the need for additional programming model and runtime overhead, improving program performance and GPU hardware utilization, and reducing developers' learning and development costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114549277B_ABST
    Figure CN114549277B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for automatically multi-stream scheduling of kernel functions based on compilation, including: reading source code, identifying the inputs and outputs of kernel functions; constructing a DAG according to the inputs and outputs; hierarchizing the DAG, and assigning appropriate streams to the kernel functions according to topological sorting; inserting event synchronization code for kernel functions with cross-stream synchronization; generating an executable file. The present invention is the first to parallelize serial GPU kernel functions through the multi-stream mechanism of the CUDA framework by means of the LLVM compiler framework, thereby improving program performance and GPU hardware utilization. The invention does not introduce additional programming models and runtime overheads, reducing the learning cost and development cost of developers. At the same time, the stream scheduling algorithm used in the invention arranges the running times of kernel functions with data dependency relationships close to each other, maximizing the memory access efficiency of the GPU for data, and thus improving program performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular, to an automatic multi-stream scheduling method for kernel functions based on compilation. Background Art

[0002] With the improvement of GPU computing power, AI applications with increasing computational volume and data throughput are more and more widely applied on GPUs. The basic computing unit of an AI application is an operator, and multiple operators are organized in a specific network structure to form a complete AI model. Therefore, the performance of an AI model depends on two parts. One is the computing performance within an operator, and the other is the resource occupancy rate, memory access efficiency, etc. determined by the network structure between operators. Since the computation and data reading / writing of a single operator in actual applications are difficult to occupy all GPU resources, if multiple operators can be made to run simultaneously on the GPU at a single moment, then more GPU resources can be used to improve program performance. To achieve the parallel effect of kernel functions (i.e., operators), developers can use the stream and event mechanisms in the CUDA framework.

[0003] The GPU supports users to use multiple streams simultaneously, and the tasks between different streams are executed asynchronously, and these tasks may be executed simultaneously at the same moment. Therefore, to ensure the ordered execution of tasks with dependencies between different streams, events can be used for cross-stream synchronization.

[0004] However, this requires users to have knowledge of streams and events in the CUDA framework for deep learning and certain heterogeneous program development experience. At the same time, it also requires users to correctly analyze the data dependencies between possibly hundreds of kernel functions in a complex program. The data dependencies between kernel functions can be divided into read-after-write, write-after-write, read-after-write, and read-after-read. Among them, read-after-write means that the variable read by the current kernel function will be modified by the previous kernel function; write-after-write means that the variable modified by the current kernel function will be modified by the previous kernel function; read-after-write means that the variable modified by the current kernel function will be read by the previous kernel function; read-after-read means that the variable read by the current kernel function will be read by the previous kernel function.

[0005] There cannot be data dependencies between kernel functions parallel at the same moment, otherwise an error of incorrect final result will occur. Stripping kernel functions out for parallelization in a complex program requires developers to pay a relatively high development cost to modify to obtain correct kernel function parallel code. These thresholds and problems have all hindered the popularization of GPUs, especially the transplantation of traditional HPC applications originally implemented on CPUs to GPUs.

[0006] To simplify the development of GPU programs and improve their performance, task graph computing systems have become the preferred choice. The task-oriented programming framework included therein abstracts each kernel function into a task and abstracts the data dependency relationship between tasks into a directed acyclic graph (DAG). Then, topological sorting is used on this graph to arrange the kernel functions without dependencies in different GPU streams to achieve parallelism. For kernel functions with dependencies, if these kernel functions are placed in different streams, event synchronization needs to be added between the dependent functions to ensure the ordered execution of the kernel functions in different streams. On the other hand, the runtime software included therein schedules kernel functions according to the real-time GPU hardware resource usage situation, determining when and on which GPU or which CPU they run, etc.

[0007] In recent years, many new task graph computing systems have emerged, but the existing technologies have the following disadvantages:

[0008] 1. Developers need to learn a new task-oriented programming framework, increasing the learning burden on developers.

[0009] 2. When developers modify serial kernel functions into multi-stream parallel kernel functions, they need to significantly adjust the existing code framework, increasing the development cost.

[0010] 3. The existing technologies introduce the real-time scheduling of kernel functions by runtime software, introducing additional runtime overhead and reducing program performance. Summary of the Invention

[0011] In view of the above defects of the existing technologies, the technical problem to be solved by the present invention is to provide a compilation-based automatic multi-stream scheduling method for kernel functions. Users only need to write serial kernel functions in C++ according to the basic functional logic, and the present invention can automatically analyze the data dependencies between kernel functions with the help of a compiler, construct the kernel functions into a DAG, and allocate the kernel functions to different streams for parallelism, thereby improving performance. This method does not introduce additional programming frameworks and runtime overhead, and maximally reduces the development cost for developers to develop and use high-performance heterogeneous programs.

[0012] To achieve the above object, the present invention provides a compilation-based automatic multi-stream scheduling method for kernel functions, including the following steps:

[0013] Step 1, read the source code and identify the inputs and outputs of the kernel functions;

[0014] Step 2, construct a DAG according to the inputs and outputs;

[0015] Step 3, hierarchize the DAG and assign appropriate streams to the kernel functions according to topological sorting;

[0016] Step 4, insert event synchronization code for kernel functions with cross-stream synchronization;

[0017] Step 5, generate an executable file.

[0018] Further, the reading of the source code and identifying the input and output of the kernel function are specifically as follows: The source code is transformed into an IR file through the LLVM compilation framework, and then the opt tool is used to process the IR file to analyze the read-only variables and write variables of each kernel function.

[0019] Further, the construction of the DAG according to the input and output is specifically as follows: Assume that the current kernel function has N incoming parameters. For each incoming parameter, take the end node of the DAG as the starting point, and use the reverse BFS algorithm to find the first kernel function that has the same incoming parameter, add this kernel function as the predecessor node of the current kernel function, and set the node represented by the current kernel function as the predecessor node of the end node in the DAG.

[0020] Further, the appropriate stream is assigned to each kernel function at each level in the DAG in a hierarchical order, specifically as follows: An increasing id is assigned to the task nodes in each layer in sequence, so that each task node has a unique id in the current layer; First, calculate the assigned stream for those task nodes that have no predecessor nodes; Then, traverse the task nodes in the current layer in ascending order according to the number of predecessor nodes, and perform stream assignment according to the following strategy:

[0021] If a certain predecessor node of the current node is the end node of its current stream, then assign the current node to this stream;

[0022] If there are multiple predecessor nodes that meet the above conditions, then select the stream of the predecessor node with the fewest uncompleted successor nodes;

[0023] If there is no such predecessor node, use the stream calculated by the current task node.

[0024] Further, inserting event synchronization code for kernel functions with cross-stream synchronization is specifically as follows:

[0025] For all nodes whose predecessor nodes are not in the same stream as itself, in each stream that has its predecessor node, select the last predecessor node in this stream, and add code with the event trigger in the CUDA / HIP programming framework and the specified task stream in the CUDA / HIP programming framework as parameters after the kernel function call code;

[0026] Then, for the task nodes that have a dependency relationship with this predecessor node and are not in the same stream, add code with the event trigger in the CUDA / HIP programming framework and the specified task stream in the CUDA / HIP programming framework as parameters before the statement that calls the kernel function to wait for the execution of the predecessor function in different streams to complete before continuing to execute this function.

[0027] Further, the generation of the executable file specifically converts the intermediate language representation file into an executable file through LLVM.

[0028] The beneficial effects of the present invention are as follows:

[0029] The present invention is the first to parallelize serial GPU kernel functions through the multi-stream mechanism of the CUDA framework using the LLVM compiler framework, thereby improving program performance and GPU hardware utilization. This invention does not introduce additional programming models and runtime overheads, reducing the learning cost and development cost for developers. At the same time, the stream scheduling algorithm used in this invention arranges the execution times of kernel functions with data dependencies close to each other, maximizing the GPU's memory access efficiency for data and thus improving program performance.

[0030] The following will further illustrate the concept, specific structure, and technical effects of the present invention in conjunction with the accompanying drawings to fully understand the purpose, features, and effects of the present invention. Description of the Drawings

[0031] Figure 1 is the overall optimization process flowchart of the present invention.

[0032] Figure 2 is the code screenshot of the depth-first search of kernel functions using the use-definition chain of the present invention.

[0033] Figure 3 is the DAG flowchart constructed according to the input and output of the present invention.

[0034] Figure 4 is the flowchart of the breadth-first reverse traversal algorithm of the present invention.

[0035] Figure 5 is the code screenshot of the algorithm for assigning appropriate streams to each kernel function at each level of the DAG in hierarchical order of the present invention. Detailed Embodiments

[0036] The definitions of relevant variables or terms of the present invention are as follows:

[0037] Relevant variables: For any variable, the other variables that appear in the same code statement with it. This includes the return variables when the variable is an operand and all operand variables when the variable is a return variable.

[0038] Directly relevant variables: The relevant variables that appear in the same code statement as the target variable.

[0039] Indirectly relevant variables: The relevant variables that do not appear in the same code statement as the target variable but are connected through a graph formed by multiple directly relevant variables.

[0040] Definition - Use Chain: For a target variable, a linked list of statements that all use the defined value as an operand.

[0041] Use - Definition Chain: For a target variable, a linked list of statements that define the value.

[0042] Directed Acyclic Graph (DAG): A directed graph is a DAG if it is not possible to start from any vertex and return to that point after passing through several edges. Nodes represent a function or a task, and the edges between nodes represent dependencies.

[0043] Breadth - First Search (BFS): For a directed acyclic graph, at each step, starting from the current node, obtain all other nodes that can be reached by passing through only one edge, and use these nodes as the starting points for the next step.

[0044] Topological Sorting: In a directed acyclic graph, for the nodes that currently have no predecessors, use BFS to obtain the reachable nodes, then delete the edges between the node and the reachable nodes, and repeat the above process for those nodes without predecessors. Finally, obtain a task execution order queue that follows the task dependency relationship.

[0045] Hierarchicalization: For a directed acyclic graph, perform topological sorting. All nodes that can be reached at each step are nodes at the same level, thereby dividing the DAG into multiple levels.

[0046] CUDA: A heterogeneous programming framework proposed by Nvidia that allows users to write GPU computing functions, offload computing tasks to the GPU, and code for communication between the CPU and the GPU.

[0047] HIP: A heterogeneous programming framework proposed by AMD, similar in function to CUDA and compatible with CUDA.

[0048] As Figure 1 shown, the present invention provides a method for automatically multi - stream scheduling of kernel functions based on compilation, including the following steps:

[0049] Step 1: Read the source code and identify the input and output of the kernel function;

[0050] Step 2: Construct a DAG based on the input and output;

[0051] Step 3: Hierarchicalize the DAG and assign appropriate streams to the kernel function according to topological sorting;

[0052] Step 4: Insert event synchronization code for kernel functions with cross - stream synchronization;

[0053] Step 5: Generate an executable file.

[0054] In this embodiment, the source code is read, and the inputs and outputs of the kernel functions are identified. Specifically, the source code is transformed into an IR file through the LLVM compilation framework, and then the opt tool is used to process the IR file to analyze the read-only variables and write variables of each kernel function. LLVM is a set of compiler and toolchain technologies that can be used to develop the front-end of any programming language and the back-end of any instruction set architecture. LLVM is designed around a language-independent intermediate representation (IR), which serves as a portable high-level assembly language and can be optimized through various transformations in multiple passes. Opt is a tool in the LLVM toolchain responsible for optimizing the intermediate representation (IR), and the specific optimization logic is implemented by the user.

[0055] In the definition part of each kernel function, for each source variable in the function signature, perform a depth-first search as shown in Figure 2 to find all its directly / indirectly related variables according to the definition-use chain and use-definition chain in LLVM.

[0056] This algorithm traverses the use statements of each current search target variable. These statements can be ordinary assignment statements or function call statements. For an assignment statement, the variable defined by the statement is used as the new search target variable for depth-first search. For a function call statement, a depth-first search is performed on the incoming parameters corresponding to the current search target variable inside the called function. When a search target variable involved by a source variable serves as the target variable of a certain code statement (such as a data write statement) and the statement is not the definition statement of the related variable, then the source variable is regarded as a write variable of the kernel function; otherwise, it is a read variable.

[0057] In this embodiment, a DAG is constructed based on the inputs and outputs, as shown in Figure 3 Specifically, assume that the current kernel function has N incoming parameters. For each incoming parameter, starting from the end node of the DAG, use the reverse BFS algorithm to find the first kernel function that has the same incoming parameter, add this kernel function as the predecessor node of the current kernel function, and set the node represented by the current kernel function as the predecessor node of the end node in the DAG. The reverse breadth-first traversal algorithm is as shown in Figure 4 shown.

[0058] In this embodiment, each kernel function at each level of the DAG is assigned an appropriate flow in hierarchical order. The specific algorithm is as shown in Figure 5 Specifically, for the task nodes at each level, assign increasing ids in sequence so that each task node has a unique id in the current level; first calculate the allocated flow for those task nodes that have no predecessor nodes; then traverse the task nodes at the current level in ascending order of the number of predecessor nodes and perform flow allocation according to the following strategy:

[0059] If a predecessor node of the current node is the end node of its stream, the current node is assigned to the stream;

[0060] If there are multiple predecessor nodes that meet the above conditions, the stream where the predecessor node with the least unfinished successor nodes is located is selected;

[0061] If the above predecessor node does not exist, the current task node is used to calculate the assigned flow.

[0062] In this embodiment, event synchronization code is inserted into the kernel function of cross-stream synchronization, specifically:

[0063] For all nodes whose predecessor nodes are not in the same stream as themselves, in each stream with its predecessor node, select the last predecessor node in the stream, and add hipEventRecord(event,stream) after its kernel function call code hipLaunchKernel(), where event is the event trigger in the CUDA / HIP programming framework, and stream is the specified task stream under the CUDA / HIP programming framework. Among them, hipLaunchKernel() deploys the user-written function that is calculated on the GPU to the GPU to start calculation when the program is running. hipEventRecord(event,stream) inserts a monitor at the end of the specified task stream. Only when all the tasks before the monitor are executed, the state of the monitor will change, thereby changing the state of the monitors that depend on the state of the monitor, so that the tasks after these monitors can start running. The code of the aforementioned kernel function is not listed here in detail.

[0064] Then, for the task nodes that have a dependency relationship with the predecessor node and are not in the same stream, hipStreamWaitEvent(stream, event) is added before the hipLaunchKernel() statement that calls the kernel function to wait for the predecessor functions of different streams to be executed before continuing to execute this function. hipStreamWaitEvent(stream, event) inserts a monitor that depends on the specified monitor at the end of the specified task stream. When the state of the dependent monitor changes, the state of the monitor will also change, thereby starting the execution of subsequent tasks in the task stream. The code of the aforementioned kernel function is not listed here in detail.

[0065] In this embodiment, generating the executable file specifically involves converting the intermediate language representation file into the executable file through LLVM.

[0066] Through the above steps, the user only needs to write serial kernel functions, and this work can automatically identify the read-only variables and modified variables of each kernel function through the LLVM compiler, so as to construct a dependency graph DAG between kernel functions, and then use a heuristic scheduling algorithm to schedule the parallelizable kernel functions into different task streams. Different task streams can execute the tasks within their respective streams in parallel, so as to maximize the utilization of GPU resources and improve program performance.

[0067] The present invention is the first to parallelize serial GPU kernel functions through the multi-stream mechanism of the CUDA framework by means of the LLVM compiler framework, so as to improve program performance and GPU hardware utilization. The invention does not introduce additional programming models and runtime overheads, reducing the learning cost and development cost of developers. At the same time, the stream scheduling algorithm used in the invention arranges the running times of kernel functions with data dependencies close to each other, and improves the memory access efficiency of the GPU for data as much as possible, so as to improve program performance.

[0068] The preferred specific embodiments of the present invention have been described in detail above. It should be understood that those of ordinary skill in the art can make many modifications and variations according to the concept of the present invention without creative work. Therefore, all technical solutions that can be obtained by those skilled in the art in the technical field of the present invention based on the concept of the present invention through logical analysis, reasoning or limited experiments on the basis of the prior art should be within the protection scope determined by the claims.

Claims

1. An automatic multi-stream scheduling method for kernel functions based on compilation, characterized in that, It includes the following steps: Step 1, read the source code and identify the input and output of the kernel function; Step 2, construct a DAG based on the input and output; Step 3, hierarchize the DAG and assign an appropriate stream to the kernel function according to the topological sort; Step 4, insert event synchronization code for kernel functions with cross-stream synchronization; Step 5, generate an executable file; Inserting event synchronization code for the kernel function with cross-stream synchronization specifically is as follows: For all nodes whose predecessor nodes are not in the same stream as itself, in each stream that has its predecessor node, select the last predecessor node in that stream, and add code with the event trigger in the CUDA / HIP programming framework and the specified task stream in the CUDA / HIP programming framework as parameters after its kernel function call code; Then for task nodes that have a dependency relationship with the predecessor node and are not in the same stream, add code with the event trigger in the CUDA / HIP programming framework and the specified task stream in the CUDA / HIP programming framework as parameters before the statement that calls the kernel function of it to wait for the predecessor function in different streams to finish execution before continuing to execute this function.

2. The automatic multi-stream scheduling method for kernel functions based on compilation according to claim 1, wherein Reading the source code and identifying the input and output of the kernel function specifically is as follows: transform the source code into an IR file through the LLVM compilation framework, and then use the opt tool to process the IR file to analyze the read-only variables and write variables of each kernel function.

3. The automatic multi-stream scheduling method for kernel functions based on compilation according to claim 1, wherein: Constructing the DAG based on the input and output specifically is as follows: assume that the current kernel function has N input parameters, take the end node of the DAG as the starting point for each input parameter, use the reverse BFS algorithm to find the first kernel function that has the same input parameter, add this kernel function as the predecessor node of the current kernel function, and set the node represented by the current kernel function as the predecessor node of the end node in the DAG.

4. A compilation-based automatic multi-stream scheduling method for kernel functions according to claim 1, characterized in that, Assigning an appropriate stream to each kernel function in each layer of the DAG in hierarchical order specifically is as follows: assign an increasing id to the task nodes in each layer so that each task node has a unique id in the current layer; first calculate the assigned stream for those task nodes that have no predecessor nodes; Then traverse the task nodes in the current layer in ascending order of the number of predecessor nodes and perform stream assignment according to the following strategy: If a certain predecessor node of the current node is the end node of its stream, then assign the current node to that stream; If there are multiple predecessor nodes that meet the above conditions, then select the stream where the predecessor node with the fewest unfinished successor nodes is located; If there is no such predecessor node, use the stream calculated for the current task node.

5. A method for automatically multi-stream scheduling of kernel functions based on compilation, characterized in that: Generating the executable file specifically is to transform the intermediate language representation file into an executable file through LLVM.

Citation Information

Patent Citations

  • Data stream processing method and related equipment

    CN111090464A