Method, computing device, medium and program product for executing kernel function

By establishing the resource usage connection relationship between thread bundles during kernel function execution, controlling the exit time of thread bundles, solving the problem of resource waste and code rewriting caused by idle computing units, and achieving efficient resource utilization and task scheduling.

CN120371482AActive Publication Date: 2025-07-25SHANGHAI BIREN TECH CO LTD

Patent Information

Application Number
CN202510863846.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-25
Publication Date
2025-07-25
Estimated Expiration
2045-06-25

AI Technical Summary

Technical Problem

The prior art has problems with waste of resources of artificial intelligence chips and difficulty in dynamic allocation of tasks when executing kernel functions, especially the waste of resources and high code rewriting caused by idle computing units during task switching.

Method used

By establishing a common resource usage connection relationship between multiple thread bundles, the exit time of the thread bundle is controlled to ensure that other thread bundles complete resource usage before the current thread bundle exits, and then tasks are continuously executed on the same computing unit to avoid idle computing unit.

Benefits of technology

It significantly improves the resource utilization rate of artificial intelligence chips, reduces the difficulty of execution costs and dynamic allocation of tasks, and maintains the reusability of the original code.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120371482A_ABST
    Figure CN120371482A_ABST
Patent Text Reader

Abstract

The invention relates to a method, a computing device, a medium and a program product for executing a kernel function. The method comprises the following steps: establishing a common use connection relationship of a plurality of thread bundles related to the same task for a common resource; and controlling the exit time of the current thread bundle of the current task in the plurality of tasks at least based on the established common use connection relationship so as to enable the current thread bundle to exit before the current thread bundle exits. All other thread bundles of the current task, which have a common use connection relationship with the current thread bundle, finish the use of corresponding public resources; and in response to determining that the current thread bundle exits, starting a corresponding thread bundle of a next task, the current task and the next task being executed by the same computing unit. According to the method, the utilization rate of the artificial intelligence chip can be remarkably improved, and the execution cost and the difficulty of dynamic task allocation can be effectively reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present invention generally relate to the field of artificial intelligence, and more particularly to a method for executing a kernel function, a computing device, a computer-readable storage medium, and a computer program product. Background Art

[0002] When executing a kernel function, the original task is usually divided into multiple tasks and handed over to multiple compute units (CUs) for execution. Each compute unit needs to execute several tasks. It should be understood that a warp is the most basic execution unit, and each task on an artificial intelligence chip usually consists of several warps. The artificial intelligence chip may be, for example but not limited to: a graphics processing unit (GPU), a general purpose computing on graphics processing units (GPGPU), or a tensor processing unit (TPU), etc. When executing a kernel function, generally, after each warp in several warps of the current task is completed, the current task can be switched to the next task. It should be understood that different warps have different responsibilities, some warps will end very early, and some warps will start late. When a compute unit executes multiple tasks, there is a compute unit idle (or "execution bubble", i.e., Gap) between the front and back tasks.

[0003] Traditional solutions for executing a kernel function mainly include two types. In the first solution, no processing is performed on the existing compute unit idle. In the second solution, the original task is rewritten so that the number of tasks is the same as the number of compute units. For example, only one task is executed on one CU, and several subtasks are processed inside the task, so that there is no task switching, and thus there is no compute unit idle between the front and back task switches. For the first solution, since no processing is performed on the compute unit idle between the front and back tasks, it will cause waste of resources of the artificial intelligence chip. Especially for the case where the execution time of the task itself is short, the compute unit idle caused by different warps may reach 10-20% of the warp execution time. For the second solution, on the one hand, rewriting the original task will result in a large amount of code being rewritten, thus resulting in a high migration cost; on the other hand, since the number of dynamically available execution units is uncertain, it is difficult to dynamically allocate tasks.

[0004] In summary, the deficiencies of the traditional method for executing a kernel function are as follows: it is difficult to reduce the execution cost and the difficulty of task dynamic allocation while avoiding waste of artificial intelligence chip resources. Summary of the Invention

[0005] The present invention provides a method for executing a kernel function, a computing device, a computer-readable storage medium, and a computer program product, which can not only significantly improve the utilization rate of artificial intelligence chip resources, but also effectively reduce the execution cost and the difficulty of task dynamic allocation.

[0006] According to a first aspect of the present invention, there is provided a method for executing a kernel function. The method includes: establishing a common usage connection relationship of multiple warps related to the same task for a common resource, where the same task is any one of multiple tasks for executing a kernel function; controlling, at least based on the established common usage connection relationship, the exit time of the current warp of the current task among the multiple tasks, so that all other warps of the current task that have a common usage connection relationship with the current warp complete the usage of the corresponding common resource before the current warp exits; and in response to determining the exit of the current warp, starting the corresponding warp of the next task, where the current task and the next task are executed by the same computing unit, and the same computing unit is any one of multiple computing units for executing a kernel function.

[0007] In some embodiments, starting the corresponding warp of the next task includes: making the corresponding warp of the next task not use the common resources that the other warps of the current task still need to access.

[0008] In some embodiments, establishing a common usage connection relationship of multiple warps related to the same task for a common resource includes: numbering the multiple warps based on the original code of the same task according to the exit order of the multiple warps related to the same task; and marking the usage status of each warp in the multiple warps for the common resource, so as to establish a common usage connection relationship of the multiple warps for the common resource.

[0009] In some embodiments, controlling, at least based on the established common usage connection relationship, the exit time of the current warp of the current task among the multiple tasks includes: in response to determining that each warp of the current task has completed the usage of the corresponding common resource, notifying other warps that have a common usage connection relationship for the corresponding common resource.

[0010] In some embodiments, controlling the exit time of the current warp of the current task among the multiple tasks is at least based on the established co-usage connection relationship and includes: in response to determining that each warp of the current task has completed using the corresponding common resource, notifying other warps of the current task.

[0011] In some embodiments, notifying other warps having a co-usage connection relationship for the corresponding common resource includes: in response to determining that each warp of the current task has completed using the corresponding common resource, sending indication information for indicating the completion of the use of the corresponding common resource to other warps having a co-usage connection relationship for the corresponding common resource.

[0012] In some embodiments, controlling the exit time of the current warp of the current task among the multiple tasks is at least based on the established co-usage connection relationship and further includes: in response to determining that the current warp of the current task has completed the current task and received the indication information for the completion of the use of the corresponding common resource sent by all other warps having a co-usage connection relationship, allowing the current warp of the current task to exit; and in response to determining that the current warp of the current task has not completed the current task or has not received the indication information for the completion of the use of the corresponding common resource sent by all other warps having a co-usage connection relationship, not allowing the current warp of the current task to exit.

[0013] In some embodiments, establishing a co-usage connection relationship for a common resource among multiple warps related to the same task includes: confirming whether a signal indicating that the concurrent mode switch has been turned on is detected; in response to confirming that the signal indicating that the concurrent mode switch has been turned on is detected, establishing a co-usage connection relationship for a common resource among multiple warps related to the same task; in response to confirming that the signal indicating that the concurrent mode switch has been turned on is not detected, sending an instruction to turn on the concurrent mode switch so that each computing unit in the concurrent mode can concurrently execute two warp groups.

[0014] In some embodiments, starting the corresponding warp of the next task includes: in response to confirming that the current computing unit in the concurrent mode needs to receive the second warp group when the first warp group has not completed the task, confirming whether the head warp of the first warp group has exited; in response to confirming that the head warp of the first warp group has exited, allowing the head warp of the second warp group to be dispatched to the current computing unit for execution; confirming whether the non-head warps of the first warp group have exited; and in response to confirming that the non-head warps of the first warp group have exited, allowing the non-head warps of the second warp group to be dispatched to the current computing unit for execution.

[0015] According to a second aspect of the present invention, there is also provided a computing device. The computing device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the computing device to execute the method of the first aspect of the present invention.

[0016] According to a third aspect of the present invention, there is also provided a computer-readable storage medium. A computer program is stored on the computer-readable storage medium, and when the computer program is executed by a machine, it executes the method of the first aspect of the present invention.

[0017] According to a fourth aspect of the present invention, there is also provided a computer program product, including a computer program which, when executed by a machine, executes the method of the first aspect of the present invention.

[0018] The present invention enables the warp corresponding to the subsequent task to start execution without waiting for all warps of the previous task to completely end, thereby enabling the previous and subsequent tasks for kernel function execution to be continuously and overlappedly executed on a computing unit, without any idle computing unit between the previous and subsequent tasks, thus significantly improving the utilization rate of the artificial intelligence chip. At the same time, the present invention still follows the classic task splitting logic of the artificial intelligence chip, resulting in strong task dynamic scheduling ability, without the need to largely modify the kernel function code, and thus facilitating the reuse of the original code. Therefore, the present invention can not only significantly improve the utilization rate of artificial intelligence chip resources, but also reduce the execution cost and the difficulty of task dynamic allocation.

[0019] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In combination with the accompanying drawings and with reference to the following detailed description, the above and other features, advantages and aspects of the embodiments of the present invention will become more obvious. In the drawings, the same or similar reference numerals denote the same or similar elements.

[0021] Figure 1 A schematic diagram schematically shows a computing device for implementing a method for executing a kernel function according to an embodiment of the present invention.

[0022] Figure 2 A flowchart shows a method for executing a kernel function according to an embodiment of the present invention.

[0023] Figure 3 A schematic diagram shows a common usage connection relationship of multiple warps related to the same task for a common resource according to an embodiment of the present invention.

[0024] Figure 4A A schematic diagram showing a method for executing a kernel function according to an embodiment of the present invention.

[0025] Figure 4B Another schematic diagram showing a method for executing a kernel function according to an embodiment of the present invention.

[0026] Figure 5 A flowchart showing a method for controlling the exit of a current warp of a current task according to an embodiment of the present invention.

[0027] Figure 6 A flowchart showing a method for starting a corresponding warp of a next task according to an embodiment of the present invention.

[0028] Figure 7 A schematic diagram showing a method for starting a corresponding warp of a next task according to an embodiment of the present invention.

[0029] Figure 8 A schematic diagram showing a traditional method of not dealing with the idle computing units when executing a kernel function.

[0030] Figure 9 A schematic diagram showing a traditional method of controlling the number of tasks to be the same as the number of computing units when executing a kernel function.

[0031] In the respective drawings, the same or corresponding reference numerals denote the same or corresponding parts. Detailed Embodiment

[0032] The preferred embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although the preferred embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present invention will be more thorough and complete, and will fully convey the scope of the present invention to those skilled in the art.

[0033] The term "comprising" and variations thereof used herein mean open-ended inclusion, i.e., "including but not limited to". The term "based on" means "at least partially based on". The terms "an example embodiment" and "an embodiment" mean "at least one example embodiment". The term "another embodiment" means "at least one additional embodiment". The terms "first", "second", etc. may refer to different or the same objects.

[0034] As described above, the traditional solutions for executing a kernel function mainly include two types. In the first solution, no treatment is performed for the existing idle computing units. Figure 8Shows a schematic diagram of a traditional method 800 for not handling the idle computing units when executing a kernel function. As Figure 8 shown, when executing a kernel function, after the zero-th warp, the first warp, and the second warp of the previous task (for example, the zero-th task, i.e., "Task 0") have all completed their tasks and exited, the next task (for example, the first task, i.e., "Task 1") can be started. When executing Task 0, the zero-th warp of Task 0 finishes the task first, and then the first warp and the second warp finish the task successively. The zero-th warp of Task 1 is started only after the second warp of Task 0 has finished the task and exited. Therefore, there is a computing unit idle 810 (or "execution bubble", i.e., Gap) between the two consecutive tasks Task 0 and Task 1. Similarly, there is also a computing unit idle 810 between Task 1 and Task 2. In the above first scheme, since the idle computing units between consecutive tasks are not handled, it will lead to waste of artificial intelligence chip resources. Especially for the case where the execution time of the Task itself is short, the computing unit idle caused by different warps may even reach 10 - 20% of the warp execution time.

[0035] In the second scheme, it is necessary to rewrite the original task to control the number of tasks to be the same as the number of computing units. Figure 9 Shows a schematic diagram of a traditional method 900 for controlling the number of tasks to be the same as the number of computing units when executing a kernel function. For example, only one task (for example, Figure 9 the zero-th task 910 shown, i.e., "Task 0") is executed on a CU, and several subtasks (for example, subtask 0, subtask 1... subtask M, where M is a natural number) are processed inside this task, so that there is no task switching on the CU, and thus no computing unit idle between consecutive task switches. For the above second scheme, on the one hand, rewriting the original task will result in a large amount of code being rewritten, thus leading to a relatively high migration cost; on the other hand, since the number of dynamically available execution units is uncertain, it is difficult to perform dynamic task allocation.

[0036] In summary, the deficiencies of the traditional method for executing a kernel function are as follows: it is difficult to avoid wasting artificial intelligence chip resources while reducing the execution cost and the difficulty of dynamic task allocation.

[0037] To at least partially solve one or more of the above problems and other potential problems, exemplary embodiments of the present invention propose a solution for executing a kernel function. In this solution, by establishing a connection relationship for the common use of a plurality of warps related to the same task among a plurality of tasks for kernel function execution; and controlling the exit time of the current warp among the plurality of warps related to the current task in the plurality of tasks at least based on the established connection relationship for common use, so that before the current warp exits, all other warps related to the current task and having a connection relationship for common use with the current warp have completed the use of the corresponding common resource; and when it is determined that the current warp related to the current task exits, starting the corresponding warp of the next task, the present invention can enable the corresponding warp of the subsequent task to start execution without waiting for all warps of the previous task to completely end, thereby enabling the previous and subsequent tasks for kernel function execution to be continuously and overlappedly executed on a computing unit, and there is no computing unit idle between the previous and subsequent tasks, thus significantly improving the utilization rate of the artificial intelligence chip. At the same time, the present invention still follows the classic task splitting logic of the artificial intelligence chip, so that the task dynamic scheduling ability is strong, and there is no need to largely modify the kernel function code, which is beneficial to improving the reusability of the original code. Therefore, the present invention can not only significantly improve the utilization rate of the artificial intelligence chip resources, but also effectively reduce the execution cost and the difficulty of task dynamic allocation.

[0038] Figure 1 FIG. schematically shows a schematic diagram of a computing device 100 for implementing a method for executing a kernel function according to an embodiment of the present invention. As Figure 1As shown, the computing device 100 may have one or more processing units, including dedicated processing units such as a Graphics Processing Unit (GPU), a Field Programmable Gate Array (FPGA), an Application Specific Integrated Circuit (ASIC), a General-purpose computing on graphics processing units (GPGPU), or a Tensor Processing Unit (GPGPU), as well as a general-purpose processing unit such as a CPU. The computing device 100 further includes at least: a shared connection relationship establishment unit 102, an exit time control unit 104 for the current warp, and a corresponding warp startup unit 106 for the next task. It should be understood that the above-mentioned shared connection relationship establishment unit 102, the exit time control unit 104 for the current warp, and the corresponding warp startup unit 106 for the next task may be software modules, and these software modules run, for example, on one or more processing units configured in the computing device 100. The one or more processing units may be integrated with multiple computing units for executing kernel functions, or may be separately arranged on different computing devices from the multiple computing units for executing kernel functions.

[0039] Regarding the shared connection relationship establishment unit 102, it is used to establish a shared connection relationship for multiple warps related to the same task with respect to a common resource, where the same task is any one of the multiple tasks for executing kernel functions.

[0040] Regarding the exit time control unit 104 for the current warp, it is used to control the exit time of the current warp of the current task among the multiple tasks at least based on the established shared connection relationship, so that all other warps of the current task that have a shared connection relationship with the current warp complete the use of the corresponding common resource before the current warp exits.

[0041] Regarding the corresponding warp startup unit 106 for the next task, it is used to respond to determining whether the current warp exits; and in response to determining that the current warp exits, start the corresponding warp of the next task, where the current task and the next task are executed by the same computing unit, and the same computing unit is any one of the multiple computing units for executing kernel functions.

[0042] The following will be combined withFigure 2 , Figure 3 , Figure 4A and Figure 4B A method 200 for executing a kernel function for describing an embodiment of the present invention. Figure 2 A flowchart of a method 200 for executing a kernel function according to an embodiment of the present invention is shown. Figure 3 A schematic diagram of a common use connection relationship 300 of multiple warps related to the same task for a common resource according to an embodiment of the present invention is shown. Figure 4A A schematic diagram of a method for executing a kernel function according to an embodiment of the present invention is shown. Figure 4B Another schematic diagram of a method for executing a kernel function according to an embodiment of the present invention is shown. It should be understood that the method 200 can be executed, for example, at Figure 1 the described computing device 100. The method 200 may further include additional actions not shown and / or may omit the shown actions, and the scope of the present invention is not limited in this regard.

[0043] At step 202, the computing device 100 establishes a common use connection relationship of multiple warps related to the same task for a common resource. Regarding multiple warps related to the same task, it means that multiple warps serve the same task. For example, in Figure 4A , the zeroth warp 310-1 for the same current task 410, the first warp 320-1 of the current task 410, and the second warp 330-1 of the current task 410 are multiple warps related to the same task.

[0044] Regarding the method of establishing the common use connection relationship, in some embodiments, it includes, for example: the computing device 100 confirms whether a signal indicating that the concurrent mode switch has been turned on is detected; in response to confirming the detection of the signal indicating that the concurrent mode switch has been turned on, establishes a common use connection relationship of multiple warps related to the same task for a common resource; in response to confirming that the signal indicating that the concurrent mode switch has been turned on is not detected, sends an instruction to turn on the concurrent mode switch so that each computing unit in the concurrent mode can concurrently execute two warp groups.

[0045] It should be understood that the concurrent mode switch is open to the upper-layer software for multiple computing units for executing the kernel function to turn on or open. When the concurrent mode switch is turned on, each computing unit among the multiple computing units can concurrently execute two warp groups; when the concurrent mode switch is not turned on, each computing unit can only execute one warp group.

[0046] It should be understood that the execution of a kernel function generally involves multiple tasks. For example, the executed kernel function is matrix multiplication. That is, C = A * B. Wherein, A and B are respectively two input matrices for which multiplication operations are to be performed. C is, for example, an output matrix, which is, for example, a 4086x4096 matrix. In some embodiments, in order to complete matrix multiplication, the original task of the matrix multiplication kernel function is usually divided into multiple tasks (the number of tasks is, for example, 1024). Each task, for example, calculates an output block, and the size of the output block is, for example, 128x128. For example, 32 CUs can be used to complete 1024 tasks. Each CU needs to run 32 tasks.

[0047] It should be understood that the tasks before and after in the multiple kernel tasks for kernel function execution usually access common resources, so conflicts in accessing common resources need to be prevented.

[0048] In some embodiments, regarding the method for establishing a common usage connection relationship, for example, it includes: the computing device 100 numbers the multiple warps according to the exit order of the multiple warps related to the same task based on the original code of the same task; and marks the usage status of each warp in the multiple warps for the common resources, so as to establish a common usage connection relationship between the multiple warps for the common resources. For example, for the current task (which is, for example, but not limited to "Task 0"), the warps related to the current task are numbered according to the exit order of the multiple (for example, N, where N is a natural number) warps related to the current task. For example, the warps related to the current task are numbered as: the zeroth warp... the Nth warp (that is, Warp0... WarpN). In some embodiments, the common resources are, for example, but not limited to a shared memory address (SharedMemory Address). For example, the computing device 100 divides the shared memory address into address segments to determine the warps accessing each address segment. For example, the divided address segments are respectively: the shared A segment (that is, the S_A segment, and its corresponding address segment is, for example, [0, 128KB]), the shared B segment (that is, the S_B segment, and its corresponding address segment is, for example, [128KB, 160KB]), and the shared C segment (that is, the S_C segment, and its corresponding address segment is, for example, [160KB, 192KB]). Among them, the warps accessing the shared A segment include: the zeroth warp, the first warp, and the second warp (that is, corresponding to Warp 0, Warp1, Warp2 respectively). The warps accessing the shared B segment include: the zeroth warp, the first warp (that is, corresponding to Warp 0, Warp1 respectively). The warps accessing the shared C segment include: the zeroth warp, the second warp (that is, corresponding to Warp 0, Warp2 respectively). As indicated in Table 1 below.

[0049] Table 1

[0050] As Figure 3 shown, in the established co - usage connection relationship 300, the zero - th thread bundle 310 and the first thread bundle 320 have a co - usage situation for the shared B section 342. The zero - th thread bundle 310, the first thread bundle 320, and the second thread bundle 330 have a co - usage situation for the shared A section 340. The zero - th thread bundle 310 and the second thread bundle 330 have a co - usage situation for the shared C section 344.

[0051] At step 204, the computing device 100 controls the exit time of the current thread bundle of the current task among the multiple tasks at least based on the established co - usage connection relationship, so that before the current thread bundle exits, all other thread bundles of the current task that have a co - usage connection relationship with the current thread bundle have completed the use of the corresponding common resources.

[0052] As Figure 4A shown, when the zero - th thread bundle 310 - 1 of the current task 410 completes the current task at time point 420, since the computing device 100 has not received the indication information of the completion of the use of the corresponding shared resources sent by the first thread bundle 320 - 1 and the second thread bundle 330 - 1 of the current task 410 that have a co - usage connection relationship with the shared A section, the shared B section, and the shared C section, the computing device 100 does not allow the zero - th thread bundle 310 - 1 of the current task 410 to exit at time point 420. Instead, it controls the exit time of the zero - th thread bundle 310 - 1 of the current task 410 so that its exit time is delayed to time point 430. At time point 430, the zero - th thread bundle 310 - 1 of the current task 410 has received the indication information of the completion of the use of the corresponding shared resources sent by the first thread bundle and the second thread bundle that have a co - usage connection relationship with the shared A section, the shared B section, and the shared C section. Therefore, the computing device 100 allows the zero - th thread bundle 310 - 1 of the current task 410 to exit.

[0053] A method for controlling the exit time of the current warp of the current task among the multiple tasks, for example, includes various ways. In some embodiments, the method for controlling the exit time of the current warp, for example, includes: the computing device 100, in response to determining that the current warp of the current task has completed the current task and received the indication information for the completion of the use of the corresponding common resource sent by all other warps having a common use connection relationship, allows the current warp of the current task to exit; and in response to determining that the current warp of the current task has not completed the current task or has not received the indication information for the completion of the use of the corresponding common resource sent by all other warps having a common use connection relationship, does not allow the current warp of the current task to exit. The following will be combined with Figure 5 to detail the method 500 for controlling the exit time of the current warp of the current task among the multiple tasks, and will not be elaborated here.

[0054] At step 206, the computing device 100 determines whether the current warp exits.

[0055] If the computing device 100 determines that the current warp has not exited, jump to step 210 and do not start the corresponding warp of the next task. As Figure 4A shown, if the computing device 100 determines that the zero-th warp 310-1 (e.g., Warp 0) of the current task 410 (e.g., Task0) has not exited, for example, between time point 420 and time point 430, do not start the zero-th warp 310-2 (e.g., Warp 0) of the next task 460 (e.g., Task1).

[0056] At step 208, if the computing device 100 determines that the current warp exits, start the corresponding warp of the next task, where the current task and the next task are executed by the same computing unit, and the same computing unit is any one of the multiple computing units for executing the kernel function.

[0057] In some embodiments, starting the corresponding warp of the next task includes: making the corresponding warp of the next task not use the common resource that the other warps of the current task still need to access.

[0058] Method for starting a corresponding warp of the next task, which may include, for example: in response to confirming that the current computing unit in concurrent mode needs to receive a second warp group while the first warp group has not completed the task, confirming whether the head warp in the first warp group has exited; in response to confirming that the head warp in the first warp group has exited, allowing the head warp in the second warp group to be dispatched to the current computing unit for execution; confirming whether the non-head warps in the first warp group have exited; and in response to confirming that the non-head warps in the first warp group have exited, allowing the non-head warps in the second warp group to be dispatched to the current computing unit for execution. It should be understood that the head warp in the warp group is mainly used for data transfer operations; the non-head warps in the warp group are mainly used for execution operations in tensor computing kernels such as matrix multiplication and convolution calculations (e.g., mma / conv) (tcore enable kernel). Therefore, the head warp will end the task earlier than the non-head warp.

[0059] For example, if the computing device 100 determines that the zeroth warp 310-1 (e.g., Warp 0 of Task 0) of the current task 410 (e.g., Task 0) exits, start the zeroth warp 310-2 (e.g., Warp 0 of Task 1) of the next task 460 (e.g., Task 1).

[0060] As Figure 4B shown, there is an overlapping execution section 470 between the zeroth warp 310-2 of the next task 460 (e.g., Task 1) and the first warp 320-1 (e.g., Warp 1 of Task 0) and the second warp 330-1 (e.g., Warp 2 of Task 0) of the current task 410.

[0061] And so on, if the computing device 100 determines that the first warp 320-1 (e.g., Warp 1 of Task 0) of the current task 410 (e.g., Task 0) exits, start the first warp 320-2 (e.g., Warp 1 of Task 1) of the next task 460 (e.g., Task 1). If the computing device 100 determines that the second warp 330-1 (e.g., Warp 2 of Task 0) of the current task 410 (e.g., Task 0) exits, start the second warp 330-2 (e.g., Warp 1 of Task 1) of the next task 460 (e.g., Task 1).

[0062] In the above solution, the present invention enables the warp corresponding to the subsequent task to start execution without waiting for all warps of the previous task to complete, so that the previous and subsequent tasks for kernel function execution can be continuously and overlapped on a computing unit, and there is no idle computing unit between the previous and subsequent tasks, thus significantly improving the utilization rate of the artificial intelligence chip. At the same time, the present invention still follows the classical task splitting logic of the artificial intelligence chip, so that the task dynamic scheduling ability is strong, and there is no need to massively modify the kernel function code, which is conducive to improving the reusability of the original code. Therefore, the present invention can not only significantly improve the utilization rate of the artificial intelligence chip, but also effectively reduce the execution cost and the difficulty of task dynamic allocation.

[0063] The following will combine Figure 5 to describe method 500 for allowing the current warp of the current task to exit. Figure 5 FIG. shows a flowchart of method 500 for controlling the current warp of the current task to exit according to an embodiment of the present invention. It should be understood that method 500 can be executed, for example, at Figure 1 the computing device 100 described. Method 500 may also include additional actions not shown and / or may omit the actions shown, and the scope of the present invention is not limited in this regard.

[0064] At step 502, the computing device 100 numbers the multiple warps according to the exit order of the multiple warps related to the same task based on the original code of the same task.

[0065] At step 504, the computing device 100 marks the usage status of each warp in the multiple warps for the common resources, so as to establish a common usage connection relationship of the multiple warps for the common resources.

[0066] For example, as Figure 4A shown, for the shared section A, the zero-th warp 310-1 of the current task 410 has a common usage connection relationship with the first warp 320-1 of the current task 410, and the zero-th warp 310-1 of the current task 410 has a common usage connection relationship with the second warp 330-1 of the current task 410. For the shared section B, the zero-th warp 310-1 of the current task 410 has a common usage connection relationship with the first warp 320-1 of the current task 410. For the shared section C, the zero-th warp 310-1 of the current task 410 has a common usage connection relationship with the second warp 330-1 of the current task 410.

[0067] At step 506, if the computing device 100 determines that each warp of the current task has completed using the corresponding common resource, it notifies other warps having a common usage connection relationship for the corresponding common resource.

[0068] For example, if it is determined that each warp of the current task has completed using the corresponding common resource, an indication message for indicating the completion of the use of the corresponding common resource is sent to other warps that have a common use connection relationship with the corresponding common resource. As Figure 4A shown, after the zero-th warp 310-1 of the current task 410 has completed using the shared A section, a first indication message of the completion of the use of the shared A section (as indicated by the marker 412) is sent to the first warp 320-1 and the second warp 330-1 of the current task 410 that have a common use connection relationship with the shared A section. After the zero-th warp of the current task 410 has completed using the shared B section, a first indication message of the completion of the use of the shared B section (as indicated by the marker 414) is sent to the first warp 320-1 of the current task 410 that has a common use connection relationship with the shared B section; and after the zero-th warp 310-1 of the current task 410 has completed using the shared C section, a second indication message of the completion of the use of the shared C section (as indicated by the marker 416) is sent to the second warp 330-1 of the current task 410 that has a common use connection relationship with the shared C section.

[0069] It should be understood that in some other embodiments, if the computing device 100 determines that each warp of the current task has completed using the corresponding common resource, all other warps of the current task are notified, not limited to other warps that have a common use connection relationship.

[0070] At step 508, if the computing device 100 determines that the current warp of the current task has completed the current task and has received the indication messages of the completion of the use of the corresponding common resource sent by all other warps that have a common use connection relationship, the current warp of the current task is allowed to exit.

[0071] If the computing device 100 determines that the current warp of the current task has not completed the current task or has not received the indication messages of the completion of the use of the corresponding common resource sent by all other warps that have a common use connection relationship, the current warp of the current task is not allowed to exit.

[0072] As Figure 4AAs shown, the zero-thread bundle 310-1 of the current task 410 completes the current task at time point 420. However, at time point 420, the zero-thread bundle 310-1 of the current task 410 has not received the indication information for the completion of the use of the corresponding common resource sent by the first-thread bundle 320-1 and the second-thread bundle 330-1 of the current task 410 with a common usage connection relationship. Thus, the zero-thread bundle 310-1 of the current task 410 is not allowed to exit. For example, after the first-thread bundle 320-1 of the current task 410 completes the use of the shared section A, it sends the second indication information for the completion of the use of the shared section A (as indicated by the marker 422) to the zero-thread bundle 310-1 of the current task 410. After the second-thread bundle 330-1 of the current task 410 completes the use of the shared section A, it sends the third indication information for the completion of the use of the shared section A (as indicated by the marker 424) to the zero-thread bundle 310-1 of the current task 410. Further, after the first-thread bundle 320-1 of the current task 410 completes the use of the shared section B, it sends the second indication information for the completion of the use of the shared section B (as indicated by the marker 426) to the zero-thread bundle 310-1 of the current task 410. After the second-thread bundle 330-1 of the current task 410 completes the use of the shared section C, it sends the second indication information for the completion of the use of the shared section C (as indicated by the marker 428) to the zero-thread bundle 310-1 of the current task 410. At time point 430, the zero-thread bundle 310-1 of the current task 410 has received the indication information for the completion of the use of the corresponding common resource sent by the first-thread bundle and the second-thread bundle with a common usage connection relationship for the shared section A, shared section B, and shared section C. Therefore, the zero-thread bundle 310-1 of the current task 410 is allowed to exit. From Figure 4A It can be seen that the zero-thread bundle 310-1 of the current task 410 completes the current task at time point 420. However, at time point 420, the zero-thread bundle 310-1 of the current task 410 is not allowed to exit. Instead, the exit time of the zero-thread bundle 310-1 of the current task 410 is controlled to be delayed until time point 430, that is, until after receiving the indication information for the completion of the use of the corresponding common resource sent by all other thread bundles with a common usage connection relationship.

[0073] Similarly, as Figure 4A shown, at time point 440, the first-thread bundle 320-1 of the current task 410 completes the task and has received the indication information for the completion of the use of the corresponding common resource sent by the zero-thread bundle 310-1 of the current task 410 with a common usage connection relationship for the shared section A and shared section B. Therefore, the first-thread bundle 320-1 of the current task 410 is allowed to exit.

[0074] By analogy, at time point 450, the second warp bundle 330-1 of the current task 410 has completed the task and has received the indication information for the completion of the use of the corresponding common resource sent by the zeroth warp bundle 310-1 of the current task 410 that has a common usage connection relationship with the shared A section and the shared C section. Therefore, the second warp bundle 330-1 of the current task 410 is allowed to exit.

[0075] By adopting the above means, the present invention can bind the exit time of each warp bundle to the completion of the use of the corresponding common resource, so as to ensure that when each warp bundle exits, other warp bundles with a common usage connection relationship have completed the use of the corresponding common resource, thereby avoiding resource conflicts with the corresponding warp bundles of the next task.

[0076] The following will be combined with Figure 6 and Figure 7 to describe the method 600 for starting the corresponding warp bundle of the next task in the embodiments of the present invention. Figure 6 FIG. shows a flowchart of the method 600 for starting the corresponding warp bundle of the next task according to an embodiment of the present invention. Figure 7 FIG. shows a schematic diagram of the method for starting the corresponding warp bundle of the next task according to an embodiment of the present invention. It should be understood that the method 600 can be executed, for example, at the computing device 100 described in Figure 1 The method 600 may further include additional actions not shown and / or may omit the actions shown. The scope of the present invention is not limited in this regard.

[0077] At step 602, in response to confirming that the current computing unit in the concurrent mode needs to receive the second warp bundle group when the first warp bundle group has not completed the task, confirm whether the head warp bundle in the first warp bundle group has exited.

[0078] Regarding the warp bundle group, it is a warp bundle group (thread group, or workgroup for short, abbreviated as "tg") for tasks related to the kernel function.

[0079] Regarding the first warp bundle group, it is, for example, several warp bundles (Warp0...WarpN) of the current task related to the kernel function. In some embodiments, the first warp bundle group is, for example, but not limited to, Warp0...Warp3 of Task0.

[0080] Regarding the second warp bundle group, it is, for example, several warp bundles (Warp0...WarpN) of the next task related to the kernel function. In some embodiments, the second warp bundle group is, for example, but not limited to, Warp0...Warp3 of Task1.

[0081] It should be understood that the first warp group and the second warp group are two adjacent warp groups that can be executed concurrently among multiple warp groups. For example, they are tg0 and tg1 among three warp groups (tg0, tg1, and tg2), and they can also be tg1 and tg2.

[0082] As Figure 7 shown, the first warp group 710 starts to be executed at the first time point 712. The second warp group 720 needs to be received when the first warp group 710 has not completed its tasks. The first warp group 710 includes a head warp and non-head warps. The second warp group 720 also includes a head warp and non-head warps. The head warp is used for the data transfer operation related to the execution of the kernel function; the non-head warps are used for the execution operation of the kernel function. As Figure 7 shown, at the first time point 712, the head warp of the first warp group 710 is started to execute tasks. For example, if the computing device 100 confirms that the second warp group 720 needs to be received when the first warp group 710 has not completed its tasks, it further confirms whether the head warp in the first warp group 710 has exited.

[0083] At step 604, in response to confirming that the head warp in the first warp group has not exited, the head warp in the second warp group is not allowed to be dispatched to the current computing unit.

[0084] For example, between the first time point 712 and the second time point 714, the head warp of the first warp group 710 has not completed its tasks and exited. At this time, the head warp of the second warp group 720 is not allowed to be dispatched to the current computing unit.

[0085] At step 606, in response to confirming that the head warp in the first warp group has exited, the head warp in the second warp group is allowed to be dispatched to the current computing unit for execution.

[0086] For example, if the computing device 100 confirms that the head warp in the first warp group 710 has exited (for example, at the second time point 714, the head warp in the first warp group 710 ends its tasks and exits), the head warp in the second warp group 720 is allowed to be dispatched to the current computing unit for execution.

[0087] At step 608, it is confirmed whether the non-head warps in the first warp group have exited.

[0088] For example, the computing device 100 confirms whether the non-head warps in the first warp group 710 have exited (for example, at the third time point 716, the non-head warps in the first warp group 710 end their tasks and exit).

[0089] At step 610, in response to confirming that the non-head warps in the first warp group have not exited, the non-head warps in the second warp group are not allowed to be dispatched to the current computing unit.

[0090] For example, if it is confirmed that the non-head warps of the first warp group 710 have not completed their tasks and exited (e.g., between the second time point 714 and the third time point 716 in Figure 7 ), at this time, the non-head warps of the second warp group 720 are not allowed to be dispatched to the current computing unit.

[0091] At step 612, in response to confirming that the non-head warps in the first warp group have exited, the non-head warps in the second warp group are allowed to be dispatched to the current computing unit for execution.

[0092] As Figure 7 shown, if the computing device 100 confirms that the non-head warps in the first warp group have exited (e.g., at the third time point 716, the non-head warps in the first warp group 710 complete their tasks and exit), the non-head warps in the second warp group are allowed to be dispatched to the current computing unit for execution.

[0093] And so on. In some embodiments, if the computing device 100 further confirms that the head warp in the second warp group 720 has exited (e.g., at the fourth time point 722, the head warp in the second warp group 720 completes its task and exits), the head warp in the third warp group 730 is allowed to be dispatched to the current computing unit for execution. If the computing device 100 further confirms that the non-head warps in the second warp group have exited (e.g., at the fifth time point 724, the non-head warps in the second warp group 720 complete their tasks and exit), the non-head warps in the third warp group 730 are allowed to be dispatched to the current computing unit for execution.

[0094] In the above solution, the present invention can achieve concurrent execution of two warp groups in a tensor computing kernel kernel function in the same computing unit (CU), thereby significantly improving the execution efficiency of warp groups in the computing unit and reducing the running time of the kernel function. Further, the present invention can achieve the execution of a kernel function with the number of warp groups exceeding the number of computing units.

[0095] The various processes and treatments described above, such as methods 200 to 500, may be executed at a computing device. The computing device includes, for example: at least one processor (at least one graphics processor and at least one central processor); and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor. In some embodiments, methods 200 to 500 may be implemented as a computer software program or program product, which is tangibly contained in a machine-readable medium. In some embodiments, part or all of the computer program may be loaded and / or installed onto the computing device via a read-only memory (ROM) and / or a communication unit. When the computer program is loaded into a random-access memory (RAM) and executed by a GPU and a CPU, one or more actions of methods 200 to 500 described above may be executed.

[0096] The present invention may be a method, an apparatus, a system, and / or a computer program product. The computer program product may include a computer-readable storage medium having thereon computer-readable program instructions for performing various aspects of the present invention. The computer-readable storage medium may be a tangible device that can retain and store instructions used by an instruction execution device. The computer-readable storage medium may be, for example, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing.

[0097] The computer-readable program instructions described herein may be downloaded from the computer-readable storage medium to various computing / processing devices, or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. Aspects of the present invention are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each block of the flowchart illustrations and / or block diagrams, and the combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0098] These computer-readable program instructions can be provided to a central processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that the instructions, when executed by the central processing unit of the computer or other programmable data processing apparatus, create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions may also be stored in a computer-readable storage medium that causes a computer, a programmable data processing apparatus, and / or other devices to operate in a particular manner, such that the computer-readable medium storing the instructions comprises a manufacture including instructions for implementing various aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0099] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a segment of code, or a portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two consecutive blocks may in fact be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or actions, or by a combination of dedicated hardware and computer instructions.

[0100] It should be understood that various forms of the flows shown above may be used, with steps reordered, added, or deleted. For example, the steps recited in this application may be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution disclosed in this application can be achieved, and this is not limited herein.

[0101] The above specific embodiments do not constitute a limitation on the scope of protection of this application. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors.

Claims

1. A method for executing a kernel function, characterized in that, Including: Establishing a common usage connection relationship for multiple warps related to the same task with respect to a common resource; Controlling, at least based on the established common usage connection relationship, the exit time of the current warp of the current task among the multiple tasks, so that all other warps of the current task that have a common usage connection relationship with the current warp have completed the usage of the corresponding common resource before the current warp exits; and In response to determining that the current warp exits, starting the corresponding warp of the next task, where the current task and the next task are executed by the same computing unit, and the same computing unit is any one of multiple computing units for executing a kernel function.

2. The method according to claim 1, characterized in that, Starting the corresponding warp of the next task includes: Making the corresponding warp of the next task not use the common resource that other warps of the current task still need to access.

3. The method according to claim 1, wherein Establishing a common usage connection relationship for multiple warps related to the same task with respect to a common resource includes: Numbering the multiple warps based on the original code of the same task according to the exit order of the multiple warps related to the same task; and Annotating the usage status of each warp in the multiple warps for the common resource, so as to establish a common usage connection relationship for the multiple warps with respect to the common resource.

4. The method according to claim 3, wherein Controlling, at least based on the established common usage connection relationship, the exit time of the current warp of the current task among the multiple tasks includes: In response to determining that each warp of the current task has completed the usage of the corresponding common resource, notifying other warps that have a common usage connection relationship with respect to the corresponding common resource.

5. The method according to claim 3, wherein Controlling, at least based on the established common usage connection relationship, the exit time of the current warp of the current task among the multiple tasks includes: In response to determining that each warp of the current task has completed the usage of the corresponding common resource, notifying other warps of the current task.

6. The method according to claim 4, characterized in that Notifying other warps that have a common usage connection relationship with respect to the corresponding common resource includes: In response to determining that each warp of the current task has completed the usage of the corresponding common resource, sending indication information for indicating the completion of the usage of the corresponding common resource to other warps that have a common usage connection relationship with respect to the corresponding common resource.

7. The method according to claim 6, characterized in that Controlling, at least based on the established common usage connection relationship, the exit time of the current warp of the current task among the multiple tasks further includes: In response to determining that the current warp of the current task has completed the current task and has received the indication information for indicating the completion of the usage of the corresponding common resource sent by all other warps that have a common usage connection relationship, allowing the current warp of the current task to exit; and In response to determining that the current warp of the current task has not completed the current task or has not received the indication information for indicating the completion of the usage of the corresponding common resource sent by all other warps that have a common usage connection relationship, not allowing the current warp of the current task to exit.

8. The method according to claim 1, characterized in that Establishing a common usage connection relationship for multiple warps related to the same task with respect to a common resource includes: Confirming whether a signal indicating that the concurrency mode switch has been turned on is detected; In response to confirming the detection of a signal indicating that the concurrent mode switch has been turned on, establish a connection relationship for the common use of a plurality of warps related to the same task for a common resource; In response to confirming that the signal indicating that the concurrent mode switch has been turned on is not detected, send an instruction to turn on the concurrent mode switch so that each computing unit in the concurrent mode can concurrently execute two warp groups.

9. The method according to claim 1, characterized in that Starting the corresponding warp of the next task includes: In response to confirming that the current computing unit in the concurrent mode needs to receive the second warp group when the first warp group has not completed the task, confirm whether the head warp in the first warp group has exited; In response to confirming that the head warp in the first warp group has exited, allow the head warp in the second warp group to be dispatched to the current computing unit for execution; Confirm whether the non-head warps in the first warp group have exited; and In response to confirming that the non-head warps in the first warp group have exited, allow the non-head warps in the second warp group to be dispatched to the current computing unit for execution.

10. A computing device, characterized in that, Comprising: At least one processor; And A memory communicatively connected to the at least one processor; Wherein The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method according to any one of claims 1-9.

11. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium, and when the computer program is executed by a machine, it executes the method according to any one of claims 1-9.

12. A computer program product, characterized in that, Comprising a computer program, and when the computer program is executed by a machine, it executes the method according to any one of claims 1-9.

Citation Information

Patent Citations

  • Kernel function transmission method, device and equipment

    CN114493980A

  • Thread bundle execution method and device, computer equipment, readable storage medium and program product

    CN118916098A

  • Instruction scheduling device and method, processor, electronic equipment and storage medium

    CN120125419A

  • System and method of arbitrating access of threads to shared resources within a data processing system

    US20070101333A1

Cited By

  • Method for issuing tasks about kernel functions, artificial intelligence chip, computing device, medium and program product

    CN121210146A