A control method, device, and medium for RGPU task scheduling

By obtaining the target call number of Cores in RGPU task scheduling and reading the current Core usage status, dynamically scheduling Core resources, the greedy occupation problem in RGPU task scheduling is solved, and efficient non-greedy task scheduling is achieved.

CN115063285BActive Publication Date: 2025-06-27LANGCHAO ELECTRONIC INFORMATION IND CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210764565.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-30
Publication Date
2025-06-27
Estimated Expiration
2042-06-30

AI Technical Summary

Technical Problem

In the prior art, RGPUs are prone to greedy task scheduling during task scheduling, occupying all Cores, resulting in being unable to serve multiple users or applications, and non-greedy task scheduling requires human intervention and is inefficient.

Method used

By obtaining the relevant data of the task to be executed and the preset number of target calls Cores, after receiving the execution instruction, the current usage status of each Core is read from the register to determine whether the current number of idle Cores is greater than or equal to the number of target calls Cores. If so, the corresponding Core will be called to execute the task. If otherwise, the call will be terminated and an error message will be returned.

Benefits of technology

It realizes non-greedy scheduling of RGPU tasks without human intervention, avoids occupancy of all Cores, improves task scheduling efficiency, and can serve multiple users or applications at the same time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115063285B_ABST
    Figure CN115063285B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of task scheduling, and discloses a control method, device and medium for RGPU task scheduling, including: obtaining relevant data of a task to be executed and a preset target number of cores to be called, reading the current usage status of each core from a register, and determining whether the current number of idle cores is greater than or equal to the target number of cores to be called according to the target number of cores to be called and the current usage status of each core. If so, determining a target core ID corresponding to the target core to be called, and calling a target number of target cores according to the target core ID to execute the task to be executed. Thus, by obtaining the target number of cores to be called and calling the corresponding number of cores according to the target number of cores to be called when executing the task to be executed, manual intervention is not required, and it is also possible to avoid occupying all cores and causing inability to serve multiple users or applications.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of task scheduling, and in particular, to a control method, device, and medium for RGPU task scheduling. Background Art

[0002] With the continuous development of heterogeneous computing, it has become crucial to reasonably and effectively solve the problem of task allocation on heterogeneous computing devices. The software architecture and hardware architecture of common commercial Graphics Processing Units (GPUs) are both closed-source. Therefore, a general-purpose GPU, namely RGPU (RISC-V-based GPU), is realized by extending the functions of the RISC-V open-source processor based on the RISC-V open-source processor.

[0003] Inside an RGPU, the physical-level logic responsible for computing is divided into three layers: Core, Warp, and Thread. There may be multiple Cores inside an RGPU, multiple Warps inside a Core, and multiple Threads inside a Warp. Among them, the minimum scheduling unit of both commercial GPUs and RGPU is Warp.

[0004] When an RGPU performs task scheduling, when the parallel processing ability of a single RGPU is less than the total number of computing tasks, the computing tasks will definitely occupy all the Cores inside the RGPU. Among them, the parallel processing ability of a single RGPU is the product of the number of Cores, Warps, and Threads inside the RGPU. Such a task scheduling method is called greedy task scheduling. Greedy task scheduling causes a single RGPU to occupy all the Cores in the RGPU in most cases, thereby resulting in the inability to serve multiple users or multiple applications.

[0005] Currently, the scheduling strategies and management strategies of non-greedy task scheduling are both based on offline analysis, that is, the final non-greedy task scheduling is achieved by offline analyzing the task information submitted by users and the hardware configuration information of the RGPU. Such a method requires manual intervention, has low efficiency, and lacks flexibility.

[0006] Therefore, it is an urgent problem for those skilled in the art to achieve non-greedy task scheduling of RGPU without manual intervention and improve task scheduling efficiency. Summary of the Invention

[0007] The purpose of this application is to provide a control method, device, and medium for RGPU task scheduling, which can achieve non-greedy scheduling of RGPU tasks without manual intervention, and avoid occupying all the Cores in the RGPU, resulting in the inability to serve multiple users or multiple applications.

[0008] To solve the above technical problems, the present application provides a control method for RGPU task scheduling, including:

[0009] Obtain relevant data of the task to be executed and the preset target number of cores to be called;

[0010] After receiving the execution instruction, read the current usage status of each core from the register;

[0011] According to the target number of cores to be called and the current usage status of each core, determine whether the current number of idle cores is greater than or equal to the target number of cores to be called;

[0012] If so, determine the target core ID corresponding to the target core to be called, and call the target number of target cores according to the target core ID to execute the task to be executed;

[0013] If not, end the call and return a call error message.

[0014] Preferably, after determining the target core ID corresponding to the target core to be called and calling the target number of target cores according to the target core ID to execute the task to be executed, it further includes:

[0015] Update the current usage status corresponding to each target core in the register.

[0016] Preferably, after calling the target number of target cores, it further includes:

[0017] Store the maximum value of the core IDs corresponding to each currently called target core.

[0018] Preferably, executing the task to be executed includes:

[0019] Determine the number of tasks to be executed corresponding to each target core;

[0020] According to the number of tasks to be executed corresponding to each target core, call the warps corresponding to each target core to execute the task to be executed to obtain an execution result.

[0021] Preferably, the calling the warps corresponding to each target core to execute the task to be executed to obtain an execution result includes:

[0022] Determine the number of threads in each warp;

[0023] According to the number of threads in each warp and the number of warps corresponding to the target core, determine the target number of times each warp needs to execute the corresponding task to be executed.

[0024] Call each of the said Warps to execute the to-be-executed tasks for the corresponding target number of times to obtain an execution result.

[0025] Preferably, the relevant data of the to-be-executed tasks includes at least the initial data for calculation, and the address and length corresponding to the initial data.

[0026] Preferably, after executing the to-be-executed tasks, it further includes:

[0027] Release the said target Core, and update the usage status corresponding to each of the said target Cores in the register.

[0028] To solve the above technical problems, the present application also provides a control device for RGPU task scheduling, including:

[0029] An acquisition module, configured to acquire the relevant data of the to-be-executed tasks and the preset target number of calling Cores;

[0030] A reading module, configured to read the current usage status of each Core from the register after receiving an execution instruction;

[0031] A processing module, configured to determine whether the current number of idle Cores is greater than or equal to the target number of calling Cores according to the target number of calling Cores and the current usage status of each Core;

[0032] If so, determine the target Core ID corresponding to the target calling Core, and call the target number of target Cores according to the target Core ID to execute the to-be-executed tasks;

[0033] If not, end the call and return a call error message.

[0034] An update module, configured to update the current usage status corresponding to each of the said target Cores in the register.

[0035] A storage module, configured to store the maximum value of the Core IDs corresponding to each of the currently called target Cores.

[0036] A determination module, configured to determine the number of to-be-executed tasks corresponding to each of the said target Cores;

[0037] A call module, configured to call the Warps corresponding to each of the said target Cores to execute the to-be-executed tasks according to the number of to-be-executed tasks corresponding to each of the said target Cores to obtain an execution result.

[0038] Among them, the calling the Warps corresponding to each of the said target Cores to execute the to-be-executed tasks to obtain an execution result includes:

[0039] Determine the number of threads in each of the Warps;

[0040] Based on the number of threads in each of the Warps and the number of Warps corresponding to the target Core, determine the target number of times each of the Warps needs to execute the corresponding task to be executed;

[0041] Invoke each of the Warps to execute the task to be executed the corresponding target number of times to obtain an execution result.

[0042] A release module, used to release the target Core and update the usage status corresponding to each of the target Cores in the register.

[0043] To solve the above technical problems, the present application also provides a control device for RGPU task scheduling, including a memory for storing a computer program;

[0044] A processor, used to implement the steps of the control method for RGPU task scheduling when executing the computer program.

[0045] To solve the above technical problems, the present application also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the control method for RGPU task scheduling are implemented.

[0046] A control method for RGPU task scheduling provided by the present invention includes: obtaining relevant data of a task to be executed and a preset target number of calling Cores, after receiving an execution instruction, reading the current usage status of each Core from a register, and based on the target number of calling Cores and the current usage status of each Core, determining whether the current number of idle Cores is greater than or equal to the target number of calling Cores. If it is greater than or equal to the target number of calling Cores, determine the target Core ID corresponding to the target number of calling Cores, and call the target number of target Cores according to the target Core ID to execute the task to be executed. If it is less than the target number of calling Cores, end the call and return a call error message. It can be seen that in the technical solution provided by the present application, when obtaining relevant data of the task to be executed, the target number of calling Cores is obtained at the same time. When executing the task to be executed, the corresponding number of Cores can be retrieved for use according to the target number of calling Cores, without manual intervention, and it can also avoid the greedy task scheduling that occupies all Cores and causes the inability to serve multiple users or applications, improving the efficiency of RGPU task scheduling.

[0047] In addition, the present application also provides a control device and medium for RGPU task scheduling, corresponding to the above control method for RGPU task scheduling, with the same effect. Brief Description of the Drawings

[0048] To more clearly illustrate the embodiments of the present application, the accompanying drawings required for use in the embodiments will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other accompanying drawings can be obtained based on these drawings without creative efforts.

[0049] Figure 1 It is a flowchart of a control method for RGPU task scheduling provided by an embodiment of the present application;

[0050] Figure 2 It is a structural diagram of a control device for RGPU task scheduling provided by an embodiment of the present application;

[0051] Figure 3 It is a structural diagram of a control device for RGPU task scheduling provided by another embodiment of the present application. Detailed Embodiments

[0052] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the protection scope of the present application.

[0053] The core of the present application is to provide a control method, device and medium for RGPU task scheduling. When obtaining the relevant data of the task to be executed, the preset target number of cores to be called is obtained at the same time, and the corresponding number of cores is called according to the target number of cores to be called, so as to avoid occupying all cores and causing the inability to provide services for multiple users or applications.

[0054] To enable those skilled in the art to better understand the solution of the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0055] With the continuous development of heterogeneous computing, how to reasonably and effectively solve the problem of task allocation on heterogeneous computing devices has become crucial. The software architecture and hardware architecture of common commercial Graphics Processing Units (GPUs) are both closed-source. Therefore, a general-purpose GPU, namely RGPU (RISC-V-based GPU), is realized by extending the functions of the RISC-V open-source processor based on the RISC-V open-source processor.

[0056] Inside an RGPU, the physical-level logic responsible for computing is divided into three layers: Core, Warp, and Thread. There may be multiple Cores inside an RGPU, multiple Warps inside a Core, and multiple Threads inside a Warp. Among them, the minimum scheduling unit for both commercial GPUs and RGPU is Warp.

[0057] When an RGPU performs task scheduling, the number of ideal Cores cN in a single RGPU is cN = Q / (NT * NW), where Q is the number of tasks to be executed, cN is the number of ideal Cores, NT is the number of Threads in a Warp, and NW is the number of Warps in a Core. For example, when the number of tasks to be executed Q = 128, cN = Q / (NT * NW) = 128 / (4 * 4) = 8. At this time, each Core can complete all the computing tasks by executing once. When the actual number of Cores NC is less than cN, the number of Cores cn actually used for computing is cn = min{cN, NC}.

[0058] When the number of tasks to be executed Q can be divided evenly by the actual number of Cores NC, the number of tasks assigned to each Core is Q / NC. Otherwise, the total number of tasks that the current Core needs to execute is int(Q / NC) + Q - NC * int(Q / NC).

[0059] In fact, the total number of tasks that an RGPU can execute at one time is the product of the number of Cores, the number of Warps in each Core, and the number of Threads in each Warp. When performing task scheduling, the number of tasks to be executed is often greater than the total number of tasks that an RGPU can execute at one time. At this time, the computing tasks will definitely occupy all the Cores inside the RGPU, and even the Warps need to execute tasks repeatedly. Such a task scheduling method is called greedy task scheduling. Greedy task scheduling causes a single RGPU to occupy all the Cores in the RGPU in most cases, thus resulting in the inability to serve multiple users or multiple applications.

[0060] Currently, both the scheduling strategy and management strategy of non-greedy task scheduling are based on offline analysis, that is, the final non-greedy task scheduling is achieved by offline analyzing the task information submitted by users and the hardware configuration information of the RGPU. Such a method requires manual intervention, has low efficiency, and lacks flexibility.

[0061] To achieve non-greedy task scheduling for the RGPU without manual intervention, avoid occupying all Cores and causing the inability to serve multiple users or applications, and improve task scheduling efficiency, the embodiments of the present application provide a control method for RGPU task scheduling. When obtaining a task to be executed, the target number of Core calls is obtained at the same time, thereby limiting the number of Cores called to execute the current task, and thus achieving non-greedy scheduling.

[0062] Figure 1 It is a flowchart of a control method for RGPU task scheduling provided by the embodiments of the present application. As Figure 1 shown, the method includes:

[0063] S10: Obtain the relevant data of the task to be executed and the preset target number of Core calls.

[0064] When the RGPU executes a task, the source code is divided into two parts. One part runs on the HOST side, and the other part runs on the Device side. Among them, the source code on the HOST side is mainly responsible for generating or obtaining the initial data required for the computing task, as well as parameters such as the address and length corresponding to the initial data. That is, the relevant data of the task to be executed may include, but is not limited to, the initial data for computing, and the address and length corresponding to the initial data. In addition, it is responsible for transmitting the initial data required for the computing task and the program instruction data on the Device side to the Device side, and for controlling and starting the RGPU calculation, etc.

[0065] To avoid occupying all Cores in the RGPU, therefore, when the HOST side transmits the initial data required for the computing task, the preset target number of Core calls is obtained, and the target number of Core calls is transmitted to the Device side together with the initial data required for the computing task.

[0066] The source code on the Device side (i.e., the RGPU side) describes the specific tasks to be executed. When compiling to generate the final Device Program, the pre-compiled Device Parallel Runtime Lib will be linked with the source code on the Device side to generate the final Device Program, that is, the binary code finally executed on the Device side. This binary code contains the Device Parallel Runtime Lib, thereby realizing task scheduling.

[0067] S11: After receiving the execution instruction, read the current usage status of each Core from the register.

[0068] After obtaining the execution instruction, the current usage status of each Core is read sequentially from the least significant bit to the most significant bit of the register. It can be understood that the register (CSR_CORE_STATUS) in the RGPU can be used to store and flag whether each Core in the RGPU is occupied. For example, if the bit width of the register is 4 bits, each bit corresponds to the status of a Core. Among them, 0 is used to indicate that the Core is not occupied, and 1 is used to indicate that the Core is occupied. For example, 0101 means that the 1st and 3rd Cores are occupied.

[0069] It should be noted that each Core has a corresponding Core ID, which is used to represent the position of the Core in the memory CSR_CORE_STATUS. For example, when the Core ID is 1, it represents that the current Core is stored in the first bit of the memory. It should also be noted that the Core ID can be a number, a letter, or a character, and this application does not limit this. For the convenience of task allocation for each Core, numbers are preferably used.

[0070] Thus, the upper-layer function in the RGPU can directly read the value of the register CSR_CORE_STATUS, and then determine the number of actually schedulable Cores and the corresponding Core IDs in the current RGPU.

[0071] S12: According to the target number of called Cores and the current usage status of each Core, determine whether the current number of idle Cores is greater than or equal to the target number of called Cores. If so, go to step S13; if not, go to step S14.

[0072] S13: Determine the target Core ID corresponding to the target number of called Cores, and call the target number of target Cores according to the target Core ID to execute the task to be executed.

[0073] S14: End the call and return a call error message.

[0074] In a specific implementation, when executing the Device Program binary program, based on the obtained target number of cores to be called and the current usage status of each core read from the register, it is determined whether the number of currently idle cores is greater than or equal to the target number of cores to be called. If the number of currently idle cores is greater than or equal to the target number of cores to be called, it is determined that the number of currently idle cores can meet the target number of cores required for scheduling. Thus, the target core ID corresponding to the target cores to be called is determined, and the target number of target cores is called based on the target core ID to execute the task to be executed. If the number of currently idle cores is less than the target number of cores to be called, the call is ended, and an error message for the call is returned.

[0075] After calling the target number of cores to execute the task, the bit information corresponding to the target cores to be called in the register is updated. In addition, the maximum value of the core IDs corresponding to each target core during the call is stored, so that when calling the core next time, the bit information after this core ID can be obtained.

[0076] When executing the task to be executed, first determine the number of tasks to be executed by each target core, then determine the number of warps existing in each target core and the number of threads corresponding to each warp, and determine the target number of times each warp needs to execute the task based on the number of warps and the number of threads in each warp. Finally, based on this target number of times, the corresponding cores and warps are called to execute the task to be executed to obtain the execution result.

[0077] After executing the task to be executed, the resources of the target cores are released, and at the same time, the relevant information of each core in the register is updated, that is, the usage status corresponding to each target core is updated.

[0078] The control method for RGPU task scheduling provided by the embodiments of the present application includes: obtaining the relevant data of the task to be executed and the preset target number of cores to be called. After receiving the execution instruction, read the current usage status of each core from the register, and determine whether the current number of idle cores is greater than or equal to the target number of cores to be called according to the target number of cores to be called and the current usage status of each core. If it is greater than or equal to the target number of cores to be called, determine the target core ID corresponding to the target number of cores to be called, and call the target number of target cores according to the target core ID to execute the task to be executed. If it is less than the target number of cores to be called, end the call and return the call error message. It can be seen that in the technical solution provided by the present application, when obtaining the relevant data of the task to be executed, the target number of cores to be called is obtained at the same time. When executing the task to be executed, the corresponding number of cores can be called according to the target number of cores to be called without manual intervention, and the greedy task scheduling that occupies all cores can be avoided, resulting in the inability to serve multiple users or applications, and the efficiency of RGPU task scheduling is improved.

[0079] On the basis of the above embodiments, in order not to affect the next task from calling the idle cores in the RGPU, after executing the task to be executed, update the current usage status corresponding to each target core in the register.

[0080] In addition, in order for the next task to quickly call the idle cores in the RGPU, after calling the target number of target cores, store the maximum core ID Core ID max corresponding to each of the currently called target cores. Thus, when other cores are called, obtain the value after the register value corresponding to this core ID in order to obtain the idle cores. For the sake of easy understanding, an example will be given below.

[0081] For example, the number of target - called cores is 2. The values obtained by sequentially reading the current usage status of each core from the least - significant bit to the most - significant bit of the register are 0000101. According to the principle from right to left, it is determined that the cores corresponding to the first bit and the third bit are occupied, that is, the cores with Core ID 1 and 3 are occupied, and the remaining cores are not occupied. According to the principle from right to left and the number of target - called cores, the cores corresponding to the second bit and the fourth bit should be obtained, that is, the cores with Core ID 2 and 4 are used as target cores for calling. After the call, the corresponding positions of the target cores in the register are set to 1, that is, the second bit and the fourth bit are set to 1, which is used to indicate that the current second bit and the fourth bit have been borrowed. At this time, Core ID max = 4 is stored. If there are other users or applications calling cores, they can directly call from the cores after Core ID max = 4.

[0082] The control method for RGPU task scheduling provided by the embodiments of the present application, after determining the target Core ID corresponding to the target - called core and calling the target number of target cores according to the target Core ID to execute the task to be executed, updates the current usage status of each target core in the register, thereby avoiding affecting the free cores in the RGPU for the next task call. In addition, after calling the target number of target cores, the maximum value of the Core ID corresponding to each current - called target core is stored, which is convenient for the next task to quickly call the free cores in the RGPU.

[0083] In a specific implementation, when calling the target number of target cores according to the target Core ID to execute the task to be executed, the number of tasks to be executed assigned to each target core is determined through the Device Program binary program, that is, the number of tasks to be executed corresponding to each target core is determined.

[0084] Specifically, the number of tasks to be processed cT assigned to each target core is cT = Q / N, where Q is the total number of tasks to be executed and N is the number of target - called cores. That is, when it is determined that there are cores in the memory that meet the number of target - called cores, regardless of the Core ID, the number of tasks cT assigned to each core is Q / N. For example, when Q = 128 and N = 2, the number of tasks to be processed cT assigned to each target core is cT = 128 / 2 = 64.

[0085] After determining the number of tasks to be executed corresponding to each target Core, determine the number of threads in each Warp corresponding to the target Core. Based on the number of threads in each Warp and the number of Warps corresponding to the target Core, determine the target number of times each Warp needs to execute the corresponding tasks to be executed.

[0086] Specifically, the number of tasks to be executed allocated to each current target Core is cT, and the number of threads in each Warp corresponding to the target Core is NT. If each thread only executes once, then NQ = int(cT / NT) Warps are required, where NQ is the number of Warps corresponding when each thread only executes the task once. For example, when cT = 32 and NT = 4, if each thread only executes once, then NQ = 8 Warps are required.

[0087] However, the number of Warps corresponding to each target Core may be less than NQ, then the number of times each Warp corresponding to the target Core needs to execute is int(cT / NT) / NW, where NW is the actual number of Warps in the target Core, and NT is the number of threads in each Warp corresponding to the target Core. It should be noted that the Warps with Warp ID less than int(cT / NT) % NW need to execute one more time, where Warp ID is used to mark each Warp, and Warp ID is a natural number greater than 0. For example, when cT = 33, NT = 4, and NW = 4, each Warp executes 2 times. However, the Warps with Warp ID less than int(cT / NT) % NW need to execute one more time to complete the tasks to be executed allocated to the target Core.

[0088] Furthermore, after determining the target number of times each Warp needs to execute the corresponding tasks to be executed, call each Warp to execute the tasks to be executed the corresponding number of times to obtain the execution results.

[0089] After the task is executed, it is also necessary to determine whether there is a remaining task rT, that is, whether there is a task to be executed that has not been processed. Among them, rT = cT - int(cT / NT)*NT. If the remaining task rT is not enough to fill all the threads NT in a single Warp, that is, all the threads NT in a Warp cannot be allocated at least 1 task, then the function spawn_kernel_remaining_callback of the runtime parallel library is called for processing. The control method for RGPU task scheduling provided by the embodiments of the present application, after obtaining the target Core based on the obtained target call Core number, determines the number of tasks to be executed corresponding to each target Core and the number of threads in each Warp, and determines the target number of times each Warp needs to execute the corresponding task to be executed according to the number of threads in each Warp and the number of Warps corresponding to the target Core. Finally, each Warp is called to execute the task to be executed the corresponding number of times to obtain the execution result. Thus, based on the target call Core number, the number of Cores called by the currently executed task is controlled, avoiding occupying all Cores and causing inability to serve other users or applications, thereby improving the efficiency of RGPU task scheduling.

[0090] It can be understood that when the HOST side executes the HOST Program, it will obtain or generate the initial data for calculating the specific task, as well as the address and length corresponding to the initial data (that is, the starting address, length, etc. of the initial data in the RGPU), and other data. That is to say, the relevant data of the task to be executed obtained by the HOST side may include, but is not limited to, the initial data for calculation, and the address and length corresponding to the initial data.

[0091] After executing the task according to the relevant data of the task to be executed and the target call Core number, release the target Core, and update the usage status corresponding to each target Core in the register, so as to avoid errors when the next task calls the target Core, reducing the efficiency and reliability of RGPU task scheduling.

[0092] The control method for RGPU task scheduling provided by the embodiments of the present application, after executing the task to be executed, releases the target Core, and updates the usage status corresponding to each target Core in the register, so as to avoid errors when the next task calls the target Core, improving the efficiency and reliability of RGPU task scheduling.

[0093] In the above embodiments, the control method for RGPU task scheduling is described in detail. The present application also provides an embodiment corresponding to a control device for RGPU task scheduling. It should be noted that the present application describes the embodiments of the device part from two perspectives, one is from the perspective of functional modules, and the other is from the perspective of hardware structure.

[0094] Figure 2 The structure diagram of a control device for RGPU task scheduling provided by an embodiment of this application is as follows Figure 2 As shown, the device includes:

[0095] An acquisition module 10, configured to acquire relevant data of a task to be executed and a preset target number of cores to be called.

[0096] A reading module 11, configured to read the current usage status of each core from a register after receiving an execution instruction.

[0097] A processing module 12, configured to determine whether the current number of idle cores is greater than or equal to the target number of cores to be called according to the target number of cores to be called and the current usage status of each core. If so, determine the target core ID corresponding to the target number of cores to be called, and call the target number of target cores according to the target core ID to execute the task to be executed. If not, end the call and return a call error message.

[0098] Preferably, the control device for RGPU task scheduling provided by an embodiment of this application further includes:

[0099] A determination module, configured to determine the number of tasks to be executed corresponding to each target core.

[0100] A call module, configured to call the Warp corresponding to each target core to execute the task to be executed according to the number of tasks to be executed corresponding to each target core to obtain an execution result.

[0101] Among them, calling the Warp corresponding to each target core to execute the task to be executed to obtain an execution result includes:

[0102] Determine the number of threads in each Warp.

[0103] Determine the target number of times each Warp needs to execute the corresponding task to be executed according to the number of threads in each Warp and the number of Warps corresponding to the target core.

[0104] Call each Warp to execute the task to be executed the corresponding target number of times to obtain an execution result.

[0105] A release module, configured to release the target core and update the usage status corresponding to each target core in the register.

[0106] Since the embodiments in the device part correspond to the embodiments in the method part, for the embodiments in the device part, please refer to the description of the embodiments in the method part, which will not be elaborated here.

[0107] The control device for RGPU task scheduling provided by the embodiments of the present application includes: obtaining the relevant data of the task to be executed and the preset target number of Core to be called. After receiving the execution instruction, read the current usage status of each Core from the register, and determine whether the current number of idle Cores is greater than or equal to the target number of Core to be called according to the target number of Core to be called and the current usage status of each Core. If it is greater than or equal to the target number of Core to be called, determine the target Core ID corresponding to the target number of Core to be called, and call the target number of target Cores according to the target Core ID to execute the task to be executed. If it is less than the target number of Core to be called, end the call and return the call error message. It can be seen that in the technical solution provided by the present application, when obtaining the relevant data of the task to be executed, the target number of Core to be called is obtained at the same time. When executing the task to be executed, the corresponding number of Cores can be called according to the target number of Core to be called without manual intervention, and it can also avoid the greedy task scheduling that occupies all Cores and cannot serve multiple users or applications, improving the efficiency of RGPU task scheduling.

[0108] Figure 3 The structural diagram of a control device for RGPU task scheduling provided by another embodiment of the present application is as Figure 3 shown. The control device for RGPU task scheduling includes: a memory 20 for storing computer programs;

[0109] a processor 21 for implementing the steps of the control method for RGPU task scheduling as mentioned in the above embodiments when executing the computer programs.

[0110] The control device for RGPU task scheduling provided in this embodiment may include but is not limited to smart phones, tablet computers, laptop computers or desktop computers, etc.

[0111] Among them, the processor 21 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 21 may be implemented in at least one hardware form of a digital signal processor (DSP), a field-programmable gate array (FPGA), or a programmable logic array (PLA). The processor 21 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the wake state, also known as the central processing unit (CPU); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 21 may be integrated with a graphics processing unit (GPU), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 21 may further include an artificial intelligence (AI) processor, and the AI processor is used to process computational operations related to machine learning.

[0112] The memory 20 may include one or more computer-readable storage media, and the computer-readable storage media may be non-transitory. The memory 20 may further include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash storage devices. In this embodiment, the memory 20 is at least used to store the following computer program 201. After the computer program is loaded and executed by the processor 21, it can implement the relevant steps of the control method for RGPU task scheduling disclosed in any of the foregoing embodiments. In addition, the resources stored in the memory 20 may further include an operating system 202 and data 203, etc., and the storage method may be temporary storage or permanent storage. Among them, the operating system 202 may include Windows, Unix, Linux, etc. The data 203 may include, but is not limited to, the relevant data involved in the control method for RGPU task scheduling.

[0113] In some embodiments, the control device for RGPU task scheduling may further include a display screen 22, an input / output interface 23, a communication interface 24, a power supply 25, and a communication bus 26.

[0114] Those skilled in the art can understand that Figure 3 the structure shown in

[0115] The control device for RGPU task scheduling provided by the embodiment of the present application includes a memory and a processor. When the processor executes the program stored in the memory, it can implement the following method: the control method for RGPU task scheduling.

[0116] When the control device for RGPU task scheduling provided by the embodiment of the present application obtains the relevant data of the task to be executed, it simultaneously obtains the number of target call Cores. When executing the task to be executed, it can call the corresponding number of Cores for use according to the number of target call Cores, without manual intervention, and can also avoid the greedy task scheduling that occupies all Cores and causes the inability to serve multiple users or applications, improving the efficiency of RGPU task scheduling.

[0117] Finally, the present application also provides an embodiment corresponding to a computer-readable storage medium. A computer program is stored on the computer-readable storage medium, and when the computer program is executed by the processor, it implements the steps recorded in the above method embodiments.

[0118] It can be understood that if the method in the above embodiments is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and executes all or part of the steps of the methods described in the various embodiments of the present application. The aforementioned storage media include: USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs, etc., which can store program codes.

[0119] The computer-readable storage medium provided by the embodiments of the present application includes: obtaining relevant data of a task to be executed and a preset target number of cores to be called. After receiving an execution instruction, reading the current usage status of each core from a register, and determining whether the current number of idle cores is greater than or equal to the target number of cores to be called according to the target number of cores to be called and the current usage status of each core. If it is greater than or equal to the target number of cores to be called, determining the target core ID corresponding to the target number of cores to be called, and calling the target number of target cores according to the target core ID to execute the task to be executed. If it is less than the target number of cores to be called, ending the call and returning a call error message. It can be seen that in the technical solution provided by the present application, when obtaining relevant data of a task to be executed, the target number of cores to be called is obtained at the same time. When executing the task to be executed, the corresponding number of cores can be called according to the target number of cores to be called for use, without manual intervention, and it can also avoid the greedy task scheduling that occupies all cores and causes the inability to serve multiple users or applications, improving the efficiency of RGPU task scheduling.

[0120] The above has introduced in detail a control method, device, and medium for RGPU task scheduling provided by the present application. The embodiments in the specification are described in a progressive manner, and the key points of each embodiment are the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the description of the method part. It should be noted that for those of ordinary skill in the art in the technical field of the present application, without departing from the principle of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the protection scope of the claims of the present application.

[0121] It should also be noted that in this specification, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover a non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the phrase "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the said element.

Claims

1. A control method for RGPU task scheduling, characterized in that It includes: Obtain the relevant data of the task to be executed and the preset target number of Core to be called; After receiving the execution instruction, read the current usage status of each Core from the register; According to the target number of Core to be called and the current usage status of each Core, determine whether the current number of idle Cores is greater than or equal to the target number of Core to be called; If so, determine the target Core ID corresponding to the target Core to be called, and call the target number of target Cores according to the target Core ID to execute the task to be executed; If not, end the call and return a call error message; Wherein, the register is a register in the RGPU; the Core is a Core in the RGPU; Wherein, executing the task to be executed includes: Determine the number of tasks to be executed corresponding to each of the target Cores; According to the number of tasks to be executed corresponding to each of the target Cores, call the Warp corresponding to each of the target Cores to execute the task to be executed to obtain an execution result; Wherein, the calling the Warp corresponding to each of the target Cores to execute the task to be executed to obtain an execution result includes: Determine the number of threads in each of the Warps; According to the number of threads in each of the Warps and the number of Warps corresponding to the target Core, determine the target number of times each of the Warps needs to execute the corresponding task to be executed; Call each of the Warps to execute the task to be executed the target number of times to obtain an execution result.

2. The control method for RGPU task scheduling according to claim 1, wherein After determining the target Core ID corresponding to the target Core to be called and calling the target number of target Cores according to the target Core ID to execute the task to be executed, it further includes: Update the current usage status corresponding to each of the target Cores in the register.

3. The control method for RGPU task scheduling according to claim 2, wherein After calling the target number of target Cores, it further includes: Store the maximum value of the Core ID corresponding to each of the currently called target Cores.

4. The control method for RGPU task scheduling according to claim 1, characterized in that, The relevant data of the task to be executed at least includes the initial data for calculation, as well as the address and length corresponding to the initial data.

5. The control method for RGPU task scheduling according to any one of claims 1 to 4, characterized in that, After executing the task to be executed, it further includes: Release the target Core and update the usage status corresponding to each of the target Cores in the register.

6. A control device for RGPU task scheduling, characterized in that, It includes: An acquisition module, configured to obtain the relevant data of the task to be executed and the preset target number of Core to be called; A reading module, configured to read the current usage status of each Core from the register after receiving the execution instruction; A processing module, configured to determine whether the current number of idle Cores is greater than or equal to the target number of Core to be called according to the target number of Core to be called and the current usage status of each Core; If so, determine the target Core ID corresponding to the target Core to be called, and call the target number of target Cores according to the target Core ID to execute the task to be executed; If not, end the call and return a call error message; Wherein, the register is a register in the RGPU; the Core is a Core in the RGPU; Wherein, it further includes: A determination module, configured to determine the number of tasks to be executed corresponding to each of the target Cores; An invocation module, configured to, according to the number of tasks to be executed corresponding to each of the target Cores, invoke the Warps corresponding to each of the target Cores to execute the tasks to be executed to obtain an execution result; Wherein, the invocation module is configured to: determine the number of threads in each of the Warps; according to the number of threads in each of the Warps and the number of Warps corresponding to the target Core, determine the target number of times for each of the Warps to execute the corresponding tasks to be executed; and invoke each of the Warps to execute the tasks to be executed the corresponding target number of times to obtain an execution result.

7. A control device for RGPU task scheduling, characterized in that, It includes a memory for storing a computer program; A processor, configured to implement the steps of the control method for RGPU task scheduling according to any one of claims 1 to 5 when executing the computer program.

8. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, the steps of the control method for RGPU task scheduling according to any one of claims 1 to 5 are implemented.

Citation Information

Patent Citations

  • Method, device and equipment for dispatching GPU resource and computer readable storage medium

    CN108363623A