Task processing method and device and related equipment
By matching and dispatching to different computing units on the processor core according to the task type, the problem of idle computing units is solved, and resource utilization and task processing efficiency are improved.
Patent Information
- Application Number
- CN202311844559.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-28
- Publication Date
- 2025-07-01
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Some computing units on the processor core are idle for a long time, resulting in low resource utilization and inability to make full use of the processor core resources.
Through the task processing device, the task is scheduled to the corresponding queue according to the type of the task, so as to ensure that each computing unit can perform appropriate tasks.
It effectively avoids long-term idleness of computing units, improves resource utilization of processor cores, and realizes hardware-level concurrent execution of multiple tasks, thereby improving overall task processing efficiency.
Smart Images

Figure CN120234103A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of task processing, and in particular, to a task processing method, apparatus, and related devices. Background Art
[0002] Currently, parallel processing of tasks can effectively improve the processing efficiency of multiple tasks and has become the mainstream task processing method. Specifically, when running an application program, multiple threads can be used to call the resources in the processor core to execute the tasks related to the application program in parallel. Among them, the resources in the processor core can include various computing units, such as computing units for matrix calculation, computing units for vector calculation, computing units for scalar calculation, etc.
[0003] In actual application scenarios, the computing units on the processor core are used to execute tasks, and the tasks to be executed need to queue up in a queue waiting for scheduling. Usually, the tasks in the queue are scheduled to the processor core in the order in the queue, so as to use the corresponding computing units on the processor core to execute the task. In this way, there are often some computing units on the processor core that are not scheduled to tasks, resulting in the long-term idle state of the computing unit, and further causing the resources on the processor core not to be fully utilized. Summary of the Invention
[0004] This application provides a task processing method to improve the overall efficiency of a processor core in executing multiple tasks. In addition, this application also provides a task processing apparatus, a processor, a computing device, a computer-readable storage medium, and a computer program product.
[0005] In a first aspect, this application provides a task processing method, which is applied to a processor. The processor includes at least one processor core, and the at least one processor core includes a first processor core. Moreover, the first processor core includes a first computing unit and a second computing unit, and the first computing unit and the second computing unit are different types of computing units. This method can be executed by a task processing apparatus. During the process of processing tasks, the task processing apparatus obtains a first task from multiple tasks to be executed. For example, the first task that can be currently executed can be determined according to the execution dependencies between multiple tasks, and further obtains the type of the first task. The type can be, for example, a matrix calculation type, a vector calculation type, or a scalar calculation type, etc. Thus, when the type of the first task matches the type of the first computing unit, the first task is scheduled to the first queue corresponding to the first computing unit, and when the type of the first task matches the type of the second computing unit, the first task is scheduled to the second queue corresponding to the second computing unit.
[0006] Since during the process of task execution, the task processing device can schedule appropriate types of tasks for different types of computing units, multiple computing units on the processor core can execute multiple different types of tasks in parallel under the scheduling of the task processing device, thereby avoiding as much as possible that some computing units on the processor core are idle for a long time, and thus improving the resource utilization rate of the processor core; at the same time, it can achieve concurrent execution of multiple tasks at the hardware level, thereby improving the overall efficiency of the processor core in executing multiple tasks.
[0007] In a possible implementation manner, the task processing device may specifically be a scheduler running in the processor, that is, the task processing method may be executed by the scheduler. The scheduler is configured with a target queue, and the target queue stores multiple tasks to be executed. Thus, when obtaining the first task, specifically, the scheduler obtains the first task from the target queue, and when scheduling the first task, the scheduler may specifically store the first task in the first queue corresponding to the first computing unit when the type of the first task matches the type of the first computing unit, or store the first task in the second queue corresponding to the second computing unit when the type of the first task matches the type of the second computing unit. In this way, the scheduler can achieve global scheduling of tasks and improve the reliability of task scheduling.
[0008] In a possible implementation manner, the task processing device may specifically be the first computing unit in the first processor core, that is, the task processing method may be executed by the first computing unit. Thus, when obtaining the first task, the first computing unit may specifically obtain the first task from the multiple tasks to be executed after completing the tasks in the first queue, and when scheduling the first task, the first computing unit may specifically store the first task in the first queue corresponding to the first computing unit when the type of the first task matches the type of the first computing unit, or store the first task in the second queue corresponding to the second computing unit when the type of the first task matches the type of the second computing unit. In this way, the computing unit can achieve scheduling of tasks and improve the flexibility of task scheduling.
[0009] In a possible implementation manner, the multiple tasks to be executed are tasks respectively executed by multiple threads included in a process of an application. In this way, multiple threads can be used to achieve parallel execution of the multiple tasks, thereby effectively improving the processing efficiency for the multiple tasks.
[0010] In a possible implementation manner, the types of the first computing unit and the second computing unit include a matrix computing unit, a scalar computing unit, or a vector computing unit. That is, each computing unit may be any one of a matrix computing unit, a scalar computing unit, or a vector computing unit.
[0011] In a possible implementation, during the process of processing multiple tasks, the task processing device may further obtain a second task from the multiple tasks to be executed. The second task includes multiple operators. Taking the first operator and the second operator as an example. The task processing device may divide the second task into multiple subtasks, including a first subtask and a second subtask. Among them, the first subtask includes the first operator, and the second subtask includes the second operator. Thus, when scheduling tasks, the task processing device schedules the first subtask to the first queue corresponding to the first computing unit. At this time, the type of the first subtask matches the type of the first computing unit. At the same time, the second subtask is scheduled to the second queue corresponding to the second computing unit. At this time, the type of the second subtask matches the type of the second computing unit. In this way, the task processing device can divide the task into multiple subtasks and execute the multiple subtasks in parallel through multiple computing units, thereby improving the overall execution efficiency of the second task.
[0012] In a possible implementation, the task processing device may specifically be the first computing unit in the first processor core, and the first task has been scheduled to the first queue corresponding to the first computing unit. During the process of processing the task, when the first computing unit cannot execute the first task in the first queue, for example, the first computing unit is a computing unit for matrix calculation, while the first task is actually a vector calculation task, etc. At this time, the first computing unit can determine the actual type to which the first task belongs, and when the actual type to which the first task belongs matches the type of the second computing unit, the first computing unit schedules the first task to the second queue corresponding to the second computing unit so that the second computing unit can execute the first task. In this way, when the first computing unit cannot execute the first task (that is, a task scheduling error occurs), the first task can be promptly scheduled to other computing units for execution, thereby improving the execution efficiency of the first task.
[0013] In a possible implementation, in addition to the first processor core, the processor may further include a second processor core, and the first processor core and the second processor core may be processor cores of the same type, or the first processor core and the second processor core may be processor cores of different types, such as the first processor core is a large core and the second processor core is a small core, etc.
[0014] Second aspect, the present application provides a task processing device, which is applied to a processor. The first processor core in the processor includes a first computing unit and a second computing unit, and the first computing unit and the second computing unit are different types of computing units. The task processing device includes: an obtaining module, configured to obtain a first task from a plurality of tasks to be executed; obtain the type of the first task; a scheduling module, configured to schedule the first task to a first queue corresponding to the first computing unit when the type of the first task matches the type of the first computing unit; and schedule the first task to a second queue corresponding to the second computing unit when the type of the first task matches the type of the second computing unit.
[0015] In a possible implementation manner, the plurality of tasks to be executed are tasks respectively executed by a plurality of threads included in a process of an application.
[0016] In a possible implementation manner, the types of the first computing unit and the second computing unit include a matrix computing unit, a scalar computing unit, or a vector computing unit.
[0017] In a possible implementation manner, the obtaining module is further configured to obtain a second task from the plurality of tasks to be executed, and the second task includes a first operator and a second operator. The task processing device further includes: a partitioning module, configured to partition the second task into a plurality of subtasks, the plurality of subtasks including a first subtask and a second subtask, the first subtask including the first operator, and the second subtask including the second operator; the scheduling module is further configured to schedule the first subtask to the first queue, where the type of the first subtask matches the type of the first computing unit; and schedule the second subtask to the second queue, where the type of the second subtask matches the type of the second computing unit.
[0018] In a possible implementation manner, the task processing device is specifically the first computing unit, and if the first task is scheduled to the first queue corresponding to the first computing unit, the task processing device further includes: a determining module, configured to determine the actual type to which the first task belongs when the first computing unit cannot execute the first task in the first queue; at this time, the scheduling module is further configured to schedule the first task to the second queue when the actual type to which the first task belongs matches the type of the second computing unit.
[0019] In a possible implementation manner, the processor further includes a second processor core; the first processor core and the second processor core are processor cores of the same type, or the first processor core and the second processor core are processor cores of different types.
[0020] Third aspect, the present application provides a processor, which is configured to execute the task processing method in the first aspect or any one of the implementation manners of the first aspect.
[0021] Fourthly, the present application provides a computing device, which includes a processor and a memory. The processor and the memory communicate with each other. The processor is configured to execute instructions stored in the memory, so that the computing device executes the task processing method as described in the first aspect or any implementation manner of the first aspect. It should be noted that the memory may be integrated into the processor or independent of the processor. The computing device may further include a bus. Among them, the processor is connected to the memory through the bus. Among them, the memory may include a readable memory and a random access memory.
[0022] Fifthly, the present application provides a computer-readable storage medium, in which instructions are stored. When the instructions are run on a computing device, the computing device is caused to execute the operation steps of the task processing method described in the above-mentioned first aspect or any implementation manner of the first aspect.
[0023] Sixthly, the present application provides a computer program product containing instructions. When the computer program product is run on a computing device, the computing device is caused to execute the operation steps of the task processing method described in the above-mentioned first aspect or any implementation manner of the first aspect.
[0024] Based on the implementation manners provided in the above aspects of the present application, further combinations can be made to provide more implementation manners. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1 It is a schematic structural diagram of an exemplary data processing system provided by the present application;
[0026] Figure 2 It is a schematic diagram of creating multiple workers for multiple computing units on a single processor core provided by the present application;
[0027] Figure 3 It is another schematic diagram of creating multiple workers for multiple computing units on a single processor core provided by the present application;
[0028] Figure 4 It is a schematic flowchart of a task processing method provided by the present application;
[0029] Figure 5 It is a schematic diagram of the execution dependency between multiple tasks provided by the present application;
[0030] Figure 6 It is a schematic diagram of executing tasks before and after considering the task and the computing type of the computing unit provided by the present application;
[0031] Figure 7 It is a schematic flowchart of another task processing method provided by the present application;
[0032] Figure 8 Structural schematic diagram of a task processing device provided for this application;
[0033] Figure 9 Hardware structural schematic diagram of a computing device provided for this application. Specific implementation manners
[0034] Terms such as "first", "second", etc. in the description and claims of this application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that such terms can be interchanged under appropriate circumstances, which is only a way of distinguishing objects with the same attributes when describing the embodiments of this application.
[0035] Next, the technical solutions in this application will be described in conjunction with the drawings provided in this application.
[0036] Refer to Figure 1 , which shows a structural schematic diagram of a data processing system. As Figure 1 shown, the data processing system 10 includes an application layer 101, a hardware layer 102, and a task processing device 200.
[0037] Among them, the application layer 101 includes at least one application, such as an image recognition application, etc., Figure 1 and an application 1 is taken as an example for illustrative purposes. And, the application 1 can include at least one programming model, Figure 1 and programming models 1 to 3 are taken as examples for illustration. Exemplarily, the programming model can specifically be a message passing interface (MPI) model, a shared memory parallel programming (Open Multi-Processing, OpenMP) model, or a SYCL model (which is a high-level programming model of the Open Computing Language (OpenCL)), etc., or can be other types of programming models.
[0038] The hardware layer 102 includes multiple processor cores, and each processor core can execute thread tasks. Moreover, the multiple processor cores can be used to execute multiple thread tasks in parallel. The multiple processor cores in the hardware layer 102 can be processor cores of the same type; or, the multiple processor cores can be processor cores of different types. For example, some processor cores are large cores with high processing performance, and some other processor cores are small cores with low processing performance. Among them, each processor core can include multiple different types of computing units. For example, it can include a computing unit for scalar calculation, a computing unit for vector calculation, a computing unit for matrix calculation, etc. Each computing unit is a hardware unit with computing ability on the processor core. In practical applications, this computing unit can also be referred to as a computing engine implemented based on hardware.
[0039] Furthermore, the hardware layer 102 can also include other devices, such as Figure 1 the network interface controller (NIC), memory, accelerator shown, or it can be other types of devices. Exemplarily, the memory can be direct memory access (DMA), etc. The accelerator can be a graphics processing unit (GPU), etc.
[0040] The task processing device 200 can include multiple workers, and the multiple workers are used to schedule the processor cores in the hardware layer 102 to process multiple thread tasks in parallel. Among them, the worker can be implemented through a thread. For example, the task processing device 200 can create multiple threads for Application 1 to execute multiple thread tasks. In practical applications, the worker can also be implemented through other means such as a process, and this is not limited. Exemplarily, the task processing device 200 can be implemented through software. For example, the task processing device 200 can be specifically referred to as a runtime scheduler, and this runtime scheduler can run on the processor core, etc.
[0041] During the running process of Application 1, it can generate at least one task based on the user's operation on this Application 1, such as a data search task, etc. For each task, the programming model can edit the task into one or more process tasks, and there can be a dependency relationship between the multiple process tasks. For example, the execution of Process Task A among the multiple process tasks depends on the execution result of Process Task B (that is, Process Task B needs to be executed first). Then, the programming model can generate multiple computations Figure 1 , and each node in this computation Figure 1 represents a process task, and this computation Figure 1Directed edges between different nodes are used to indicate the dependency relationships between different process tasks. Next, for each process task, the orchestration model can edit the process task into multiple thread tasks, and there are dependency relationships between the multiple thread tasks. For example, the execution of thread task A among the multiple thread tasks depends on thread task B being executed first. Thus, the programming model can generate calculations based on the dependency relationships between the multiple thread tasks. Figure 2 In an actual application scenario, each process task can correspond to a calculation. Figure 2 In this way, the programming model can utilize the calculation. Figure 1 and the calculation. Figure 2 to achieve the orchestration of process tasks and thread tasks.
[0042] Based on the orchestration results of the programming model for process tasks and thread tasks (for example, it can be the above calculations. Figure 1 and the calculation. Figure 2 ), the task processing device can schedule the processor cores in the hardware layer 102 to execute in parallel the multiple thread tasks included in each process task. Among them, the programming model can output the multiple thread tasks; correspondingly, the task processing device 200 can add the multiple thread tasks to a queue, so that subsequently, according to the order of the thread tasks in the queue, each thread task can be scheduled to the corresponding computing unit for execution.
[0043] Specifically, the task processing device 200 can create multiple threads for each of the multiple types of computing units included in each processor core. Each thread can be a worker, used to schedule a type of computing unit on the processor core to execute the thread task, and different workers can schedule different computing units to execute different types of thread tasks.
[0044] For example, as Figure 2As shown, when a single processor core includes a computing unit 1 for scalar computation, a computing unit 2 for vector computation, a computing unit 3 for vector computation, and a computing unit 4 for matrix computation, the task processing device 200 can create 4 workers for each computing unit in this processor core, namely worker 1, worker 2, worker 3, and worker 4. Among them, worker 1 is used to schedule the computing unit 1 to process thread tasks of scalar computation type, such as data movement tasks, etc.; worker 2 and worker 3 are respectively used to schedule the computing unit 2 and the computing unit 3 to process thread tasks of vector computation type, such as data access tasks, etc.; worker 4 is used to schedule the computing unit 4 to process thread tasks of matrix computation type, such as convolution computation tasks, etc. Moreover, each computing unit can be configured with a queue. For example, the computing unit 1 can be configured with queue 1, and the computing unit 4 is configured with queue 4, etc. The queue of each computing unit is used to store the thread tasks that the computing unit needs to execute. Specifically, it can be that the worker utilizes the computing unit to execute the thread tasks in this queue.
[0045] Furthermore, in addition to various computing units and queues, the processor core can also include other hardware such as caches and registers. Thus, each worker can also use multiple hardware such as computing units, caches, and registers to process the same thread task simultaneously.
[0046] In Figure 2 the shown processor core, the computing units used by different workers can be independent of each other, that is, a single computing unit is only allowed to be used by one worker. Or, in Figure 3 the shown processor core, different workers can use multiple computing units, that is, a single computing unit is allowed to be used by multiple workers.
[0047] In this way, for each processor core in the hardware layer 102, the task processing device 200 can create multiple workers for this processor core.
[0048] After creating multiple workers, for the multiple thread tasks obtained through orchestration in the programming model, the task processing device 200 can schedule these multiple thread tasks to the computing units on one or more processor cores to execute these multiple thread tasks in parallel, so as to improve the execution efficiency of these multiple thread tasks. Among them, when the task processing device 200 schedules each thread task to the corresponding computing unit for execution, it can, according to the type of the thread task, schedule the thread task to the queue corresponding to the computing unit on the processor core used to execute the thread tasks of this type, so that the computing unit executes the thread tasks in this queue.
[0049] For example, the task processing device 200 can pre-create worker 1, worker 2, and worker 3 for the computing unit 1 for scalar calculation, the computing unit 2 for vector calculation, and the computing unit 3 for matrix calculation on the processor core A, and create worker 4, worker 5, and worker 6 for the computing unit 4 for scalar calculation, the computing unit 5 for vector calculation, and the computing unit 6 for matrix calculation on the processor core B. Among them, worker 1 and worker 4 are responsible for processing tasks of the scalar calculation type, worker 2 and worker 5 are responsible for processing tasks of the vector calculation type, and worker 3 and worker 6 are responsible for processing tasks of the matrix calculation type.
[0050] Suppose the multiple tasks to be executed include a first task. The first task can be, for example, a thread task generated by a programming model. Suppose the first task is a thread task of the matrix calculation type. The task processing device 200 can, according to the type of the first task, schedule the first task to the queue 3 corresponding to the computing unit 3 on the processor core A. Thus, worker 3 can use the computing unit 3 to execute the first task in queue 3. Alternatively, the task processing device 200 can, according to the type of the first task, schedule the first task to the queue 6 corresponding to the computing unit 6 on the processor core B. Thus, worker 6 can use the computing unit 6 to execute the first task in queue 6.
[0051] Similarly, for other tasks among the multiple tasks to be executed, the task processing device 200 can also, according to the types of other tasks, schedule other tasks to the queues corresponding to the computing units on the processor core A or the processor core B that are used to execute tasks of this type, so that the computing units can subsequently execute the tasks in the queues.
[0052] Since during the process of executing tasks, the task processing device 200 will, according to the type of (thread) task, schedule the task to the queue of the computing unit that matches this type, so that the computing unit can execute the task in this queue. Since usually, the multiple tasks to be executed are usually tasks of different types, therefore, multiple computing units on the processor core can usually be assigned tasks and execute these tasks, which enables multiple computing units on the processor core to execute multiple tasks of different types in parallel, thereby avoiding as much as possible that some computing units on the processor core are idle for a long time, and thus can make full use of the resources on the processor core and improve the resource utilization rate of the processor core.
[0053] Moreover, different workers are responsible for processing tasks of different computing types, which enables each worker to usually have sufficient computing unit resources to execute tasks, avoiding the problem of low execution efficiency of multiple tasks caused by multiple workers simultaneously using the same computing unit on the same processor core to process multiple tasks of the same type, achieving concurrent execution of multiple tasks at the hardware level, and thus being able to improve the overall efficiency of the processor core in executing multiple tasks.
[0054] It should be noted that the data processing system 10 shown above is only for illustrative purposes and is not used for limitation. For example, in the actual application scenario, the data processing system 10 also includes parts such as the kernel of the operating system ( Figure 1 not shown in the figure). For another example, in other data processing systems, there may be multiple servers, and each server can adopt the Figure 1 many-core architecture shown above and execute tasks based on the above method. At this time, the data processing system can be applied to the supercomputing scenario. For another example, in other data processing systems, the application layer 101 may include a larger number of applications, and the types and quantities of programming models included in different applications may vary; or, the hardware layer 102 may also include other types or other quantities of hardware. Or, in other data processing systems, the application layer 101 may not include a programming model, so that the applications in the application layer 101 can directly generate tasks and send them to the task processing device 200. Figure 1
[0055] For ease of understanding, the embodiments of the task processing method provided in the present application will be described below with reference to the accompanying drawings.
[0056] See Figure 4 Figure 4 which is a schematic flowchart of a task processing method provided in an embodiment of the present application. This method can be applied to the Figure 1 data processing system 10 described above, or can be applied to other applicable data processing systems. For ease of description, this embodiment takes the data processing system 10 shown in Figure 1 as an example for illustrative description.
[0057] Among them, Figure 4 the task processing method shown specifically may include:
[0058] S401: The task processing device 200 obtains the to-be-executed task 1 from among multiple to-be-executed tasks.
[0059] In this embodiment, after a programming model (such as programming model 1, etc.) edits a corresponding plurality of thread tasks and a computation graph corresponding to the plurality of thread tasks for each process task, the programming model can determine the thread tasks that can be executed currently according to the dependency relationships between different thread tasks indicated by the computation graph.
[0060] For example, the computation graph generated by the programming model can be as Figure 5 shown. Among them, Figure 5 it includes a plurality of nodes, and each node is used to indicate a thread task; the directed edges between different nodes are used to indicate the execution dependencies between different thread tasks. For example, the directed edge between node 1 and node 3 is used to indicate that the execution of the thread task identified by node 3 depends on the thread task identified by node 1 being executed first. Then, the task processing device 200 can determine at least one thread task that can be executed currently according to the execution dependencies between the plurality of thread tasks indicated by the computation graph. For example, the thread tasks indicated by the leaf nodes (such as nodes 4, 6, 8, or 9) in the computation graph can be determined as the thread tasks that can be executed currently.
[0061] Then, the programming model sends at least one thread task that can be executed currently to the kernel of the operating system through a corresponding interface, so as to request the kernel of the operating system to schedule resources to execute the at least one thread task. Correspondingly, the task processing device 200 can continuously monitor the interface (which can be one or more); and when the programming model outputs at least one executable thread task through this interface, the task processing device 200 can intercept the thread task sent by the programming model to the kernel, so that the task processing device 200 can perform resource scheduling for the thread task subsequently, such as scheduling a processor core to execute the thread task, etc.
[0062] Alternatively, after the programming model programs a plurality of thread tasks and generates a computation graph, the task processing device 200 can detect the thread tasks that can be executed currently according to the dependency relationships between different thread tasks indicated by the computation graph, that is, detect the thread tasks that do not depend on other thread tasks being executed first currently.
[0063] Assume that there are executable thread tasks currently, and the following will refer to this thread task as task 1.
[0064] Furthermore, the task processing device 200 can also combine the priorities of each task (used to indicate the priority degree of task execution) to determine the task 1 to be executed currently. For example, in Figure 5In the shown computation graph, if the tasks indicated by nodes 4, 6, 8, and 9 are the tasks that can be executed currently, then, the task processing device 200 can determine, from these 4 tasks, that the task indicated by node 8 is the task to be preferentially executed currently (the priority of the task indicated by node 8 is relatively high), that is, the task 1 described in step S401.
[0065] When specifically implemented, multiple (thread) tasks generated by the programming model can carry indication information of the priority of this task. For example, a technician / user can define that a task containing a data movement operator is executed with a relatively high priority, and a task containing a data access operator is executed with a relatively low priority, etc. Thus, when the programming model generates multiple tasks, it can add priorities to these tasks.
[0066] Alternatively, the priorities of multiple tasks generated by the programming model can be obtained by the task processing device 200 through analysis. For example, when the task processing device 200 obtains multiple (thread) tasks sent by the programming model, it can also obtain an identifier for indicating the computation type to which this task belongs. Thus, the task processing device 200 can determine the priority of task execution according to the task type indicated by this identifier. For example, when the type of the task is a scalar computation type, the task processing device 200 can determine that the priority of this task execution is relatively high, while when the type of the task is a matrix computation type, the task processing device 200 can determine that the priority of this task execution is relatively low. Then, the task processing device 200 can add indication information for indicating the high or low priority to each task. For example, the task processing device 200 can add a priority identifier to the task. When the value of this priority identifier is "high" or "1", it is used to indicate that the priority of this task execution is relatively high, while when the value of this priority identifier is "low" or "0", it is used to indicate that the priority of this task execution is relatively low.
[0067] Correspondingly, after the task processing device 200 obtains multiple tasks, it can determine, from the multiple tasks, the task that can be executed currently and has a relatively high priority according to the priority of each task execution. Assume that the determined task is the task 1 described in step S401.
[0068] S402: The task processing device 200 obtains the type of task 1.
[0069] After determining the task 1 to be executed currently, the task processing device 200 can obtain the type of the task 1. Among them, the thread tasks generated by the programming model can be tasks of multiple types. Exemplarily, the type of the task refers to the computing type to which the task belongs. For example, it can be a scalar computing type, a vector computing type, or a matrix computing type, etc., or it can be other types. Among them, a task of the scalar computing type refers to a task whose data computing process is mainly scalar computing. For example, it can be a data movement task, etc. A task of the vector computing type refers to a task whose data computing process is mainly vector computing. For example, it can be a data access task, etc. A task of the matrix computing type refers to a task whose data computing process is mainly matrix computing. For example, it can be a convolution computing task, etc.
[0070] Exemplarily, the task processing device 200 can determine the type of the task according to the operator included in the task 1. For example, when the task 1 includes a data access operator, it can be determined that the type of the task 1 is the scalar computing type; while when the task 1 includes a data movement operator, it can be determined that the type of the task 1 is the scalar computing type.
[0071] Alternatively, during the process of generating the task 1 by the programming model, a type identifier 1 can be added to the task 1, and this type identifier 1 is used to indicate the computing type to which the task 1 belongs. In this way, after obtaining the task 1, the task processing device 200 can determine the type of the task 1 according to the type identifier 1 carried by the task 1.
[0072] S403: When the type of the task 1 matches the type of the computing unit 1 on the processor core, the task processing device 200 schedules the task 1 to the queue 1 corresponding to the computing unit 1.
[0073] S404: When the type of the task 1 matches the type of the computing unit 2 on the processor core, the task processing device 200 schedules the task 1 to the queue 2 corresponding to the computing unit 2.
[0074] The processor core usually includes multiple different computing units, and different computing units are suitable for executing different types of tasks. Therefore, after obtaining the task 1 and the type of the task 1, the task processing device 200 can match the type of the task 1 with the types of the respective computing units on the processor core, and schedule the task 1 to the queue configured by the appropriate computing unit. Specifically, when the type of the task 1 matches the type of the computing unit 1, the task 1 is scheduled to the queue 1 of the computing unit 1, and when the type of the task 1 matches the type of the computing unit 2, the task 1 is scheduled to the queue 2 of the computing unit 2.
[0075] In this way, each computing unit can execute the tasks in the queue corresponding to the computing unit under the scheduling of the worker.
[0076] In this embodiment, the task processing device 200 may pre-create multiple workers for multiple computing units on the processor core 1, and each worker is responsible for scheduling one type of computing unit on the processor core 1 to execute tasks. In an actual application scenario, the number of each type of computing unit on the processor core 1 may be one or more.
[0077] Specifically, before processing the tasks issued by the programming model, the task processing device 200 may pre-administer the hardware in the hardware layer 102. For example, all the processor cores in the application layer 102 may be administered during the initialization of the task processing device 200 to obtain the hardware description information of each processor core included in the hardware layer 102. Then, the task processing device 200 may determine the multiple computing units included in each processor core according to the hardware description information. Thus, the task processing device 200 may create a worker for each type of computing unit on the processor core. For example, create worker 1 for computing unit 1 on the processor core 1, create worker 2 for computing unit 2 on the processor core 1, etc., and establish a mapping relationship between the worker and the computing unit. Next, the task processing device 200 may add a type identifier to each worker according to the type of the computing unit corresponding to each worker to indicate the type of tasks that the worker can handle when scheduling the computing unit.
[0078] Then, the task processing device 200 may determine which computing unit's type matches the type of task 1. Taking the type of task 1 matching the type of the first computing unit as an example, the task processing device 200 may determine worker 1 for processing task 1 from the multiple pre-created workers, and according to the mapping relationship between worker 1 and computing unit 1 on the processor core 1, use worker 1 to schedule computing unit 1 on the processor core 1 to execute task 1.
[0079] For the convenience of understanding and explanation, the following uses the task processing device 200 using worker 1 to process task 1 to introduce the task processing process.
[0080] In a first possible implementation manner, the task processing device 200 may be, for example, a scheduler running in the processor, and the scheduler may be implemented by software, for example, a program running in the processor. At this time, the scheduler may determine worker 1 from the multiple workers created for the processor core 1 according to the type identifier 1 of task 1, and the type identifier of worker 1 matches the type identifier 1 of task 1.
[0081] The scheduler can be configured with a target queue, which can be, for example, a queue that supports the first in first out (FIFO) principle. Thus, the scheduler can add multiple thread tasks to be executed and the type identifier of each thread task to the target queue according to the execution dependencies between multiple thread tasks (and the priorities of the tasks to be executed). Then, the scheduler schedules the multiple thread tasks in the target queue to different computing units for execution in sequence. Correspondingly, the scheduler can obtain task 1 and the type identifier 1 of task 1 from the queue.
[0082] Then, the scheduler can schedule each thread task in the target queue one by one according to the order of the thread tasks in the target queue. Taking the currently scheduled thread task as task 1 as an example, the scheduler can match the type identifier 1 of task 1 with the type identifiers of each worker and determine the worker with a successful match as worker 1. In practical applications, the successfully matched worker 1 is currently in an idle state, that is, the worker 1 is not currently using the computing unit to execute tasks.
[0083] It can be understood that since the same functional computing units can be included on different processor cores, such as computing units for scalar calculation, there can be workers responsible for processing tasks of the calculation type 1 indicated by the type identifier 1 among the multiple workers created by the scheduler for each processor core in advance. Therefore, after obtaining the type identifier 1, the scheduler can determine whether the worker corresponding to processor core 1 for executing tasks of this calculation type 1 is in an idle state. If so, the scheduler can determine this worker as worker 1. If not, the scheduler can continue to determine whether the worker corresponding to processor core 2 for executing tasks of this calculation type 1 is in an idle state. If so, the scheduler can determine the worker corresponding to processor core 2 as worker 1. If not, the scheduler continues to determine whether the workers corresponding to the remaining processor cores for executing tasks of calculation type 1 are in an idle state. And so on, until the scheduler determines worker 1. In this embodiment, worker 1 is set as an example of the worker created for the computing unit 1 on processor core 1 for illustration.
[0084] When the calculation type indicated by the type identifier 1 matches the actual calculation type to which task 1 belongs, the scheduler schedules task 1 to queue 1 corresponding to the computing unit 1 responsible for scheduling by worker 1, so that worker 1 schedules the computing unit 1 to execute task 1 in queue 1, as Figure 4 shown.
[0085] In an actual application scenario, the computing type indicated by the type identifier 1 added by the programming model for task 1 may be the same as the computing type to which task 1 actually belongs, or may be inconsistent with the computing type to which task 1 belongs (such as the programming model marking the wrong type identifier for task 1, etc.). Therefore, before using worker 1 to execute task 1, the scheduler can first determine whether the computing type indicated by the type identifier 1 matches the computing type to which task 1 actually belongs. For example, it can be determined whether the computing type to which task 1 actually belongs is the computing type indicated by the type identifier 1 according to the operators included in task 1, etc. If they match, it indicates that the type identifier marked by the programming model for task 1 is correct, then the scheduler can store task 1 in queue 1 configured for computing unit 1; correspondingly, worker 1 can retrieve task 1 from this queue 1 and call computing unit 1 on processor core 1 to execute this task 1. If they do not match, it indicates that the type identifier marked by the programming model for task 1 is wrong, then the scheduler can determine the computing type to which task 1 actually belongs according to the operators included in task 1, and store this task 1 in the queue corresponding to the computing unit on processor core 1 or other processor cores that can be used to execute this task, so that task 1 can be executed by other computing units, thereby ensuring the execution efficiency of task 1 by scheduling task 1 to a suitable computing unit.
[0086] In actual application, in addition to the scheduler being able to verify whether the type identifier 1 marked by the programming model for task 1 is correct, it can also be the worker (or rather, the computing unit) that verifies whether the type identifier 1 of task 1 is correct. In specific implementation, the scheduler can first add task 1 to queue 1 configured for computing unit 1 according to the type identifier 1. Then, during the running process, worker 1 can retrieve task 1 from queue 1 and determine whether the computing type indicated by the type identifier 1 matches the computing type to which task 1 actually belongs. If they match, then worker 1 can schedule computing unit 1 on processor core 1 to execute this task 1. If they do not match, then worker 1 can schedule this task 1 from queue 1 to queue 2 allocated for computing unit 2. Among them, the computing type to which the tasks that computing unit 2 can handle belongs is the same as the computing type to which this task 1 actually belongs; and, computing unit 2 is currently in a state of not executing tasks. Thus, after worker 2 determines that the computing type indicated by the type identifier 1 matches the computing type to which this task 1 actually belongs, it schedules this computing unit 2 to execute this task 1, as Figure 4 shown. Among them, worker 2 is the worker created by task processing device 200 for computing unit 2 on processor core 1.
[0087] Further, if the computing unit 2 on the processor core 1 is currently in a state of executing a task, the worker 1 can schedule task 1 to the queue corresponding to the computing unit on other processor cores, and the types of tasks that the computing unit on the other processor cores can execute are the same as the type of task 1. For example, the worker 1 can determine whether the computing unit 3 on the processor core 2 that can be used to execute task 1 is in a state of executing a task. If not, the worker 1 can schedule task 1 to the computing unit 3 on the processor core 2 so that the computing unit 3 on the processor core 2 can execute task 1. If so, the worker 1 can continue to judge the computing units on other processor cores that are used to execute this type of task in the above-mentioned manner.
[0088] In practical applications, each worker may only be able to call one type of computing unit on the processor core. Alternatively, each worker can call multiple types of computing units on the processor core. For example, when task 1 simultaneously includes multiple scalar calculation type operators a and a matrix calculation type operator b, when the worker 1 is calling the computing unit 1 for scalar calculation to execute task 1, when it comes to executing operator b, the worker 1 can call the computing unit 4 for matrix calculation on the processor core 1 to execute operator b in task 1. Correspondingly, for the worker 4 created by the task processing device 200 for the computing unit 4, in addition to being able to call the computing unit 4 to process matrix calculation type tasks, it can also call the computing unit 1 to process the scalar calculation type operators included in this task.
[0089] In the second possible implementation manner, it can also be the worker (or the computing unit running this worker) that schedules the task 1 to be executed. At this time, during the process of scheduling task 1, the task processing device can specifically be the worker 1 (or the computing unit 1 running this worker). The task processing device 200 can add the multiple tasks to be executed and the type identifier of each task to the target queue according to the execution dependencies (and priorities) between the multiple tasks to be executed. In this way, after the worker 1 schedules the computing unit 1 to execute the task, it can take out a task from the target queue. Taking the currently taken out task as task 1 as an example. Then, the worker 1 can determine whether the type identifier 1 of the task matches the type identifier of the worker 1. If it matches, the worker 1 can store task 1 in the queue 1 corresponding to the computing unit 1; if it does not match, the worker 1 can store task 1 in the queue corresponding to other computing units according to this type identifier 1.
[0090] In an actual application scenario, after Task 1 is scheduled to Queue 1, it may be that Computing Unit 1 cannot execute Task 1 in this Queue 1. For example, if Task 1 is actually a matrix calculation type task, while Computing Unit 1 is a computing unit for scalar calculation, then Computing Unit 1 cannot execute this Task 1. Therefore, in a further possible implementation, after Task 1 is scheduled to Queue 1, when Computing Unit 1 cannot execute Task 1 in this Queue 1, Computing Unit 1 can determine the actual type to which Task 1 belongs, and determine the computing unit that can be used to execute Task 1 according to the actual type to which Task 1 belongs. Assuming that when Computing Unit 1 determines that the actual type to which Task 1 belongs matches the type of Computing Unit 2, Computing Unit 1 can schedule Task 1 to Queue 2 corresponding to Computing Unit 2, so that Computing Unit 2 executes Task 1. In this way, when Computing Unit 1 cannot execute Task 1, Task 1 can be promptly scheduled to other computing units for execution, thereby improving the execution efficiency of Task 1.
[0091] As an implementation example, before executing Task 1, Worker 1 can also first determine whether the calculation type indicated by Type Identifier 1 matches the actual calculation type to which Task 1 belongs. If they match, Worker 1 can call Computing Unit 1 to execute Task 1 in Queue 1. If they do not match, Worker 1 can determine the actual calculation type to which Task 1 belongs according to the operators included in Task 1, and schedule Task 1 to Queue 2 corresponding to Computing Unit 2, which is a computing unit on Processor Core 1 or other processor cores that can be used to execute Task 1, so that Computing Unit 2 executes Task 1, thereby ensuring the execution efficiency of Task 1 by scheduling Task 1 to a suitable computing unit.
[0092] Similarly, for other tasks among multiple tasks, such as Task 2, Task Processing Device 200 can continue to obtain Task 2 and the type of Task 2 from the multiple tasks to be executed by referring to the above method. Then, according to the type of Task 2, Task 2 is scheduled to the queue corresponding to the corresponding computing unit with a matching type. For example, when Task Processing Device 200 is a scheduler, the scheduler can take out a new task from the target queue, that is, Task 2, and schedule Task 2 to the queue corresponding to the corresponding computing unit according to the type of Task 2. Among them, the process of Task Processing Device 200 scheduling and using the computing unit to execute Task 2 is similar to the above process of scheduling and executing Task 1, and will not be elaborated here.
[0093] Thus, during the process of task execution, the task processing device 200 will utilize workers with matching computing types to schedule the computing units corresponding to the computing type on the processor core 1 or other processor cores to execute the task. In this way, for multiple thread tasks of different types (such as the above-mentioned task 1 and task 2), the task processing device can utilize different workers to call different types of computing units on the processor core 1 to execute different types of thread tasks in parallel, thereby avoiding as much as possible that some computing units on the processor core 1 are idle for a long time, and thus improving the resource utilization rate of the processor core 1. As Figure 6 shown, the task processing device 200 can simultaneously utilize the computing unit 1 and the computing unit 2 on the processor core 1 to concurrently execute task 1 and task 2, avoiding the situation where the resources of the computing unit 2 are idle during the execution of task 1 and causing task 2 to be in a waiting state for a long time.
[0094] In addition, in this embodiment, different workers can be responsible for processing tasks of different computing types, which enables each worker to usually have sufficient computing unit resources to execute tasks, avoiding the problem that multiple workers use the same computing unit on the processor core 1 to process multiple tasks simultaneously, resulting in low execution efficiency of these multiple tasks. In this way, the overall efficiency of the processor core 1 for parallel execution of multiple tasks can be improved, thereby improving the business operation performance of application 1.
[0095] When the task processing device 200 identifies an executable task from multiple (thread) tasks, after determining that worker 1 has completed the execution of task 1, the task processing device 200 can update the execution status of task 1 in the computation graph (such as marking it as "executed", etc.), and continue to obtain one or more new executable tasks from the computation graph, and add the new task to the target queue. At the same time, the task processing device 200 can continue to take out tasks from the target queue, and allocate the task to the corresponding worker for execution according to the type identifier of the taken-out task.
[0096] It should be noted that the above Figure 4 illustrated method embodiment is only for illustrative purposes. Based on the Figure 4 shown method flow, the process of the task processing device 200 using multiple workers to execute tasks can also adopt the following various embodiments.
[0097] Example 1, Figure 4In the illustrated embodiment, the task processing device 200 assigns task 1 to the corresponding worker according to the type identifier added to task 1 based on the programming model. In other embodiments, after obtaining a task, the task processing device 200 can analyze the actual computing type to which the task belongs according to the operators included in the task, and assign the task to the worker responsible for processing tasks of this computing type based on the obtained computing type. In this way, after each worker is assigned a task, it can directly call the corresponding computing unit on the processor core to execute the task, without pushing the task to other workers according to the computing type to which the task belongs.
[0098] Example 2, Figure 4 In the illustrated embodiment, an example is given in which the task processing device 200 schedules each task to the queue corresponding to the corresponding computing unit according to the computational graph. In other embodiments, after the task processing device 200 generates the computational graph, each idle worker can obtain the currently executable task according to the computational graph, and when it is determined that the computing type to which the task belongs matches the computing type to which the tasks that the worker itself can process belong, the worker can obtain the task and schedule the corresponding computing unit on the processor core to execute the task. Further, after the worker finishes executing the current task, it can update the execution status of the task in the computational graph (such as marking it as "executed", etc.), and continue to obtain new unexecuted tasks from the computational graph.
[0099] The above Figure 4 In the illustrated embodiment, it is described by taking one worker to process task 1 as an example. In other embodiments, the task processing device 200 can also use multiple workers to collaboratively process the same task to improve the efficiency of processing this task. The following combines Figure 7 to describe this in detail.
[0100] See Figure 7 , which shows a schematic flowchart of another task processing method. As Figure 7 shown, this method may specifically include:
[0101] S701: The task processing device 200 obtains the task 2 to be executed.
[0102] Among them, task 2 includes multiple operators. In this embodiment, it is described by taking task 2 including operator 1 and operator 2 as an example. And, operator 1 and operator 2 are operators of the same computing type, specifically, they can be operators of the vector computing type.
[0103] Among them, the specific implementation manner for the task processing device 200 to obtain task 2 is the same as the above Figure 4In the illustrated embodiment, the implementation manner in which the task processing device 200 obtains Task 1 is similar. For specific details, reference may be made to the relevant descriptions above, and thus will not be elaborated here.
[0104] S702: The task processing device 200 divides Task 2 into Sub-task 1 and Sub-task 2. Among them, Sub-task 1 includes Operator 1 of the vector calculation type, and Sub-task 2 includes Operator 2 of the vector calculation type.
[0105] In this embodiment, after obtaining Task 2, the task processing device 200 can split this Task 2 into multiple sub-tasks, and each sub-task includes one operator.
[0106] As an implementation example, the task processing device 200 can block the execution logic of Task 2, and the execution of each block can be used as a sub-task of Task 2, thereby realizing splitting Task 2 into multiple sub-tasks. Among them, each block of execution logic can be an operator in Task 1. Thus, the task processing device 200 can split this Task 2 into multiple sub-tasks with the operator as the granularity.
[0107] Among them, based on Sub-task 1 and Sub-task 2 obtained by dividing Task 2, they can be executed in parallel, that is, the execution of Sub-task 1 and the execution of Sub-task 2 can be independent of each other.
[0108] S703: The task processing device 200 determines the types of Sub-task 1 and Sub-task 2.
[0109] Exemplarily, the task processing device 200 can determine the type of each sub-task according to the operator included in each sub-task, that is, determine the calculation type to which the sub-task belongs. That is, both Sub-task 1 and Sub-task 2 belong to the vector calculation type.
[0110] S704: The task processing device 200 stores Sub-task 1 into Queue 3 corresponding to the computing unit 3 scheduled by Worker 3 according to the type of Sub-task 1, and stores Sub-task 2 into Queue 4 corresponding to the computing unit 4 scheduled by Worker 4 according to the type of Sub-task 2.
[0111] Among them, both Worker 3 and Worker 4 are workers for executing tasks of the same vector calculation type.
[0112] In specific implementation, the task processing device 200 can respectively create corresponding workers for various computing units on the processor core 1 and for various computing units on the processor core 2 in advance, and add corresponding identifiers to each worker, where the identifier is used to indicate the computing type to which the task that the worker can call the computing unit belongs. For example, assume that the processor core 1 includes a computing unit 1 for scalar calculation, a computing unit 2 for matrix calculation, and a computing unit 3 for vector calculation. Then, the task processing device 200 can create a worker 1 for the computing unit 1, a worker 2 for the computing unit 2, and a worker 3 for the computing unit 3, and add identifiers of the computing type to each worker respectively. At the same time, the processor core 2 also includes a computing unit 4 for vector calculation, a computing unit 5 for matrix calculation, and a computing unit 6 for scalar calculation. Then, the task processing device 200 can create a worker 4 for the computing unit 4, a worker 5 for the computing unit 5, and a worker 6 for the computing unit 6, and add identifiers of the computing type to each worker respectively.
[0113] After determining the types of the subtask 1 and the subtask 2, the task processing device 200 can match the types of each subtask with the types of multiple pre-created workers. Specifically, it can be the type identifier of the subtask that is matched with the type identifier of the worker. Assume that the type of the subtask 1 matches the type of the worker 3, and the type of the subtask 2 matches the type of the worker 4. Then, the task processing device 200 can schedule the subtask 1 to the queue corresponding to the computing unit 3, and schedule the subtask 2 to the queue corresponding to the computing unit 4.
[0114] S705: The task processing device 200 uses the worker 3 to schedule the computing unit 3 on the processor core 1 to execute the subtask 1, and uses the worker 4 to schedule the computing unit 4 on the processor core 2 to execute the subtask 2.
[0115] In this way, the task processing device 200 can improve the overall execution efficiency of the task 2 by splitting the task 2 into multiple subtasks and using multiple workers to schedule the computing units on different processor cores to execute the multiple subtasks in parallel. And when the computing unit 4 on the processor core 2 is idle, it can cooperate with the computing unit 3 on the processor core 1 to execute the same task, thereby improving the resource utilization rate of the processor core 2.
[0116] It should be noted that Figure 7 The shown task processing method is only used as an exemplary illustration and is not used for limitation. For example, in other embodiments, when the thread task issued by the programming model includes multiple operators of different computing types, the task processing device 200 can also refer to the above Figure 7The method flow shown uses different types of computing units on one or more processor cores to execute different subtasks included in the task in parallel, thereby improving the overall efficiency of task execution.
[0117] It should be noted that other reasonable combinations of steps that can be thought of by those skilled in the art based on the above description also fall within the protection scope of this application. Secondly, those skilled in the art should also be familiar that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily essential for this application.
[0118] The above combination Figures 1 to 7 introduces the task processing method provided by the embodiments of this application. Next, the structures of the task processing device and the computing device provided by the embodiments of this application will be introduced with reference to the accompanying drawings.
[0119] See Figure 8 , which shows a schematic structural diagram of a task processing device. Figure 8 The task processing device 800 shown is applied to a processor. The processor includes at least one processor core. The first processor core included in the at least one processor core includes a first computing unit and a second computing unit, and the first computing unit and the second computing unit are different types of computing units.
[0120] As Figure 8 shown, the task processing device 800 includes:
[0121] An acquisition module 801, configured to acquire a first task from multiple tasks to be executed; acquire the type of the first task.
[0122] A scheduling module 802, configured to schedule the first task to the first queue corresponding to the first computing unit when the type of the first task matches the type of the first computing unit; schedule the first task to the second queue corresponding to the second computing unit when the type of the first task matches the type of the second computing unit.
[0123] In a possible implementation manner, the multiple tasks to be executed are tasks respectively executed by multiple threads included in a process of an application.
[0124] In a possible implementation manner, the types of the first computing unit and the second computing unit include a matrix computing unit, a scalar computing unit, or a vector computing unit.
[0125] In a possible implementation manner, the acquisition module 801 is further configured to acquire a second task from the multiple tasks to be executed, and the second task includes a first operator and a second operator.
[0126] The task processing device 800 further includes:
[0127] A partitioning module 803, configured to partition a second task into multiple subtasks, the multiple subtasks including a first subtask and a second subtask, the first subtask including a first operator, and the second subtask including a second operator;
[0128] The scheduling module 802 is further configured to:
[0129] Schedule the first subtask to a first queue, where the type of the first subtask matches the type of the first computing unit;
[0130] Schedule the second subtask to a second queue, where the type of the second subtask matches the type of the second computing unit.
[0131] In a possible implementation manner, the task processing device 800 is specifically a first computing unit, and the first task is scheduled to a first queue corresponding to the first computing unit;
[0132] The task processing device 800 further includes:
[0133] A determination module 804, configured to determine the actual type to which the first task belongs when the first computing unit cannot execute the first task in the first queue;
[0134] The scheduling module 802 is further configured to schedule the first task to the second queue when the actual type to which the first task belongs matches the type of the second computing unit.
[0135] In a possible implementation manner, the processor further includes a second processor core; the first processor core and the second processor core are processor cores of the same type, or the first processor core and the second processor core are processor cores of different types.
[0136] Since Figure 8 the task processing device 800 shown corresponds to the task processing device 200 in the above Figure 4 or Figure 7 shown embodiment, therefore Figure 8 For the specific implementation manner of the task processing device 800 shown and the technical effects thereof, refer to the relevant descriptions in the above Figure 4 or Figure 7 shown embodiment, which will not be elaborated herein.
[0137] Figure 9 FIG. is a schematic hardware structure diagram of a computing device 900 provided by the present application. The computing device 900 can, for example, implement the task processing device 200 in the above Figure 4 or Figure 7 shown embodiment, etc.
[0138] As Figure 9As shown, the computing device 900 includes a processor 901, a memory 902, and a communication interface 903. Among them, the processor 901, the memory 902, and the communication interface 903 communicate through a bus 904, and can also achieve communication through other means such as wireless transmission. The memory 902 is used to store instructions, and the processor 901 is used to execute the instructions stored in the memory 902. Further, the computing device 900 may further include a memory unit 905, and the memory unit 905 can be connected to the processor 901, the storage medium 902, and the communication interface 903 through the bus 904. Among them, the memory 902 stores program codes, and the processor 901 can call the program codes stored in the memory 902 to perform the following operations:
[0139] Obtain a first task from multiple tasks to be executed;
[0140] Obtain the type of the first task;
[0141] When the type of the first task matches the type of the first computing unit included in the first processor core, schedule the first task to the first queue corresponding to the first computing unit;
[0142] When the type of the first task matches the type of the second computing unit included in the first processor core, schedule the first task to the second queue corresponding to the second computing unit, where the first computing unit and the second computing unit are computing units of different types.
[0143] It should be understood that in this embodiment, the processor 901 may be a CPU, and the processor 901 may also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete device components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc.
[0144] The memory 902 may include a read-only memory and a random access memory, and provide instructions and data to the processor 901. The memory 902 may further include a non-volatile random access memory.
[0145] The memory 902 can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchlink DRAM (SLDRAM), and direct rambus RAM (DR RAM).
[0146] The communication interface 903 is used to communicate with other devices connected to the computing device 900. In addition to including a data bus, the bus 904 can also include a power bus, a control bus, a status signal bus, etc. However, for the sake of clarity, all kinds of buses are labeled as the bus 904 in the figure.
[0147] It should be understood that the computing device 900 according to the embodiment of the present application can correspond to the task processing device 800 in the embodiment of the present application, and can correspond to the method executed by the task processing device 200 in the method shown in Figure 4 Or Figure 7 The above and other operations and / or functions implemented by the computing device 900 are respectively for implementing the Figure 4 Or Figure 7 The corresponding method flow of the corresponding method in, for the sake of brevity, will not be elaborated here.
[0148] Embodiments of the present application also provide a computer-readable storage medium. The computer-readable storage medium may be any available medium that can be stored by a computing device or a data storage device such as a data center that contains one or more available media. The available medium may be a magnetic medium (e.g., a floppy disk, a hard disk, a magnetic tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive), etc. The computer-readable storage medium includes instructions that direct the computing device to execute the above-mentioned task processing method.
[0149] Embodiments of the present application also provide a computer program product. The computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computing device, they generate, in whole or in part, the processes or functions described in the embodiments of the present application.
[0150] The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, or data center to another website, computer, or data center by wire (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wirelessly (e.g., infrared, wireless, microwave, etc.).
[0151] The computer program product may be a software installation package. In the case where any of the above-mentioned task processing methods is required, the computer program product can be downloaded and executed on a computing device.
[0152] The above embodiments can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded or executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from a website, computer, server, or data center to another website, computer, server, or data center via wired (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that contains one or more collections of available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state drive.
[0153] The terms used in the above embodiments are only for the purpose of describing specific embodiments and are not intended to limit the present application. As used in the specification and claims of the present application, the singular forms "a", "an", "the", "above", "said", "this", and "such" are also intended to include the form of "one or more", unless the context clearly indicates otherwise. It should also be understood that in the embodiments of the present application, "one or more" means one, two, or more than two; the character " / " generally indicates an "or" relationship between the associated objects before and after. In the embodiments of the present application, "simultaneously" means within the same time period, including the case of being at the same moment.
[0154] The reference to "one embodiment" or "some embodiments" etc. described in this specification means that a specific feature, structure, or characteristic described in connection with the embodiment is included in one or more embodiments of the present application. Thus, the statements "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc. that appear in different places in this specification are not necessarily all referring to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways. The terms "comprising", "including", "having", and their variants all mean "including but not limited to", unless otherwise specifically emphasized in other ways.
[0155] As described above, it is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.
Claims
1. A task processing method, applied to a processor, characterized in that The first processor core in the processor includes a first computing unit and a second computing unit, and the first computing unit and the second computing unit are different types of computing units; The method includes: Obtain a first task from multiple tasks to be executed; Obtain the type of the first task; When the type of the first task matches the type of the first computing unit, schedule the first task to a first queue corresponding to the first computing unit; When the type of the first task matches the type of the second computing unit, schedule the first task to a second queue corresponding to the second computing unit.
2. The method according to claim 1, characterized in that, The method is executed by a scheduler running in the processor, and the scheduler is configured with a target queue that stores the multiple tasks to be executed; The obtaining a first task from multiple tasks to be executed includes: The scheduler obtains the first task from the target queue; The scheduling the first task to the first queue corresponding to the first computing unit includes: The scheduler stores the first task into the first queue; The scheduling the first task to the second queue corresponding to the second computing unit includes: The scheduler stores the first task into the second queue.
3. The method according to claim 1, characterized in that, The method is executed by the first computing unit; The obtaining a first task from multiple tasks to be executed includes: After the first computing unit finishes executing the tasks in the first queue, obtain the first task from the multiple tasks to be executed; The scheduling the first task to the first queue corresponding to the first computing unit includes: The first computing unit stores the first task into the first queue; The scheduling the first task to the second queue corresponding to the second computing unit includes: The first computing unit stores the first task into the second queue.
4. The method according to any one of claims 1 to 3, characterized in that The multiple tasks to be executed are tasks respectively executed by multiple threads included in a process of an application.
5. The method according to any one of claims 1 to 4, characterized in that, The types of the first computing unit and the second computing unit include a matrix computing unit, a scalar computing unit, or a vector computing unit.
6. The method according to any one of claims 1 to 5, characterized in that, The method further includes: Obtain a second task from multiple tasks to be executed, where the second task includes a first operator and a second operator; Divide the second task into multiple subtasks, the multiple subtasks include a first subtask and a second subtask, the first subtask includes the first operator, and the second subtask includes the second operator; Schedule the first subtask to the first queue, and the type of the first subtask matches the type of the first computing unit; Schedule the second subtask to the second queue, and the type of the second subtask matches the type of the second computing unit.
7. The method according to any one of claims 1 to 6, characterized in that, The method is executed by the first computing unit, and the first task is scheduled to the first queue; The method further includes: When the first computing unit cannot execute the first task in the first queue, the first computing unit determines the actual type to which the first task belongs; When the type to which the first task actually belongs matches the type of the second computing unit, the first computing unit schedules the first task to the second queue.
8. A task processing device, applied to a processor, characterized in that, The first processor core in the processor includes a first computing unit and a second computing unit, and the first computing unit and the second computing unit are computing units of different types; The device includes: An obtaining module, configured to obtain a first task from multiple tasks to be executed; and obtain the type of the first task; A scheduling module, configured to schedule the first task to the first queue corresponding to the first computing unit when the type of the first task matches the type of the first computing unit; and schedule the first task to the second queue corresponding to the second computing unit when the type of the first task matches the type of the second computing unit.
9. The device according to claim 8, characterized in that The multiple tasks to be executed are tasks respectively executed by multiple threads included in a process of an application.
10. The device according to claim 8 or 9, characterized in that, The types of the first computing unit and the second computing unit include a matrix computing unit, a scalar computing unit, or a vector computing unit.
11. The device according to any one of claims 8 to 10, wherein The obtaining module is further configured to obtain a second task from multiple tasks to be executed, and the second task includes a first operator and a second operator; The device further includes: A dividing module, configured to divide the second task into multiple subtasks, the multiple subtasks including a first subtask and a second subtask, the first subtask including the first operator, and the second subtask including the second operator; The scheduling module is further configured to schedule the first subtask to the first queue, where the type of the first subtask matches the type of the first computing unit; and schedule the second subtask to the second queue, where the type of the second subtask matches the type of the second computing unit.
12. The device according to any one of claims 8 to 11, characterized in that, The device is the first computing unit, and the first task is scheduled to the first queue; The device further includes: A determining module, configured to determine the type to which the first task actually belongs when the first computing unit cannot execute the first task in the first queue; The scheduling module is further configured to schedule the first task to the second queue when the type to which the first task actually belongs matches the type of the second computing unit.
13. A computing device, characterized in that, Including a processor and a memory; The processor is configured to execute instructions stored in the memory, so that the computing device executes the method according to any one of claims 1 to 7.
14. A computer-readable storage medium, characterized in that, Including instructions, when running on a computing device, causing the computing device to execute the method according to any one of claims 1 to 7.
15. A computer program product comprising instructions, characterized in that, When running on at least one computing device, causing the at least one computing device to execute the method according to any one of claims 1 to 7.