Task processing method and apparatus, and related device

By setting up different types of computing units on the processor core and scheduling according to the task type, the problem of idle computing units is solved, and efficient utilization of processor core resources and improvement of task execution efficiency is achieved.

WO2025138965A1PCT designated stage expired Publication Date: 2025-07-03HUAWEI TECH CO LTD

Patent Information

Application Number
PCT/CN2024/115320
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-28
Filing Date
2024-08-29
Publication Date
2025-07-03

AI Technical Summary

Technical Problem

In the prior art, due to mismatch of task scheduling in the processor core, some computing units are idle for a long time, and the resource utilization rate is low, which affects the overall task execution efficiency of the processor core.

Method used

By setting up different types of computing units on the processor core and scheduling them to the corresponding queue according to the task type, we ensure that each computing unit performs the appropriate task type, and realizes reasonable allocation of tasks and parallel processing.

Benefits of technology

It improves the resource utilization rate of the processor core, improves the overall execution efficiency of multiple tasks, avoids the idleness of computing units, and improves the task processing capability of the processor core.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024115320_03072025_PF_FP_ABST
    Figure CN2024115320_03072025_PF_FP_ABST
Patent Text Reader

Abstract

A task processing method and apparatus and a related device, relating to the technical field of task processing. The method is applied to a processor, the processor comprises a first processor core, and the first processor core comprises a first computing unit and a second computing unit of different types. During task processing, a first task is acquired from among a plurality of tasks to be executed, and the type of the first task is acquired, so that when the type of the first task matches the type of the first computing unit, the first task is scheduled to a first queue corresponding to the first computing unit, and when the type of the first task matches the type of the second computing unit, the first task is scheduled to a second queue corresponding to the second computing unit. In this way, different types of computing units can be assigned appropriate types of tasks, so that the plurality of computing units on the processor core can execute different types of tasks in parallel, thereby improving the resource utilization rate of the processor core and improving the overall efficiency of the processor core in executing a plurality of tasks.
Need to check novelty before this filing date? Find Prior Art

Description

Task processing method, device and related equipment

[0001] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office on December 28, 2023, with application number 202311844559.8 and application name “Task processing methods, devices and related equipment”, the entire contents of which are incorporated by reference into this application. Technical Field

[0002] The present application relates to the field of task processing technology, and in particular to a task processing method, apparatus, and related equipment. Background Art

[0003] Currently, parallel processing can effectively improve the processing efficiency of multiple tasks and has become a mainstream task processing method. Specifically, when running an application, multiple threads can be used to call resources in the processor core to execute application-related tasks in parallel. The resources in the processor core can include various computing units, such as computing units for matrix calculations, vector calculations, and scalar calculations.

[0004] In real-world applications, tasks are executed using the compute units on a processor core. These tasks are queued in a queue and await scheduling. Typically, tasks in the queue are dispatched to the processor core in the order they appear in the queue, allowing the corresponding compute units on that core to execute the tasks. This often results in some compute units on a processor core not being assigned tasks, leaving them idle for extended periods of time and underutilizing the resources on that core.

[0005] Summary of the Invention

[0006] The present application provides a task processing method to improve the overall efficiency of a processor core in executing multiple tasks. In addition, the present application also provides a task processing apparatus, a processor, a computing device, a computer-readable storage medium, and a computer program product.

[0007] In the first aspect, the present application provides a task processing method, which is applied to a processor, wherein the processor includes at least one processor core, the at least one processor core includes a first processor core, and the first processor core includes a first computing unit and a second computing unit, and the first computing unit and the second computing unit are computing units of different types. The method can be executed by a task processing device. In the process of processing tasks, the task processing device obtains a first task from multiple tasks to be executed, such as determining the first task that can be currently executed according to the execution dependency between the multiple tasks, and further obtaining the type of the first task, which type can be, for example, a matrix calculation type, a vector calculation type, or a scalar calculation type, so that when the type of the first task matches the type of the first computing unit, the first task is scheduled to the first queue corresponding to the first computing unit, and when the type of the first task matches the type of the second computing unit, the first task is scheduled to the second queue corresponding to the second computing unit.

[0008] Since, during the execution of tasks, the task processing device can schedule appropriate types of tasks for different types of computing units, multiple computing units on the processor core can execute multiple tasks of different types in parallel under the scheduling of the task processing device. This can avoid as much as possible that some computing units on the processor core are idle for a long time, thereby improving the resource utilization of the processor core; at the same time, it can realize the concurrent execution of multiple tasks at the hardware level, thereby improving the overall efficiency of the processor core in executing multiple tasks.

[0009] In one possible implementation, the task processing device may be a scheduler running in a processor, that is, the task processing method may be executed by the scheduler, the scheduler being configured with a target queue, and the target queue storing a plurality of tasks to be executed, so that when obtaining a first task, the scheduler specifically obtains the first task from the target queue, and when scheduling the first task, the scheduler may specifically store the first task in a first queue corresponding to the first computing unit when the type of the first task matches the type of the first computing unit, or store the first task in a second queue corresponding to the second computing unit when the type of the first task matches the type of the second computing unit. In this way, the scheduler may implement global scheduling of tasks, thereby improving the reliability of task scheduling.

[0010] In one possible implementation, the task processing device may specifically be the first computing unit in the first processor core, that is, the task processing method may be executed by the first computing unit, so that when the first computing unit obtains the first task, it may specifically obtain the first task from multiple tasks to be executed after executing the tasks in the first queue, and when scheduling the first task, the first computing unit may specifically store the first task in the first queue corresponding to the first computing unit when the type of the first task matches the type of the first computing unit, or store the first task in the second queue corresponding to the second computing unit when the type of the first task matches the type of the second computing unit. In this way, the computing unit can implement task scheduling, thereby improving the flexibility of task scheduling.

[0011] In a possible implementation, the multiple tasks to be executed are tasks that are respectively executed by multiple threads included in a process of an application. In this way, multiple threads can be used to implement parallel execution of the multiple tasks, thereby effectively improving the processing efficiency of the multiple tasks.

[0012] In a possible implementation, the first computing unit and the second computing unit may be a matrix computing unit, a scalar computing unit, or a vector computing unit. That is, each computing unit may be any one of a matrix computing unit, a scalar computing unit, or a vector computing unit.

[0013] In one possible implementation, in the process of processing multiple tasks, the task processing device can also obtain a second task from the multiple tasks to be executed, and the second task includes multiple operators, taking the first operator and the second operator as an example. The task processing device can divide the second task into multiple subtasks, and the multiple subtasks include a first subtask and a second subtask, wherein the first subtask includes a first operator and the second subtask includes a second operator, so that when scheduling tasks, the task processing device schedules the first subtask to the first queue corresponding to the first computing unit, at which time the type of the first subtask matches the type of the first computing unit; at the same time, the task processing device schedules the second subtask to the second queue corresponding to the second computing unit, at which time the type of the second subtask matches the type of the second computing unit. In this way, the task processing device can improve the overall execution efficiency of the second task by dividing the task into multiple subtasks and executing the multiple subtasks in parallel through multiple computing units.

[0014] In a possible embodiment, the task processing device can specifically be the first computing unit in the first processor core, and the first task has been scheduled to the first queue corresponding to the first computing unit. In the process of processing tasks, when the first computing unit is unable to execute the first task of the first queue, such as the first computing unit is a computing unit for matrix calculation, and the first task is actually a vector calculation task, etc., at this time, the first computing unit can determine the actual type of the first task, and when the actual type of the first task matches the type of the second computing unit, the first computing unit schedules the first task to the second queue corresponding to the second computing unit so that the second computing unit can execute the first task. In this way, when the first computing unit is unable to execute the first task (that is, an error occurs in task scheduling), the first task can be promptly scheduled to other computing units for execution, thereby improving the efficiency of the execution of the first task.

[0015] In one possible implementation, the processor may include, in addition to the first processor core, a second processor core, and the first processor core and the second processor core may be processor cores of the same type, or the first processor core and the second processor core may be processor cores of different types, such as the first processor core is a large core and the second processor core is a small core.

[0016] In the second aspect, the present application provides a task processing device, which is applied to a processor, wherein the first processor core in the processor includes a first computing unit and a second computing unit, and the first computing unit and the second computing unit are computing units of different types; the task processing device includes: an acquisition module, which is used to obtain a first task from multiple tasks to be executed; obtain the type of the first task; a scheduling module, which is used to schedule the first task to a first queue corresponding to the first computing unit when the type of the first task matches the type of the first computing unit; and schedule the first task to a second queue corresponding to the second computing unit when the type of the first task matches the type of the second computing unit.

[0017] In a possible implementation, the multiple tasks to be executed are tasks respectively executed by multiple threads included in a process of an application.

[0018] In a possible implementation, types of the first computing unit and the second computing unit include a matrix computing unit, a scalar computing unit, or a vector computing unit.

[0019] In one possible embodiment, the acquisition module is further used to acquire a second task from multiple tasks to be executed, and the second task includes a first operator and a second operator; the task processing device also includes: a division module, used to divide the second task into multiple subtasks, the multiple subtasks include a first subtask and a second subtask, the first subtask includes a first operator, and the second subtask includes a second operator; the scheduling module is further used to schedule the first subtask to the first queue, and the type of the first subtask matches the type of the first computing unit; and schedule the second subtask to the second queue, and the type of the second subtask matches the type of the second computing unit.

[0020] In a possible embodiment, the task processing device is specifically a first computing unit, and the first task is scheduled to the first queue corresponding to the first computing unit. The task processing device also includes: a determination module, which is used to determine the actual type of the first task when the first computing unit cannot execute the first task in the first queue; at this time, the scheduling module is also used to schedule the first task to the second queue when the actual type of the first task matches the type of the second computing unit.

[0021] In a possible implementation, the processor further includes a second processor core; the first processor core and the second processor core are processor cores of the same type, or the first processor core and the second processor core are processor cores of different types.

[0022] In a third aspect, the present application provides a processor for executing the task processing method in the first aspect or any implementation manner of the first aspect.

[0023] In a fourth aspect, the present application provides a computing device comprising a processor and a memory. The processor and the memory communicate with each other. The processor is configured to execute instructions stored in the memory so that the computing device performs a task processing method as described in the first aspect or any one of the implementations of the first aspect. It should be noted that the memory may be integrated into the processor or may be independent of the processor. The computing device may further comprise a bus. The processor is connected to the memory via the bus. The memory may comprise a readable memory and a random access memory.

[0024] In a fifth aspect, the present application provides a computer-readable storage medium, which stores instructions. When the computer-readable storage medium is run on a computing device, the computing device executes the operating steps of the task processing method described in the first aspect or any implementation of the first aspect.

[0025] In a sixth aspect, the present application provides a computer program product comprising instructions, which, when executed on a computing device, enables the computing device to execute the operating steps of the task processing method described in the first aspect or any one of the implementations of the first aspect.

[0026] Based on the implementation methods provided in the above aspects, this application can also be further combined to provide more implementation methods. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] FIG1 is a schematic diagram of the structure of an exemplary data processing system provided by the present application;

[0028] FIG2 is a schematic diagram of an exemplary process of creating multiple workers for multiple computing units on a single processor core provided by the present application;

[0029] FIG3 is a schematic diagram of another exemplary method of creating multiple workers for multiple computing units on a single processor core provided by the present application;

[0030] FIG4 is a flowchart of a task processing method provided by this application;

[0031] FIG5 is a schematic diagram of execution dependencies between multiple tasks provided by this application;

[0032] FIG6 is a schematic diagram of the present application providing consideration of tasks and calculation types of calculation units before and after the execution of tasks;

[0033] FIG7 is a flowchart of another task processing method provided by the present application;

[0034] FIG8 is a schematic structural diagram of a task processing device provided by the present application;

[0035] FIG9 is a schematic diagram of the hardware structure of a computing device provided in this application. DETAILED DESCRIPTION

[0036] The terms "first," "second," and so on, in the specification and claims of this application and the accompanying drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate and are merely used to describe the manner in which objects with the same attributes are distinguished in the embodiments of this application.

[0037] The technical solution in this application will be described below in conjunction with the drawings provided in this application.

[0038] 1 , which shows a schematic diagram of the structure of a data processing system. As shown in FIG1 , the data processing system 10 includes an application layer 101 , a hardware layer 102 , and a task processing device 200 .

[0039] The application layer 101 includes at least one application, such as an image recognition application, and is described in FIG1 using application 1 as an example. Furthermore, application 1 may include at least one programming model, and FIG1 uses programming models 1 to 3 as an example. Exemplarily, the programming model may be a message passing interface (MPI) model, a shared memory parallel programming (OpenMP) model, or a SYCL model (a high-level programming model for the open computing language (OpenCL)), or other types of programming models.

[0040] The hardware layer 102 includes multiple processor cores, each of which can execute thread tasks, and the multiple processor cores can be used to execute multiple thread tasks in parallel. The multiple processor cores in the hardware layer 102 can be processor cores of the same type; alternatively, the multiple processor cores can be processor cores of different types, such as some processor cores being large cores with higher processing performance, while others being small cores with lower processing performance. Each processor core can include multiple different types of computing units, such as computing units for scalar calculations, computing units for vector calculations, and computing units for matrix calculations. Each computing unit is a hardware unit on the processor core with computing capabilities. In actual applications, the computing unit can also be referred to as a hardware-based computing engine.

[0041] Furthermore, the hardware layer 102 may also include other devices, such as a network interface controller (NIC), memory, accelerator, or other types of devices, as shown in FIG1 . For example, the memory may be a direct memory access (DMA) device. The accelerator may be a graphics processing unit (GPU), for example.

[0042] The task processing device 200 may include multiple workers, which are used to schedule the processor cores in the hardware layer 102 to process multiple thread tasks in parallel. Among them, the workers can be implemented by threads. For example, the task processing device 200 can create multiple threads for the application 1 to execute multiple thread tasks. In actual application, the workers can also be implemented by other means such as processes, and this is not limited to this. Exemplarily, the task processing device 200 can be implemented by software. For example, the task processing device 200 can be specifically referred to as a runtime scheduler, and the runtime scheduler can run on the processor core, etc.

[0043] During the execution of application 1, at least one task, such as a data search task, may be generated based on user operations on application 1. For each task, the programming model can edit the task into one or more process tasks. These process tasks may have dependencies, such as the execution of process task A within the process tasks being dependent on the execution result of process task B (i.e., process task B must be executed first). The programming model can then generate multiple computation graphs 1 based on the dependencies between these multiple process tasks. Each node in computation graph 1 represents a process task, and directed edges between different nodes in computation graph 1 indicate the dependencies between different process tasks. Next, for each process task, the orchestration model can edit the process task into multiple thread tasks. These thread tasks have dependencies, such as the execution of thread task A within the thread tasks being dependent on the execution of thread task B being executed first. Based on the dependencies between these thread tasks, the programming model can then generate computation graph 2. In actual application scenarios, each process task may correspond to a computation graph 2. In this way, the programming model can use computation graphs 1 and 2 to orchestrate process tasks and thread tasks.

[0044] Based on the programming model's orchestration results for process tasks and thread tasks (e.g., computation graphs 1 and 2 described above), the task processing device can schedule processor cores in the hardware layer 102 to execute the multiple thread tasks included in each process task in parallel. The programming model can output these multiple thread tasks; accordingly, the task processing device 200 can add these multiple thread tasks to a queue so that each thread task can be subsequently scheduled for execution on a corresponding computing unit according to the order in which the thread tasks are queued.

[0045] In specific implementation, the task processing device 200 can create multiple threads for the various computing units included in each processor core. Each thread can serve as a worker to schedule a computing unit on the processor core to execute thread tasks. Different workers can schedule different computing units to execute different types of thread tasks.

[0046] For example, as shown in Figure 2, when a single processor core includes a computing unit 1 for scalar computation, a computing unit 2 for vector computation, a computing unit 3 for vector computation, and a computing unit 4 for matrix computation, the task processing apparatus 200 can create four workers for each computing unit in the processor core, namely Worker 1, Worker 2, Worker 3, and Worker 4. Worker 1 is used to schedule computing unit 1 to process scalar computation thread tasks, such as data movement tasks; Worker 2 and Worker 3 are used to schedule computing unit 2 and computing unit 3, respectively, to process vector computation thread tasks, such as data access tasks; and Worker 4 is used to schedule computing unit 4 to process matrix computation thread tasks, such as convolution tasks. Furthermore, each computing unit can be configured with a queue, such as Queue 1 for computing unit 1 and Queue 4 for computing unit 4. Each computing unit's queue is used to store the thread tasks required to be executed by that computing unit. Specifically, a worker can use that computing unit to execute the thread tasks in the queue.

[0047] Furthermore, in addition to including multiple computing units and queues, the processor core may also include other hardware such as cache and registers, so that each worker can also use multiple hardware such as computing units, cache and registers to process the same thread task at the same time.

[0048] In the processor core shown in Figure 2, the computing units used by different workers can be independent of each other, that is, a single computing unit can only be used by one worker. Alternatively, in the processor core shown in Figure 3, different workers can use multiple computing units, that is, a single computing unit can be used by multiple workers.

[0049] In this way, for each processor core in the hardware layer 102 , the task processing device 200 can create multiple workers for the processor core.

[0050] After creating multiple workers, the task processing device 200 can schedule the multiple thread tasks obtained through the programming model to be executed in parallel by computing units on one or more processor cores, thereby improving the execution efficiency of the multiple thread tasks. Specifically, when scheduling each thread task to be executed on a corresponding computing unit, the task processing device 200 can schedule the thread task to a queue corresponding to a computing unit on the processor core that is used to execute thread tasks of that type, based on the type of the thread task, so that the computing unit can execute the thread tasks in the queue.

[0051] For example, the task processing device 200 may pre-create workers 1, 2, and 3 for computing unit 1 for scalar calculations, computing unit 2 for vector calculations, and computing unit 3 for matrix calculations on processor core A, and create workers 4, 5, and 6 for computing unit 4 for scalar calculations, computing unit 5 for vector calculations, and computing unit 6 for matrix calculations on processor core B. Workers 1 and 4 are responsible for processing scalar calculation tasks, workers 2 and 5 are responsible for processing vector calculation tasks, and workers 3 and 6 are responsible for processing matrix calculation tasks.

[0052] Assume that the multiple tasks to be executed include a first task. This first task can be, for example, a thread task generated by a programming model. Assume that the first task is a thread task of the matrix calculation type. The task processing device 200 can, based on the type of the first task, schedule the first task to queue 3 corresponding to computing unit 3 on processor core A. Thus, worker 3 can use computing unit 3 to execute the first task in queue 3. Alternatively, the task processing device 200 can schedule the first task to queue 6 corresponding to computing unit 6 on processor core B based on the type of the first task. Thus, worker 6 can use computing unit 6 to execute the first task in queue 6.

[0053] Similarly, for other tasks among the multiple tasks to be executed, the task processing device 200 can also schedule other tasks to processor core A or processor core B according to the type of other tasks, and put them into the queue corresponding to the computing unit for executing tasks of this type, so that the computing unit can subsequently execute the tasks in the queue.

[0054] During the execution of a task, the task processing device 200 will schedule the task to a queue of a computing unit that matches the type of the (thread) task, so that the computing unit can execute the task in the queue. Since, in general, the multiple tasks to be executed are usually tasks of different types, the multiple computing units on the processor core can usually be assigned to tasks and execute the tasks, which enables the multiple computing units on the processor core to execute multiple tasks of different types in parallel, thereby avoiding as much as possible that some computing units on the processor core are idle for a long time, thereby making full use of the resources on the processor core and improving the resource utilization of the processor core.

[0055] In addition, different workers are responsible for processing tasks of different computing types, which ensures that each worker usually has sufficient computing unit resources to perform tasks, avoiding the problem of multiple workers using the same computing unit on the same processor core to process multiple tasks of the same type at the same time, resulting in low execution efficiency of the multiple tasks. This enables hardware-level concurrent execution of multiple tasks, thereby improving the overall efficiency of the processor core in executing multiple tasks.

[0056] It is worth noting that the data processing system 10 shown in FIG1 is only an exemplary illustration and is not intended to be limiting. For example, in actual application scenarios, the data processing system 10 also includes parts such as the kernel of the operating system (not shown in FIG1 ). For another example, in other data processing systems, multiple servers may be included, each of which may adopt the multi-core architecture shown in FIG1 and perform tasks based on the above method. In this case, the data processing system can be applied to supercomputing scenarios. For another example, in other data processing systems, the application layer 101 may include a larger number of applications, and the types and quantities of programming models included in different applications may vary; or, the hardware layer 102 may also include other types or other quantities of hardware. Alternatively, in other data processing systems, the application layer 101 may not include a programming model, so that the application in the application layer 101 can directly generate tasks and send them to the task processing device 200.

[0057] For ease of understanding, an embodiment of the task processing method provided in this application is described below with reference to the accompanying drawings.

[0058] Referring to FIG4 , FIG4 is a flow chart illustrating a task processing method provided in an embodiment of the present application. This method can be applied to the data processing system 10 shown in FIG1 , or can be applied to other applicable data processing systems. For ease of explanation, this embodiment is described using the data processing system 10 shown in FIG1 as an example.

[0059] The task processing method shown in FIG4 may specifically include:

[0060] S401 : The task processing apparatus 200 obtains a task 1 to be executed from a plurality of tasks to be executed.

[0061] In this embodiment, after the programming model (such as programming model 1, etc.) edits each process task to obtain corresponding multiple thread tasks and the calculation graph corresponding to the multiple thread tasks, the programming model can determine the thread tasks that can be currently executed based on the dependency relationship between different thread tasks indicated by the calculation graph.

[0062] For example, the computation graph generated by the programming model can be shown in FIG5. FIG5 includes a plurality of nodes, each node is used to indicate a thread task; the directed edges between different nodes are used to indicate the execution dependencies between different thread tasks. For example, the directed edge between node 1 and node 3 is used to indicate that the execution of the thread task identified by node 3 depends on the thread task identified by node 1 to be executed first. Then, the task processing device 200 can determine at least one thread task that can be currently executed based on the execution dependencies between the multiple thread tasks indicated by the computation graph, such as determining the thread task indicated by the leaf node (such as node 4, 6, 8, or 9) in the computation graph as the thread task that can be currently executed.

[0063] The programming model then sends at least one currently executable thread task to the operating system kernel via a corresponding interface, requesting the operating system kernel to schedule resources to execute the at least one thread task. Accordingly, the task processing device 200 can continuously monitor the interface (which can be one or more); and when the programming model outputs at least one executable thread task via the interface, the task processing device 200 can intercept the thread task sent by the programming model to the kernel, so that the task processing device 200 can subsequently perform resource scheduling for the thread task, such as scheduling a processor core to execute the thread task.

[0064] Alternatively, after the programming model obtains multiple thread tasks through programming and generates a computational graph, the task processing device 200 can detect the thread tasks that can be currently executed based on the dependency relationship between different thread tasks indicated by the computational graph, that is, detect the thread tasks that are currently not dependent on other thread tasks to be executed first.

[0065] Assume that there is currently an executable thread task, which is referred to as task 1 below.

[0066] Furthermore, the task processing device 200 can also determine the currently executed task 1 based on the priority of each task (used to indicate the priority of task execution). For example, in the computation graph shown in FIG5 , the tasks indicated by nodes 4, 6, 8, and 9 are currently executable tasks. Then, the task processing device 200 can determine, from among these four tasks, that the task indicated by node 8 is the currently executed priority task (the task indicated by node 8 has a higher priority), i.e., task 1 described in step S401.

[0067] In specific implementation, the multiple (thread) tasks generated by the programming model can carry indication information of the priority of the task. For example, technicians / users can define that tasks containing data movement operators are executed with a higher priority, and define that tasks containing data access operators are executed with a lower priority, etc., so that the programming model can add priority to the task when generating multiple tasks.

[0068] Alternatively, the priorities of the multiple tasks generated by the programming model can be obtained by the task processing device 200 through analysis. For example, when the task processing device 200 obtains the multiple (thread) tasks sent by the programming model, it can also obtain an identifier indicating the calculation type to which the task belongs, so that the task processing device 200 can determine the priority of the task to be executed based on the task type indicated by the identifier. For example, when the task type is a scalar calculation type, the task processing device 200 can determine that the task is executed with a higher priority, while when the task type is a matrix calculation type, the task processing device 200 can determine that the task is executed with a lower priority. Then, the task processing device 200 can add indication information for each task to indicate the priority level. For example, the task processing device 200 can add a priority identifier to the task. When the value of the priority identifier is "high" or "1", it indicates that the task is executed with a higher priority, and when the value of the priority identifier is "low" or "0", it indicates that the task is executed with a lower priority.

[0069] Accordingly, after acquiring multiple tasks, the task processing device 200 can determine the currently executable task with a higher priority from the multiple tasks according to the execution priority of each task. It is assumed that the determined task is task 1 described in step S401.

[0070] S402: The task processing device 200 obtains the type of task 1.

[0071] After determining the task 1 to be executed currently, the task processing device 200 can obtain the type of task 1. Among them, the thread tasks generated by the programming model can be tasks of multiple types. Exemplarily, the type of task refers to the calculation type to which the task belongs, for example, it can be a scalar calculation type, a vector calculation type, or a matrix calculation type, etc., or it can be other types. Among them, a task of scalar calculation type refers to a task whose data calculation process is mainly scalar calculation, such as a data movement task, etc. A task of vector calculation type refers to a task whose data calculation process is mainly vector calculation, such as a data access task, etc. A task of matrix calculation type refers to a task whose data calculation process is mainly matrix calculation, such as a convolution calculation task, etc.

[0072] Exemplarily, the task processing apparatus 200 may determine the type of the task based on the operators included in task 1. For example, when task 1 includes a data access operator, the type of task 1 may be determined to be a scalar calculation type; and when task 1 includes a data movement operator, the type of task 1 may be determined to be a scalar calculation type.

[0073] Alternatively, when generating Task 1, the programming model may add a type identifier 1 to Task 1, where the type identifier 1 indicates the type of computation to which Task 1 belongs. Thus, after acquiring Task 1, the task processing apparatus 200 can determine the type of Task 1 based on the type identifier 1 carried by Task 1.

[0074] S403 : When the type of task 1 matches the type of computing unit 1 on the processor core, the task processing device 200 schedules task 1 to queue 1 corresponding to computing unit 1 .

[0075] S404 : When the type of task 1 matches the type of computing unit 2 on the processor core, the task processing device 200 schedules task 1 to queue 2 corresponding to computing unit 2 .

[0076] A processor core usually includes multiple different computing units, and different computing units are suitable for executing different types of tasks. Therefore, after obtaining task 1 and the type of task 1, the task processing device 200 can match the type of task 1 with the types of each computing unit on the processor core, and schedule task 1 to the queue configured by the appropriate computing unit. Specifically, when the type of task 1 matches the type of computing unit 1, task 1 is scheduled to queue 1 of computing unit 1, and when the type of task 1 matches the type of computing unit 2, task 1 is scheduled to queue 2 of computing unit 2.

[0077] In this way, each computing unit can execute the tasks in the queue corresponding to the computing unit under the scheduling of the worker.

[0078] In this embodiment, the task processing device 200 can pre-create multiple workers for the various computing units on the processor core 1. Each worker is responsible for scheduling a computing unit on the processor core 1 to execute a task. In actual application scenarios, the number of each computing unit on the processor core 1 can be one or more.

[0079] In a specific implementation, before processing tasks issued by the programming model, the task processing device 200 can pre-manage the hardware in the hardware layer 102. For example, during the initialization of the task processing device 200, all processor cores in the application layer 102 can be managed to obtain hardware description information for each processor core included in the hardware layer 102. The task processing device 200 can then determine the various computing units included in each processor core based on the hardware description information. Thus, the task processing device 200 can create a worker for each computing unit on the processor core, such as creating worker 1 for computing unit 1 on processor core 1, creating worker 2 for computing unit 2 on processor core 1, and so on, and establish a mapping relationship between workers and computing units. Next, the task processing device 200 can add a type identifier to each worker based on the type of computing unit corresponding to each worker to indicate the type of task that the worker can process when scheduling the computing unit.

[0080] Then, the task processing device 200 can determine which computing unit type matches the type of task 1. Taking the example of matching the type of task 1 with the type of the first computing unit, the task processing device 200 can determine the worker 1 for processing task 1 from the multiple pre-created workers, and based on the mapping relationship between worker 1 and computing unit 1 on processor core 1, use the worker 1 to schedule computing unit 1 on processor core 1 to execute task 1.

[0081] For ease of understanding and explanation, the task processing process is described below by taking the task processing device 200 using worker 1 to process task 1.

[0082] In a first possible implementation, task processing device 200 may be, for example, a scheduler running on a processor. The scheduler may be implemented in software, such as a program running on the processor. In this case, the scheduler may determine, based on type identifier 1 of task 1, worker 1 from among multiple workers created for processor core 1, where the type identifier of worker 1 matches type identifier 1 of task 1.

[0083] The scheduler can be configured with a target queue, which can be, for example, a queue that supports the first-in-first-out (FIFO) principle. Thus, the scheduler can add multiple thread tasks to be executed and the type identifier of each thread task to the target queue according to the execution dependencies between the multiple thread tasks (and the priority of the tasks to be executed). The scheduler then sequentially schedules the multiple thread tasks in the target queue to different computing units for execution. Accordingly, the scheduler can obtain task 1 and the type identifier 1 of task 1 from the queue.

[0084] The scheduler then schedules each thread task in the target queue one by one according to the order in which the thread tasks are in the target queue. For example, if the currently scheduled thread task is Task 1, the scheduler matches Task 1's type identifier 1 with the type identifiers of each worker and identifies the worker with the successful match as Worker 1. In practice, Worker 1, which successfully matches, is currently idle; that is, it is not currently executing tasks using a compute unit.

[0085] It will be appreciated that since different processor cores can include computing units with the same functionality, such as computing units for scalar computations, the scheduler can create multiple workers for each processor core, each of which is responsible for processing tasks of computation type 1 indicated by type identifier 1. Therefore, after obtaining type identifier 1, the scheduler can determine whether the worker corresponding to processor core 1, which is responsible for executing tasks of computation type 1, is idle. If so, the scheduler can identify that worker as worker 1. If not, the scheduler can continue to determine whether the worker corresponding to processor core 2, which is responsible for executing tasks of computation type 1, is idle. If so, the scheduler can identify the worker corresponding to processor core 2 as worker 1. If not, the scheduler can continue to determine whether the workers corresponding to the remaining processor cores, which are responsible for executing tasks of computation type 1, are idle. This process continues in this manner until the scheduler identifies worker 1. In this embodiment, worker 1 is set as the worker created for computation unit 1 on processor core 1 for illustration purposes.

[0086] When the computing type indicated by type identifier 1 matches the computing type to which task 1 actually belongs, the scheduler schedules task 1 to queue 1 corresponding to computing unit 1 that worker 1 is responsible for scheduling, so that worker 1 schedules computing unit 1 to execute task 1 in queue 1, as shown in Figure 4.

[0087] In actual application scenarios, the calculation type indicated by the type identifier 1 added by the programming model to task 1 may be consistent with the actual calculation type of task 1, or it may be inconsistent with the calculation type of task 1 (for example, the programming model has marked the wrong type identifier for task 1, etc.). Therefore, before using worker 1 to execute task 1, the scheduler can first determine whether the calculation type indicated by type identifier 1 matches the actual calculation type of task 1. For example, it can be verified based on the operators included in task 1 whether the actual calculation type of task 1 is the calculation type indicated by type identifier 1. If they match, it indicates that the type identifier marked by the programming model for task 1 is correct, and the scheduler can store task 1 in queue 1 configured for computing unit 1; accordingly, worker 1 can take task 1 out of queue 1 and call computing unit 1 on processor core 1 to execute task 1. If there is a mismatch, it means that the type identifier marked by the programming model for task 1 is incorrect. The scheduler can determine the actual computing type of task 1 based on the operators included in task 1, and store task 1 in the queue corresponding to the computing unit on processor core 1 or other processor cores that can be used to execute the task, so that other computing units can execute task 1, thereby ensuring the execution efficiency of task 1 by scheduling task 1 to the appropriate computing unit.

[0088] In practical applications, in addition to the scheduler verifying the correctness of the type identifier 1 assigned to task 1 by the programming model, a worker (or compute unit) can also verify the correctness of task 1's type identifier 1. In specific implementations, the scheduler can first add task 1 to queue 1 configured for compute unit 1 based on type identifier 1. Then, during execution, worker 1 can retrieve task 1 from queue 1 and determine whether the computation type indicated by type identifier 1 matches the actual computation type of task 1. If so, worker 1 can schedule compute unit 1 on processor core 1 to execute task 1. If not, worker 1 can schedule task 1 from queue 1 to queue 2 assigned to compute unit 2. The computation type of tasks that compute unit 2 can process is consistent with the actual computation type of task 1, and compute unit 2 is currently not executing any tasks. Therefore, after determining that the computation type indicated by type identifier 1 matches the actual computation type of task 1, worker 2 schedules compute unit 2 to execute task 1, as shown in Figure 4. Among them, worker 2 is a worker created by the task processing device 200 for the computing unit 2 on the processor core 1.

[0089] Furthermore, if computing unit 2 on processor core 1 is currently executing a task, worker 1 can dispatch task 1 to the queue corresponding to the computing unit on another processor core, where the computing unit on that other processor core can execute the same type of task as task 1. For example, worker 1 can determine whether computing unit 3 on processor core 2, which is capable of executing task 1, is currently executing a task. If not, worker 1 can dispatch task 1 to computing unit 3 on processor core 2, so that computing unit 3 on processor core 2 can execute task 1. If so, worker 1 can continue to determine the computing units on other processor cores that are capable of executing tasks of the same type, using the same approach as above.

[0090] In actual application, each worker can only call one computing unit on the processor core. Alternatively, each worker can call multiple computing units on the processor core. For example, when task 1 contains multiple scalar calculation type operators a and one matrix calculation type operator b at the same time, when worker 1 calls computing unit 1 for scalar calculation to execute task 1, when operator b is executed, worker 1 can call computing unit 4 for matrix calculation on processor core 1 to execute operator b in task 1. Accordingly, the worker 4 created by the task processing device 200 for the computing unit 4, in addition to being able to call computing unit 4 to process matrix calculation type tasks, can also call computing unit 1 to process the scalar calculation type operators contained in the task.

[0091] In a second possible implementation, the task 1 to be executed may be scheduled by a worker (or the computing unit running the worker). In this case, in the process of scheduling task 1, the task processing device may specifically be worker 1 (or the computing unit running the worker 1). The task processing device 200 may add the multiple tasks to be executed and the type identifier of each task to the target queue according to the execution dependencies (and priorities) between the multiple tasks to be executed. In this way, after scheduling computing unit 1 to execute the completed task, worker 1 may take out a task from the target queue, taking the currently taken out task 1 as an example. Then, worker 1 may determine whether the type identifier 1 of the task matches the type identifier of worker 1. If they match, worker 1 may store task 1 in queue 1 corresponding to computing unit 1; if they do not match, worker 1 may store task 1 in queues corresponding to other computing units based on the type identifier 1.

[0092] In actual application scenarios, it is possible that after task 1 is dispatched to queue 1, computing unit 1 is unable to execute task 1 in queue 1, for example, task 1 is actually a matrix calculation type task, while computing unit 1 is a computing unit for scalar calculation, so computing unit 1 is unable to execute task 1. Therefore, in a further possible implementation, after task 1 is dispatched to queue 1, when computing unit 1 is unable to execute task 1 in queue 1, computing unit 1 can determine the actual type of task 1 and determine the computing unit that can be used to execute task 1 based on the actual type of task 1. Assuming that computing unit 1 determines that the actual type of task 1 matches the type of computing unit 2, computing unit 1 can dispatch task 1 to queue 2 corresponding to computing unit 2 so that computing unit 2 can execute task 1. In this way, when computing unit 1 is unable to execute task 1, it can promptly dispatch task 1 to other computing units for execution, thereby improving the efficiency of task 1 execution.

[0093] As an implementation example, before executing Task 1, Worker 1 may also first determine whether the computation type indicated by Type Identifier 1 matches the actual computation type of Task 1. If so, Worker 1 may call Compute Unit 1 to execute Task 1 in Queue 1. If not, Worker 1 may determine the actual computation type of Task 1 based on the operators included in Task 1 and schedule Task 1 to Queue 2 corresponding to Compute Unit 2 on Processor Core 1 or other processor cores that can be used to execute Task 1, so that Compute Unit 2 can execute Task 1. This ensures the execution efficiency of Task 1 by scheduling Task 1 to the appropriate Compute Unit.

[0094] Similarly, for other tasks among the multiple tasks, such as Task 2, the task processing device 200 can continue to obtain Task 2 and the type of Task 2 from the multiple tasks to be executed in the same manner as described above, and thus schedule Task 2 to the queue corresponding to the corresponding computing unit of matching type according to the type of Task 2. For example, when the task processing device 200 is a scheduler, the scheduler can take out a new task, namely Task 2, from the target queue, and schedule Task 2 to the queue corresponding to the corresponding computing unit according to the type of Task 2. The process of the task processing device 200 scheduling and executing Task 2 using the computing unit is similar to the process of scheduling and executing Task 1 described above, and will not be elaborated here.

[0095] In this way, during the execution of a task, the task processing device 200 will use workers that match the computing type to schedule the computing unit corresponding to the computing type on the processor core 1 or other processor cores to execute the task. In this way, for multiple thread tasks of different types (such as the above-mentioned task 1 and task 2), the task processing device can use different workers to call different types of computing units on the processor core 1 to execute different types of thread tasks in parallel, so as to avoid as much as possible that some computing units on the processor core 1 are idle for a long time, thereby improving the resource utilization of the processor core 1. As shown in Figure 6, the task processing device 200 can simultaneously use the computing unit 1 and the computing unit 2 on the processor core 1 to execute tasks 1 and 2 concurrently, avoiding the situation in which the resources of the computing unit 2 are idle during the execution of task 1 and causing task 2 to be in a state of waiting for execution for a long time.

[0096] In addition, in this embodiment, different workers can be responsible for processing tasks of different computing types, which enables each worker to usually have sufficient computing unit resources to perform tasks, avoiding the problem of multiple workers using the same computing unit on processor core 1 to process multiple tasks at the same time, resulting in low execution efficiency of the multiple tasks. In this way, the overall efficiency of the processor core 1 in executing multiple tasks in parallel can be improved, thereby improving the business operation performance of application 1.

[0097] When the task processing device 200 identifies an executable task from multiple (threaded) tasks, and after determining that worker 1 has completed task 1, the task processing device 200 can update the execution status of task 1 in the computation graph (e.g., mark it as "executed"), and continue to obtain one or more new executable tasks from the computation graph and add the new tasks to the target queue. At the same time, the task processing device 200 can continue to extract tasks from the target queue and, based on the type identifier of the extracted task, assign the task to the corresponding worker for execution.

[0098] It should be noted that the method embodiment shown in FIG4 is merely an example. Based on the method flow shown in FIG4 , the process of the task processing device 200 using multiple workers to execute tasks may also adopt the following embodiments.

[0099] In Example 1, in the embodiment shown in FIG4 , the task processing device 200 assigns Task 1 to the corresponding worker based on the type identifier added to Task 1 according to the programming model. In other embodiments, after obtaining a task, the task processing device 200 can analyze the actual computation type of the task based on the operators included in the task, and assign the task to the worker responsible for processing tasks of that computation type based on the respectively obtained computation type. In this way, after being assigned a task, each worker can directly call the corresponding computation unit on the processor core to execute the task, without having to push the task to other workers based on the computation type of the task.

[0100] Example 2, in the embodiment shown in FIG4 , is illustrated by taking the task processing device 200 as an example to schedule each task to the queue corresponding to the corresponding computing unit according to the computing graph. In other embodiments, after the task processing device 200 generates the computing graph, each worker in an idle state can obtain the currently executable task according to the computing graph, and when it is determined that the computing type to which the task belongs matches the computing type to which the worker itself can process the task belongs, the worker can obtain the task and schedule the corresponding computing unit on the processor core to execute the task. Furthermore, after the worker completes the execution of the current task, the worker can update the execution status of the task in the computing graph (such as marking it as "executed", etc.), and continue to obtain new unexecuted tasks from the computing graph.

[0101] The embodiment shown in FIG4 is described above using one worker to process Task 1. In other embodiments, the task processing device 200 may utilize multiple workers to collaboratively process the same task, thereby improving the efficiency of processing the task. This will be described in detail below with reference to FIG7.

[0102] Referring to Figure 7, a flow chart of another task processing method is shown. As shown in Figure 7, the method may specifically include:

[0103] S701: The task processing device 200 obtains task 2 to be executed.

[0104] Task 2 includes multiple operators. In this embodiment, the description is made by taking an example where Task 2 includes Operator 1 and Operator 2. Furthermore, Operator 1 and Operator 2 are operators of the same calculation type, specifically, vector calculation type operators.

[0105] Among them, the specific implementation method of the task processing device 200 obtaining task 2 is similar to the implementation method of the task processing device 200 obtaining task 1 in the embodiment shown in Figure 4 above. For details, please refer to the above-mentioned relevant descriptions and will not be repeated here.

[0106] S702: The task processing device 200 divides task 2 into subtask 1 and subtask 2, wherein subtask 1 includes operator 1 of a vector calculation type, and subtask 2 includes operator 2 of a vector calculation type.

[0107] In this embodiment, after obtaining Task 2, the task processing device 200 may split Task 2 into multiple subtasks, each subtask including an operator.

[0108] As an implementation example, the task processing device 200 can divide the execution logic of Task 2 into blocks, with each block of execution tasks serving as a subtask of Task 2, thereby splitting Task 2 into multiple subtasks. Each block of execution logic can be an operator in Task 1, allowing the task processing device 200 to split Task 2 into multiple subtasks at the operator granularity.

[0109] Among them, subtask 1 and subtask 2 obtained by dividing task 2 can be executed in parallel, that is, the execution of subtask 1 and the execution of subtask 2 can be independent of each other.

[0110] S703 : The task processing apparatus 200 determines the type of subtask 1 and the type of subtask 2 .

[0111] For example, the task processing device 200 can determine the type of each subtask, that is, the calculation type of the subtask, based on the operators included in each subtask. That is, subtask 1 and subtask 2 both belong to the vector calculation type.

[0112] S704: The task processing device 200 stores subtask 1 in queue 3 corresponding to computing unit 3 scheduled by worker 3 according to the type of subtask 1, and stores subtask 2 in queue 4 corresponding to computing unit 4 scheduled by worker 4 according to the type of subtask 2.

[0113] Among them, Worker 3 and Worker 4 are both workers that perform the same vector calculation type task.

[0114] In a specific implementation, the task processing device 200 can pre-create corresponding workers for various computing units on the processor core 1 and for various computing units on the processor core 2, and add a corresponding identifier for each worker, which is used to indicate the type of calculation to which the worker calls the computing unit to process the task. For example, assuming that the processor core 1 includes a computing unit 1 for scalar calculation, a computing unit 2 for matrix calculation, and a computing unit 3 for vector calculation, the task processing device 200 can create worker 1 for the computing unit 1, worker 2 for the computing unit 2, and worker 3 for the computing unit 3, and add an identifier of the computing type for each worker. At the same time, the processor core 2 also includes a computing unit 4 for vector calculation, a computing unit 5 for matrix calculation, and a computing unit 6 for scalar calculation. The task processing device 200 can create worker 4 for the computing unit 4, worker 5 for the computing unit 5, and worker 6 for the computing unit 6, and add an identifier of the computing type for each worker.

[0115] After determining the types of subtasks 1 and 2, task processing device 200 can match the types of each subtask with the types of multiple pre-created workers. Specifically, the subtask type identifiers can be matched with the worker type identifiers. Assuming that the type of subtask 1 matches the type of worker 3, and the type of subtask 2 matches the type of worker 4, task processing device 200 can schedule subtask 1 to the queue corresponding to computing unit 3 and schedule subtask 2 to the queue corresponding to computing unit 4.

[0116] S705 : The task processing apparatus 200 uses worker 3 to schedule computing unit 3 on processor core 1 to execute subtask 1 , and uses worker 4 to schedule computing unit 4 on processor core 2 to execute subtask 2 .

[0117] In this way, task processing device 200 can improve the overall execution efficiency of Task 2 by splitting Task 2 into multiple subtasks and using multiple workers to schedule computing units on different processor cores to execute the multiple subtasks in parallel. Furthermore, when computing unit 4 on processor core 2 is idle, it can collaborate with computing unit 3 on processor core 1 to execute the same task, thereby improving resource utilization on processor core 2.

[0118] It is worth noting that the task processing method shown in FIG7 is merely an example and is not intended to be limiting. For example, in other embodiments, when a thread task issued by a programming model includes multiple operators of different computing types, the task processing device 200 may also refer to the method flow shown in FIG7 and utilize different types of computing units on one or more processor cores to execute the different subtasks included in the task in parallel, thereby improving the overall efficiency of task execution.

[0119] It is worth noting that other reasonable step combinations that can be thought of by those skilled in the art based on the above description also fall within the scope of protection of this application. Secondly, those skilled in the art should also be familiar with that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily required by this application.

[0120] The task processing method provided in the embodiment of the present application is introduced above with reference to FIG. 1 to FIG. 7 . Next, the structure of the task processing apparatus and computing device provided in the embodiment of the present application is introduced with reference to the accompanying drawings.

[0121] Referring to Figure 8, a structural schematic diagram of a task processing device is shown. The task processing device 800 shown in Figure 8 is applied to a processor, which includes at least one processor core. The at least one processor core includes a first processor core. The first processor core includes a first computing unit and a second computing unit. The first computing unit and the second computing unit are computing units of different types.

[0122] As shown in FIG8 , the task processing device 800 includes:

[0123] The acquisition module 801 is configured to acquire a first task from a plurality of tasks to be executed; and acquire a type of the first task;

[0124] Scheduling module 802 is used to schedule the first task to the first queue corresponding to the first computing unit when the type of the first task matches the type of the first computing unit; and to schedule the first task to the second queue corresponding to the second computing unit when the type of the first task matches the type of the second computing unit.

[0125] In a possible implementation, the multiple tasks to be executed are tasks respectively executed by multiple threads included in a process of an application.

[0126] In a possible implementation, types of the first computing unit and the second computing unit include a matrix computing unit, a scalar computing unit, or a vector computing unit.

[0127] In a possible implementation, the acquisition module 801 is further configured to acquire a second task from the plurality of tasks to be executed, where the second task includes the first operator and the second operator;

[0128] The task processing device 800 further includes:

[0129] a division module 803, configured to divide the second task into a plurality of subtasks, the plurality of subtasks including a first subtask and a second subtask, the first subtask including a first operator, and the second subtask including a second operator;

[0130] The scheduling module 802 is further configured to:

[0131] Dispatching a first subtask to a first queue, where a type of the first subtask matches a type of the first computing unit;

[0132] The second subtask is scheduled to the second queue, and the type of the second subtask matches the type of the second computing unit.

[0133] In a possible implementation, the task processing device 800 is specifically a first computing unit, and the first task is scheduled to a first queue corresponding to the first computing unit;

[0134] The task processing device 800 further includes:

[0135] A determination module 804 is configured to determine the actual type of the first task when the first computing unit cannot execute the first task in the first queue;

[0136] The scheduling module 802 is further configured to schedule the first task to the second queue when the actual type of the first task matches the type of the second computing unit.

[0137] In a possible implementation, the processor further includes a second processor core; the first processor core and the second processor core are processor cores of the same type, or the first processor core and the second processor core are processor cores of different types.

[0138] Since the task processing device 800 shown in Figure 8 corresponds to the task processing device 200 in the embodiment shown in Figure 4 or Figure 7 above, the specific implementation method of the task processing device 800 shown in Figure 8 and its technical effects can be found in the relevant description of the embodiment shown in Figure 4 or Figure 7 above, and will not be repeated here.

[0139] FIG9 is a schematic diagram of the hardware structure of a computing device 900 provided in the present application. The computing device 900 can, for example, implement the task processing apparatus 200 in the embodiment shown in FIG4 or FIG7 .

[0140] As shown in Figure 9, the computing device 900 includes a processor 901, a memory 902, and a communication interface 903. The processor 901, the memory 902, and the communication interface 903 communicate via a bus 904, and may also communicate via other means such as wireless transmission. The memory 902 is used to store instructions, and the processor 901 is used to execute the instructions stored in the memory 902. Furthermore, the computing device 900 may also include a memory unit 905, and the memory unit 905 may be connected to the processor 901, the storage medium 902, and the communication interface 903 via a bus 904. The memory 902 stores program code, and the processor 901 may call the program code stored in the memory 902 to perform the following operations:

[0141] Obtaining a first task from a plurality of tasks to be executed;

[0142] Get the type of the first task;

[0143] When the type of the first task matches the type of the first computing unit included in the first processor core, scheduling the first task to a first queue corresponding to the first computing unit;

[0144] When the type of the first task matches the type of the second computing unit included in the first processor core, the first task is scheduled to the second queue corresponding to the second computing unit, wherein the first computing unit and the second computing unit are computing units of different types.

[0145] It should be understood that in this embodiment, the processor 901 may be a CPU, or may be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete device components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc.

[0146] The memory 902 may include a read-only memory and a random access memory, and provides instructions and data to the processor 901. The memory 902 may also include a nonvolatile random access memory.

[0147] The memory 902 may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. The non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic random access memory (DRAM), synchronous DRAM (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link DRAM (SLDRAM), and direct rambus RAM (DR RAM).

[0148] The communication interface 903 is used to communicate with other devices connected to the computing device 900. In addition to the data bus, the bus 904 may also include a power bus, a control bus, and a status signal bus. However, for the sake of clarity, various buses are labeled as bus 904 in the figure.

[0149] It should be understood that the computing device 900 according to the embodiment of the present application may correspond to the task processing device 800 in the embodiment of the present application, and may correspond to the method executed by the task processing device 200 in the method shown in Figure 4 or Figure 7 in the embodiment of the present application. The above-mentioned and other operations and / or functions implemented by the computing device 900 are respectively for implementing the process of the corresponding method in Figure 4 or Figure 7. For the sake of brevity, they will not be repeated here.

[0150] The present application also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium that can be stored by a computing device, or a data storage device such as a data center that contains one or more available media. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a magnetic tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive). The computer-readable storage medium includes instructions that instruct the computing device to execute the above-mentioned task processing method.

[0151] The present application also provides a computer program product. The computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computing device, the computer program product fully or partially generates the process or function described in the present application.

[0152] The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions may be transmitted from one website, computer, or data center to another website, computer, or data center via wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means.

[0153] The computer program product may be a software installation package. When any of the aforementioned task processing methods is required, the computer program product may be downloaded and executed on a computing device.

[0154] The above embodiments can be implemented in whole or in part by software, hardware, firmware or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded or executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that contains one or more available media sets. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state drive.

[0155] The terms used in the above embodiments are only for the purpose of describing specific embodiments and are not intended to limit the present application. As used in the specification of this application and the appended claims, the singular expressions "a", "an", "said", "above", "the" and "this" are intended to also include expressions such as "one or more", unless the context clearly indicates otherwise. It should also be understood that in the embodiments of the present application, "one or more" refers to one, two or more; the character " / " generally indicates that the objects before and after are in an "or" relationship. In the embodiments of the present application. "Simultaneously" refers to the same time period, including the situation at the same moment.

[0156] References to "one embodiment" or "some embodiments" in this specification mean that a particular feature, structure, or characteristic described in conjunction with that embodiment is included in one or more embodiments of the present application. Thus, phrases such as "in one embodiment," "in some embodiments," "in other embodiments," and "in yet other embodiments" appearing in various places in this specification do not necessarily refer to the same embodiment, but rather mean "one or more but not all embodiments," unless otherwise specifically emphasized. The terms "including," "comprising," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.

[0157] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present application, and such modifications or substitutions should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.

Claims

1. A task processing method, applied to a processor, characterized in that, The first processor core in the processor includes a first computing unit and a second computing unit, and the first computing unit and the second computing unit are different types of computing units; The method includes: Obtain a first task from multiple tasks to be executed; Obtain the type of the first task; When the type of the first task matches the type of the first computing unit, schedule the first task to a first queue corresponding to the first computing unit; When the type of the first task matches the type of the second computing unit, schedule the first task to a second queue corresponding to the second computing unit.

2. The method according to claim 1, wherein The method is executed by a scheduler running in the processor, and the scheduler is configured with a target queue that stores the multiple tasks to be executed; The obtaining a first task from multiple tasks to be executed includes: The scheduler obtains the first task from the target queue; The scheduling the first task to a first queue corresponding to the first computing unit includes: The scheduler stores the first task in the first queue; The scheduling the first task to a second queue corresponding to the second computing unit includes: The scheduler stores the first task in the second queue.

3. The method according to claim 1, characterized in that, The method is executed by the first computing unit; The obtaining a first task from multiple tasks to be executed includes: After the first computing unit finishes executing the tasks in the first queue, obtain the first task from the multiple tasks to be executed; The scheduling the first task to a first queue corresponding to the first computing unit includes: The first computing unit stores the first task in the first queue; The scheduling the first task to a second queue corresponding to the second computing unit includes: The first computing unit stores the first task in the second queue.

4. The method according to any one of claims 1 to 3, characterized in that, The multiple tasks to be executed are tasks respectively executed by multiple threads included in a process of an application.

5. The method according to any one of claims 1 to 4, characterized in that, The types of the first computing unit and the second computing unit include a matrix computing unit, a scalar computing unit, or a vector computing unit.

6. The method according to any one of claims 1 to 5, characterized in that The method further includes: Obtain a second task from multiple tasks to be executed, where the second task includes a first operator and a second operator; Divide the second task into multiple subtasks, the multiple subtasks include a first subtask and a second subtask, the first subtask includes the first operator, and the second subtask includes the second operator; Schedule the first subtask to the first queue, and the type of the first subtask matches the type of the first computing unit; Schedule the second subtask to the second queue, and the type of the second subtask matches the type of the second computing unit.

7. The method according to any one of claims 1 to 6, characterized in that, The method is executed by the first computing unit, and the first task is scheduled to the first queue; The method further includes: When the first computing unit cannot execute the first task in the first queue, the first computing unit determines the actual type to which the first task belongs; When the type to which the first task actually belongs matches the type of the second computing unit, the first computing unit schedules the first task to the second queue.

8. A task processing device, applied to a processor, characterized in that The first processor core in the processor includes a first computing unit and a second computing unit, and the first computing unit and the second computing unit are computing units of different types; The device includes: An obtaining module, configured to obtain a first task from a plurality of tasks to be executed; and obtain the type of the first task; A scheduling module, configured to schedule the first task to a first queue corresponding to the first computing unit when the type of the first task matches the type of the first computing unit; and when the type of the first task matches the type of the second computing unit, Schedule the first task to a second queue corresponding to the second computing unit.

9. The device according to claim 8, characterized in that, The plurality of tasks to be executed are tasks respectively executed by a plurality of threads included in a process of an application.

10. The device according to claim 8 or 9, characterized in that The types of the first computing unit and the second computing unit include a matrix computing unit, a scalar computing unit, or a vector computing unit.

11. The device according to any one of claims 8 to 10, wherein The obtaining module is further configured to obtain a second task from a plurality of tasks to be executed, and the second task includes a first operator and a second operator; The device further includes: A dividing module, configured to divide the second task into a plurality of subtasks, the plurality of subtasks including a first subtask and a second subtask, the first subtask including the first operator, and the second subtask including the second operator; The scheduling module is further configured to schedule the first subtask to the first queue, where the type of the first subtask matches the type of the first computing unit; and schedule the second subtask to the second queue, where the type of the second subtask matches the type of the second computing unit.

12. The device according to any one of claims 8 to 11, characterized in that, The device is the first computing unit, and the first task is scheduled to the first queue; The device further includes: A determining module, configured to determine the type to which the first task actually belongs when the first computing unit cannot execute the first task in the first queue; The scheduling module is further configured to schedule the first task to the second queue when the type to which the first task actually belongs matches the type of the second computing unit.

13. A computing device, characterized in that, Includes a processor and a memory; The processor is configured to execute instructions stored in the memory to enable the computing device to execute the method according to any one of claims 1 to 7.

14. A computer-readable storage medium, characterized in that, Includes instructions that, when running on a computing device, cause the computing device to execute the method according to any one of claims 1 to 7.

15. A computer program product comprising instructions, characterized in that, When running on at least one computing device, cause the at least one computing device to execute the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Configurable heterogeneous artificial intelligence processor

    CN112463709A

  • Task scheduling method and service system

    CN113900776A

  • Operation processing method, equipment and medium

    CN114020476A

  • Task scheduling method and device

    CN115269131A

  • Operator scheduling method and device, equipment, storage medium and program product

    CN116302439A

Cited By

  • Electronic equipment and data processing method

    CN122111693A

  • Electronic device and data processing method

    CN122111693B