Method and apparatus for scheduling operators

By allocating different streams and threads to asynchronous or synchronous operators according to the operator type, the problem of inefficient execution of existing executors on hardware devices is solved, and more efficient task scheduling and execution is achieved.

CN114936096BActive Publication Date: 2025-07-25BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210659177.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-09
Publication Date
2025-07-25
Estimated Expiration
2041-11-09

AI Technical Summary

Technical Problem

When existing executors schedule and execute operators in deep learning models, they cannot effectively utilize the parallel execution capabilities of hardware devices, resulting in low operator execution efficiency. Especially when different hardware devices and deep learning models have large differences in structures, synchronous operators may cause blocking problems.

Method used

By determining whether the operator information is asynchronous or synchronous operators, different streams and threads are allocated to them respectively to maximize the advantages of parallel execution of hardware devices and avoid blockage caused by synchronous operators.

Benefits of technology

It improves the execution efficiency of operators, improves the parallel execution capabilities of hardware devices, avoids the blocking problem of synchronous operators in a single thread, and achieves more efficient task scheduling.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114936096B_ABST
    Figure CN114936096B_ABST
Patent Text Reader

Abstract

The present disclosure provides a method and apparatus for scheduling operators, belonging to the field of computer technology, and particularly relating to the field of task scheduling. The specific implementation solution includes: determining operator information related to two or more operators included in a target task being executed; and scheduling the two or more operators according to the operator information, wherein the operator information indicates whether the operators included in the two or more operators are asynchronous operators or synchronous operators.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of a Chinese patent application with an application date of November 9, 2021 and an application number of 202111323084.9. Technical Field

[0002] The present disclosure relates to the field of computer technologies, and in particular to the field of task scheduling, and specifically to a method and apparatus for scheduling operators. Background Art

[0003] In various computer application scenarios, for example, in a deep learning framework, an executor is a core component for scheduling and executing each operator included in a deep learning model. A well-designed executor can support multiple hardware devices, cover a wide range of deep learning model usage scenarios, and efficiently complete the scheduling and execution of operators.

[0004] With the booming development of hardware chips and the continuous deepening of cutting-edge deep learning models, the executor of the deep learning framework has higher requirements for aspects such as the horizontal scalability of hardware and the operator execution efficiency. However, there are certain differences in the architecture design and execution mechanism (for example, the design of a stream) of different hardware devices. In addition, the deep learning model structures in different fields are also different. For example, models in the natural language processing (NLP) field are more "wide", while models in the computer vision (CV) field are more "deep". Therefore, the requirements for the scheduling strategy of the executor are also different. Summary of the Invention

[0005] The present disclosure provides a method, apparatus, electronic device, and storage medium for scheduling operators.

[0006] According to a first aspect of the present disclosure, there is provided a method for scheduling operators, including: determining operator information related to two or more operators included in a target task being executed; and scheduling the two or more operators according to the operator information, where the operator information indicates whether the operators included in the two or more operators are asynchronous operators or synchronous operators.

[0007] According to a second aspect of the present disclosure, there is provided an operator scheduling apparatus, including: a determination unit configured to determine operator information related to two or more operators included in a target task being executed; and a scheduling unit configured to schedule the two or more operators according to the operator information, where the operator information indicates whether the operators included in the two or more operators are asynchronous operators or synchronous operators.

[0008] According to a third aspect of the present disclosure, there is provided an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the above method.

[0009] According to a fourth aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions for causing a computer to execute the above method.

[0010] According to a fifth aspect of the present disclosure, there is provided a computer program product including a computer program which, when executed by a processor, implements the above method.

[0011] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:

[0013] Figure 1 FIG. shows an exemplary schematic diagram of the execution order of an operator sequence;

[0014] Figure 2 FIG. shows an exemplary schematic diagram of the execution process of an operator sequence;

[0015] Figure 3 FIG. shows a flowchart of a method for scheduling operators according to an embodiment of the present disclosure;

[0016] Figure 4 FIG. shows a flowchart of a method for allocating a stream to a kernel according to operator information according to an embodiment of the present disclosure;

[0017] Figure 5 FIG. shows a schematic diagram of a scenario for allocating multiple streams to a kernel corresponding to an operator according to an embodiment of the present disclosure;

[0018] Figure 6 FIG. shows a flowchart of a method for assigning a thread to an operator according to operator information according to an embodiment of the present disclosure;

[0019] Figure 7 FIG. shows a schematic diagram of a scenario for assigning a thread to an operator according to an embodiment of the present disclosure;

[0020] Figure 8 FIG. shows a flowchart of a method for assigning a thread to an operator according to another embodiment of the present disclosure;

[0021] Figure 9 A schematic diagram showing a scenario of assigning threads to operators according to another embodiment of the present disclosure;

[0022] Figure 10 A flowchart showing a method of scheduling operators according to another embodiment of the present disclosure;

[0023] Figure 11 A schematic diagram showing providing a unified interface for different hardware devices according to an embodiment of the present disclosure;

[0024] Figure 12 A block diagram showing an operator scheduling device according to an embodiment of the present disclosure; and

[0025] Figure 13 A schematic block diagram showing an example electronic device that can be used to implement the embodiments of the present disclosure. Detailed implementation manners

[0026] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to assist understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0027] The terms used herein are merely for describing specific embodiments and are not intended to limit the present disclosure. The terms "including", "comprising", etc. used herein indicate the presence of the described features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.

[0028] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those of ordinary skill in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.

[0029] Figure 1 A schematic example showing the execution order of an operator sequence.

[0030] As Figure 1As shown, the target task being executed by the actuator in the electronic device includes an operator sequence composed of multiple operators feed_op, OP2, OP3, OP4, OP5, OP6, and OP7. Here, an operator refers to an independent logical computing unit in the target task (e.g., a deep learning model) being executed by the actuator. Generally speaking, a target task can be regarded as a combination of multiple operators (e.g., the convolution operators Conv2D and Conv3D in the field of vision). Each operator can have fixed input variables (also known as tensors), computing kernels (kernels), and output variables. The kernel is the logical computing implementation of an operator on a hardware device. To support multiple hardware devices (e.g., NVIDIA's GPU (CUDA), AMD's ROCM hardware (HIP), Ascend chips (NPUs), Baidu Kunlun chips (XPUs), etc.), an operator can include multiple kernels, and each kernel represents a specific implementation on a certain hardware device. In terms of hardware, an operator first needs to be executed on the central processing unit CPU (also known as the host side), and then, according to the current hardware device type (e.g., a CUDA device), different kernels are called to implement it, so as to send a corresponding kernel API (application programming interface) to the stream. According to some embodiments, an operator can also correspond to multiple kernels, and each kernel corresponds to a kind of hardware device.

[0031] As Figure 1 shown, at the start of the target task, the operator feed_op receives the data to be processed. Then, the actuator sequentially executes the operators OP2, OP3, OP4, OP5, OP6, and OP7 according to the order of the topology graph as Figure 1 shown. There can also be variable copy operations (also known as copy kernels) between operators, such as copying from the CPU to the graphics processing unit GPU (CPU -> GPU), from the GPU to the CPU (GPU -> CPU), or from the GPU to the GPU (GPU -> GPU).

[0032] When the actuator executes the current operator in the operator sequence, it generally involves the following steps. First, the actuator receives the input variable, and the variable data can be the output variable of the previous operator or the copied variable. Then, the actuator applies for the memory or video memory of the output variable. Next, it executes the computing kernel for the specific hardware device corresponding to the current operator, such as a CPU kernel or a GPU kernel, etc. Finally, it provides the variable output by the computing kernel to the next operator.

[0033] Currently, most hardware platforms related to task execution (such as acceleration hardware platforms) have introduced the mechanism of streams, that is, they support parallel execution of multiple streams simultaneously to maximize the acceleration efficiency of hardware devices. Existing executors only use one or a few streams and lack analysis of streams at the operator level, and do not effectively allocate each operator to a suitable stream, so the best stream parallel execution efficiency cannot be achieved.

[0034] Figure 2 FIG. shows an example schematic diagram of an execution process of an operator sequence.

[0035] As Figure 2 shown, at the host side, an operator sequence composed of multiple operators feed_op, OP2, OP3, OP4, OP5, H2D (Host to Device), OP6, OP7 is assigned to a host-side thread for execution. Correspondingly, the computing kernels Kernel 2, Kernel 3, Kernel 5, Kernel 6 and the copy kernels CPU->GPU, GPU->CPU, GPU->GPU corresponding to the multiple operators are allocated in a Kernel (kernel) stream.

[0036] According to whether the subsequent operator needs to wait for the current operator to finish execution before starting to execute, operators can be divided into asynchronous operators and synchronous operators. An asynchronous operator means that after the CPU side (i.e., the host side) finishes executing the current operator code, it will send a computing kernel to the stream, and can continue to execute the next operator when the computing kernel has not started to execute. A synchronous operator means that when the CPU side executes the current operator, it must wait for the computing kernel corresponding to the current operator to finish execution before it can execute the next operator. As Figure 2 shown, the white box represents a synchronous operator, and the gray box represents an asynchronous operator. Therefore, for synchronous operators, synchronous blocking situations may occur, hindering the subsequent operators from being executed by a single thread.

[0037] Figure 3 FIG. shows a flowchart of a method 300 for scheduling operators according to an embodiment of the present disclosure.

[0038] As Figure 3 shown, in step S310, operator information related to two or more operators included in the target task being executed is determined. The operator information may indicate whether the operators included in the two or more operators are asynchronous operators or synchronous operators.

[0039] In step S320, two or more operators are scheduled according to the operator information. According to an embodiment of the present disclosure, scheduling two or more operators may include one or a combination of the following operations: allocating streams to two or more kernels corresponding to the two or more operators; and assigning threads to the two or more operators.

[0040] According to an embodiment of the present disclosure, compared with a scheme of allocating all kernels to a single stream (hereinafter simply referred to as the "single-stream scheme"), by analyzing the types of operators in the target task and allocating the kernels corresponding to the operators to multiple different streams, the parallel execution advantage of the hardware device can be maximally utilized, and the execution efficiency of the operators can be effectively improved. In addition, compared with a scheme of assigning all operators to a single thread (hereinafter simply referred to as the "single-thread scheme"), by analyzing the types of operators in the target task and assigning the operators to different threads, the blocking problem of some synchronous operators in the single thread can be effectively avoided.

[0041] Figure 4 The flowchart shows a method for allocating streams to kernels according to operator information according to an embodiment of the present disclosure.

[0042] As Figure 4 shown, in step S421, the first kernel among two or more kernels corresponding to two or more operators is allocated to the first stream. The first kernel may correspond to an asynchronous operator among the two or more operators and may include, for example, a computing kernel.

[0043] In step S422, the second kernel among the two or more kernels is allocated to the second stream. The second kernel may correspond to a synchronous operator among the two or more operators and may include, for example, a copy kernel. According to one embodiment, the second stream may be further divided according to the type of the second kernel. For example, the copy kernel from device A to device B may be allocated to the third stream, and the copy kernel from device B to device A may be allocated to the fourth stream.

[0044] According to an embodiment of the present disclosure, compared with the single-stream scheme, by analyzing the types of operators in the target task and allocating the kernels corresponding to the operators to multiple different streams, the parallel execution advantage of the hardware device can be maximally utilized, and the execution efficiency of the operators can be effectively improved.

[0045] Figure 5 The schematic diagram shows a scenario of allocating multiple streams to kernels corresponding to operators according to an embodiment of the present disclosure.

[0046] According to an embodiment of the present disclosure, since the types of operators on different hardware devices (i.e., whether the operator is a synchronous operator or an asynchronous operator) are different, the stream can be allocated to the kernel corresponding to the operator according to the operator information indicating whether the operator is a synchronous operator or an asynchronous operator. For example, the compute kernel on the GPU corresponds to an asynchronous operator. Therefore, the compute kernel on the GPU can be allocated to a separate stream for sequential execution. In addition, the cross-device copy operation of variables (also known as the cross-device copy kernel) corresponds to a synchronous operator, which causes a waiting operation in the kernel stream, resulting in a blocking situation. Therefore, the cross-device copy kernel can be allocated to another separate stream for sequential execution.

[0047] As Figure 5 shown, the streams executed on the hardware device side include the Kernel stream, the D2H (device-to-host copy) stream, and the H2D (host-to-device copy) stream. The Host-side thread includes multiple operators, where the white boxes feed_op, OP4, and H2D represent synchronous operators, and the gray boxes OP2, OP3, OP5, OP6, and OP7 represent asynchronous operators. The compute kernels Kernel2, Kernel 3, Kernel 5, and Kernel 6 corresponding to the asynchronous operators can be allocated to the first stream. For example, the compute kernels Kernel 2, Kernel 3, Kernel 5, and Kernel 6 can be allocated to the Kernel stream. There may also be a copy kernel within the same device in the compute kernel. For example, the GPU->GPU copy kernel. In addition, the cross-device copy kernels CPU->GPU and GPU->CPU corresponding to the synchronous operators can be allocated to the second stream. According to some embodiments, the second stream can be further divided according to the type of the cross-device copy kernel (e.g., whether it is CPU->GPU or GPU->CPU). For example, the cross-device copy kernel GPU->CPU can be allocated to the third stream (e.g., the D2H stream), and the cross-device copy kernel CPU->GPU can be allocated to the fourth stream (e.g., the H2D stream).

[0048] Figure 6 The flowchart of the method for assigning threads to operators according to operator information according to an embodiment of the present disclosure is shown.

[0049] As Figure 6 shown, in step S621, when the operator information indicates that at least one first operator in the operator sequence composed of two or more operators is an asynchronous operator, at least one first operator is assigned to the first thread. In one embodiment, the first thread may include a thread pool composed of first threads, such as an asynchronous thread pool.

[0050] In step S622, when the operator information indicates that at least one second operator in an operator sequence composed of two or more operators is a synchronous operator, assign at least one second operator to a second thread. In one embodiment, the second thread may include a thread pool composed of second threads, such as a synchronous thread pool.

[0051] According to an embodiment of the present disclosure, compared with a single-threaded solution, by analyzing the operator types in a target task and assigning operators to different threads, the blocking problem of some synchronous operators in a single thread can be effectively avoided.

[0052] Figure 7 A schematic diagram showing a scenario of assigning threads to operators according to an embodiment of the present disclosure is shown.

[0053] According to an embodiment of the present disclosure, not only can a stream be allocated to a kernel corresponding to an operator at the hardware device end according to operator information indicating whether the operator is a synchronous operator or an asynchronous operator, but also threads can be assigned to the operator at the host end according to the operator information.

[0054] When an executor executes a target task, the number, type, and execution hardware device of the operators included in the target task are known and fixed. Therefore, when the executor executes the target task, the following analysis operations can be performed: by analyzing on which hardware device each operator in the operator sequence should be executed, select a kernel corresponding to the hardware device; analyze to which stream each kernel should be allocated; after executing the current operator in the operator sequence, analyze which thread should be assigned to the subsequent operator of the current operator (that is, whether to assign the subsequent operator to the current thread for execution or to another thread for execution); after executing the current operator, analyze which variables' video memories should be recycled.

[0055] As Figure 7 shown, threads can be assigned to operators at the host end according to operator information indicating whether the operator is a synchronous operator or an asynchronous operator. For example, the threads executed at the host end may include two threads, thread 1 and thread 2. When the first operators OP2, OP3, OP5, OP6, OP7 in an operator sequence composed of two or more operators are asynchronous operators, the first operators OP2, OP3, OP5, OP6, OP7 can be assigned to thread 1. Additionally, when the second operators feed_op, OP4, H2D in the operator sequence are synchronous operators, the second operators feed_op, OP4, H2D can be assigned to thread 2.

[0056] Those skilled in the art should understand that the fact that the threads executed at the host end include two threads is only an example, and the threads executed at the host end may also include three or more threads.

[0057] According to some embodiments, the analysis operation as described above can be run only once when the executor executes, and the results of the analysis operation are reused directly for each subsequent iteration execution, thereby minimizing the repeated analysis work of the executor and improving the execution efficiency.

[0058] Figure 8 FIG. shows a flowchart of a method for assigning threads to operators according to another embodiment of the present disclosure.

[0059] In step S821, determine whether the operator information indicates that the current operator in the operator sequence composed of two or more operators is an asynchronous operator or a synchronous operator.

[0060] If it is determined in step S821 that the operator information indicates that the current operator in the operator sequence composed of two or more operators is an asynchronous operator, then in step S822, determine whether the subsequent operator of the current operator is an asynchronous operator or a synchronous operator.

[0061] In the case where the operator information indicates that the subsequent operator of the current operator is an asynchronous operator, in step S823, retain the subsequent operator of the current operator in the current thread.

[0062] In the case where the operator information indicates that the subsequent operator of the current operator is a synchronous operator, in step S824, assign the subsequent operator of the current operator to at least one third thread different from the current thread. In one embodiment, in the case where the number of the subsequent operators is greater than l, each of the subsequent operators can be assigned to each thread in at least one third thread in a one-to-one manner.

[0063] Return to step S821. If it is determined in step S821 that the operator information indicates that the current operator in the operator sequence composed of two or more operators is a synchronous operator, then in step S825, determine whether the subsequent operator of the current operator is an asynchronous operator or a synchronous operator.

[0064] If it is determined in step S825 that the subsequent operator of the current operator is a synchronous operator, then in step S826, determine the number of subsequent operators of the current operator in the operator sequence. If it is determined in step S826 that the number of subsequent operators is greater than 1, then in step S827, retain one of the subsequent operators of the current operator in the current thread, and assign the other operators among the subsequent operators to at least one fourth thread different from the current thread. In one embodiment, in the case where the number of the other operators is greater than 1, each of the other operators can be assigned to each thread in at least one fourth thread in a one-to-one manner. If it is determined in step S826 that the number of subsequent operators is equal to 1, then in step S828, retain the subsequent operator of the current operator in the current thread.

[0065] If it is determined in step S825 that the subsequent operator of the current operator is a synchronous operator, then in step S829, the subsequent operator of the current operator is assigned to a fifth thread different from the current thread and the fourth thread.

[0066] According to an embodiment of the present disclosure, while maximizing the advantages of multi-threaded scheduling, the additional overhead of excessive task addition operations can be reduced, the costs of thread dormancy and repeated wake-up can also be reduced, and the scheduling efficiency of the executor can be improved.

[0067] Figure 9 A schematic diagram showing a scenario of assigning threads to operators according to another embodiment of the present disclosure is shown.

[0068] In a multi-threaded scheduling scheme, if new operators are generated by the subsequent operator of the current operator in the operator sequence, the generated new operators are all added to other threads, resulting in additional overhead.

[0069] Due to possible resource limitations of graphics acceleration hardware such as GPUs, according to an embodiment of the present disclosure, the subsequent asynchronous operator of the current operator that is an asynchronous operator in the operator sequence can be left to be executed in the current thread, while the subsequent synchronous operator of the current operator is assigned to a new thread for execution. As Figure 5 shown, the gray boxes GPU 1, GPU 2, GPU 3, GPU 4, GPU 5, GPU 6, GPU 7 represent asynchronous operators executed on the GPU. The operators GPU 1 to GPU 7 indicated by the gray boxes can all be assigned to an asynchronous thread or an asynchronous thread pool.

[0070] For synchronous operators executed on the CPU, one of the subsequent synchronous operators of the current operator that is a synchronous operator in the operator sequence can be left to be executed in the current thread, while the other synchronous operators in the subsequent synchronous operators are all assigned to new threads for execution. It should be noted that the "subsequent operator" described herein refers to the operator that is executed immediately after the current operator.

[0071] As Figure 9As shown, when the subsequent operator of the asynchronous operator GPU 1 in the operator sequence only includes one asynchronous operator GPU2, the asynchronous operator GPU 2 can be left to execute in the current thread. When the subsequent operators of the asynchronous operator GPU 2 in the operator sequence include one asynchronous operator GPU 3 and one synchronous operator D2H 1, the asynchronous operator GPU 3 can be left to execute in the current thread, and the synchronous operator D2H 1 is assigned to the synchronous thread 2 through the AddTask operation of adding a task. When the subsequent operator of the asynchronous operator GPU 4 in the operator sequence only includes one synchronous operator D2H 2, the synchronous operator D2H 2 can be assigned to the synchronous thread 1 through the AddTask operation of adding a task. Moreover, the subsequent synchronous operator CPU 8 of the synchronous operator D2H 2 can be retained in the current synchronous thread 1.

[0072] For the synchronous operator CPU 2, since its subsequent operators include two synchronous operators, namely H2D 1 and CPU 3, one of these two synchronous operators, such as H2D 1, can be retained in the current synchronous thread 2, and the other synchronous operator CPU 3 is assigned to a new synchronous thread 3 through the AddTask operation of adding a task. For the synchronous operator H2D1, since its subsequent operator is the asynchronous operator GPU 5, the asynchronous operator GPU5 can be added to the asynchronous thread through the AddTask operation of adding a task. For the synchronous operator CPU 3, since its subsequent operators include two synchronous operators, namely CPU 4 and CPU 5, one of these two synchronous operators, such as CPU 4, can be retained in the current synchronous thread 3, and the other synchronous operator CPU 5 is assigned to a new synchronous thread 4 through the AddTask operation of adding a task. Moreover, the subsequent synchronous operators CPU 6 and CPU 7 downstream of the synchronous operator CPU 5 can be retained in the synchronous thread 4.

[0073] For the asynchronous operator GPU 5, its subsequent asynchronous operators GPU 6 and GPU 7 can be retained in the asynchronous thread. When the subsequent operator of the asynchronous operator GPU 7 only includes one synchronous operator D2H 3, the synchronous operator D2H 3 can be assigned to the synchronous thread 2 through the AddTask operation of adding a task. Moreover, the subsequent synchronous operator CPU9 of the synchronous operator D2H 3 can be retained in the synchronous thread 2.

[0074] Although Figure 9 shows that the subsequent operator of the current operator includes one or two operators, but as needed, the subsequent operator can also include three or more operators. For the case where the subsequent operator includes three or more operators, the scheduling method is the same as that described with reference to Figure 9 and will not be elaborated here for the sake of brevity.

[0075] Figure 10 FIG. 1000 is a flowchart of a method for scheduling operators according to another embodiment of the present disclosure.

[0076] Figure 10 Steps S1010 and S1020 shown are the same as Figure 3 steps S310 and S320 in, and for the sake of brevity, the repeated description thereof will be omitted.

[0077] In step S1030, a unified interface is provided between the upper layer of the framework and the hardware device layer by registering a plurality of hardware devices included in the hardware device layer of the electronic device in the registration manager of the electronic device. In one embodiment, the upper layer of the framework may include, for example, an executor, a trainer, a predictor, and an operator. The hardware devices may include, for example, HIP, CUDA, NPU, XPU, and CPU, where HIP and CUDA adopt the same event mechanism CUDAEvent, while NPU, XPU, and CPU respectively adopt their corresponding event mechanisms NPUEvent, XPUEvent, and CPUEvent.

[0078] According to an embodiment of the present disclosure, by providing a unified interface, the module interfaces related to different hardware devices in the application framework (including but not limited to stream interfaces, event interfaces, etc.) are unified, and a convenient new hardware registration management mechanism is provided, facilitating the horizontal expansion of emerging hardware devices.

[0079] Figure 11 FIG. shows a schematic diagram of providing a unified interface for different hardware devices according to an embodiment of the present disclosure.

[0080] Currently commonly used hardware devices include CPUs of Intel, GPUs (CUDA) of NVIDIA, ROCM hardware (HIP) of AMD, Ascend chips (NPU), Baidu Kunlun chips (XPU), etc. Different hardware devices have different interface forms and support degrees for stream and event mechanisms. For example, CPU devices do not support stream and event mechanisms.

[0081] In one embodiment, a unified interface may be provided between the upper layer of the framework and the hardware device layer of the electronic device. As Figure 11 shown in, an intermediate proxy layer may be provided between the upper layer of the framework and the underlying event layer associated with the underlying hardware device. The upper layer of the framework may include, for example, an executor, a trainer, a predictor, and an operator. The hardware devices may include, for example, HIP, CUDA, NPU, XPU, and CPU, where HIP and CUDA adopt the same event mechanism CUDAEvent, while NPU, XPU, and CPU respectively adopt their corresponding event mechanisms NPUEvent, XPUEvent, and CPUEvent.

[0082] In one embodiment, a unified interface can be provided between the upper layer of the framework of the electronic device and the hardware device layer by registering multiple hardware devices included in the hardware device layer in the registration manager of the electronic device.

[0083] Figure 12 The block diagram of the operator scheduling device 1200 according to an embodiment of the present disclosure is shown.

[0084] As Figure 12 shown, the operator scheduling device 1200 includes a determination unit 1210 and a scheduling unit 1220.

[0085] The determination unit 1210 is configured to determine operator information related to two or more operators included in the target task being executed. The operator information can indicate whether the operators included in the two or more operators are asynchronous operators or synchronous operators.

[0086] The scheduling unit 1220 schedules the two or more operators according to the operator information. According to an embodiment of the present disclosure, scheduling the two or more operators can include one or a combination of the following operations: allocating streams to two or more kernels corresponding to the two or more operators; and assigning threads to the two or more operators.

[0087] In one embodiment, the operator scheduling device 1200 may further include an interface providing unit, which is configured to provide a unified interface between the upper layer of the framework and the hardware device layer by registering multiple hardware devices included in the hardware device layer of the electronic device in the registration manager of the electronic device.

[0088] In the technical solution of the present disclosure, the acquisition, storage, and application of the user's personal information involved all comply with the provisions of relevant laws and regulations, take necessary confidentiality measures, and do not violate public order and good customs.

[0089] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0090] Figure 13FIG. shows a schematic block diagram of an example electronic device 1300 that may be used to implement embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, for example, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as, for example, personal digital processors, cellular telephones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementations of the present disclosure described and / or claimed herein.

[0091] As Figure 13 shown, the device 1300 includes a computing unit 1301 that can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 1302 or a computer program loaded from a storage unit 1308 into a random access memory (RAM) 1303. In the RAM 1303, various programs and data required for the operation of the device 1300 can also be stored. The computing unit 1301, the ROM 1302, and the RAM 1303 are connected to each other via a bus 1304. An input / output (I / O) interface 1305 is also connected to the bus 1304.

[0092] Multiple components in the device 1300 are connected to the I / O interface 1305, including: an input unit 1306, such as a keyboard, a mouse, etc.; an output unit 1307, such as various types of displays, speakers, etc.; a storage unit 1308, such as a magnetic disk, an optical disk, etc.; and a communication unit 1309, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 1309 allows the device 1300 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0093] The computing unit 1301 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1301 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1301 executes the various methods and processes described above, such as the operator scheduling method. For example, in some embodiments, the operator scheduling method can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 1308. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 1300 via the ROM 1302 and / or the communication unit 1309. When the computer program is loaded into the RAM 1303 and executed by the computing unit 1301, one or more steps of the operator scheduling method described above can be executed. Alternatively, in other embodiments, the computing unit 1301 can be configured to execute the operator scheduling method in any other suitable way (e.g., by means of firmware).

[0094] The various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip systems (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including two or more programmable processors, which can be special-purpose or general-purpose programmable processors, receiving data and instructions from a storage system, two or more input devices, and two or more output devices, and transmitting the data and instructions to the storage system, the two or more input devices, and the two or more output devices.

[0095] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to the processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program code can be executed entirely on the machine, partially on the machine, as an independent software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0096] In the context of this disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0097] In order to provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, speech input, or tactile input).

[0098] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.

[0099] A computer system can include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The relationship between the client and the server is generated by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, can also be a server of a distributed system, or a server incorporating a blockchain.

[0100] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in the present disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired results of the technical solution disclosed in the present disclosure can be achieved, and no limitations are imposed herein.

[0101] The above specific embodiments do not constitute a limitation on the protection scope of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub - combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present disclosure shall be included within the protection scope of the present disclosure.

Claims

1. A method for scheduling operators, comprising: Determining operator information related to two or more operators included in a target task being executed; And Scheduling the two or more operators according to the operator information, including: Assigning threads to the two or more operators according to the operator information; Wherein the operator information indicates whether the operators included in the two or more operators are asynchronous operators or synchronous operators; The assigning threads to the two or more operators includes: When the operator information indicates that the current operator in the operator sequence composed of the two or more operators is a synchronous operator, determining whether the subsequent operator of the current operator is an asynchronous operator or a synchronous operator; When it is determined that the subsequent operator of the current operator is a synchronous operator, then: Determining the number of subsequent operators of the current operator in the operator sequence; When the number of subsequent operators of the current operator is equal to 1, retaining the subsequent operator of the current operator in the current thread; and When the number of subsequent operators of the current operator is greater than 1, retaining one of the subsequent operators of the current operator in the current thread, and assigning the other operators among the subsequent operators to at least one fourth thread different from the current thread; and When it is determined that the subsequent operator of the current operator is an asynchronous operator, assigning the subsequent operator of the current operator to a fifth thread different from the current thread and the fourth thread.

2. The method according to claim 1, wherein Scheduling the two or more operators according to the operator information includes: Allocating streams to two or more cores corresponding to the two or more operators according to the operator information.

3. The method according to claim 2, wherein Allocating streams to two or more cores corresponding to the two or more operators according to the operator information includes: Allocating a first core among the two or more cores to a first stream; and Allocating a second core among the two or more cores to a second stream, Where the first core corresponds to an asynchronous operator among the two or more operators, and the second core corresponds to a synchronous operator among the two or more operators.

4. The method according to claim 1, wherein Assigning threads to the two or more operators includes: When the operator information indicates that at least one first operator in the operator sequence composed of the two or more operators is an asynchronous operator, assigning the at least one first operator to a first thread; When the operator information indicates that at least one second operator in the operator sequence composed of the two or more operators is a synchronous operator, assigning the at least one second operator to a second thread.

5. The method according to claim 1, wherein Assigning threads to the two or more operators includes: When the operator information indicates that the current operator in the operator sequence composed of the two or more operators is an asynchronous operator, determining whether the subsequent operator of the current operator is an asynchronous operator or a synchronous operator; When the operator information indicates that the subsequent operator of the current operator is an asynchronous operator, retaining the subsequent operator of the current operator in the current thread; and When the operator information indicates that the subsequent operator of the current operator is a synchronous operator, assign the subsequent operator of the current operator to at least one third thread different from the current thread; Among them, assigning the subsequent operator of the current operator to at least one third thread different from the current thread includes: In the operator sequence, when the number of the subsequent operators is greater than 1, assign each of the subsequent operators to each thread in the at least one third thread in a one-to-one manner.

6. The method according to claim 1, wherein Assigning other operators among the subsequent operators of the current operator to at least one fourth thread different from the current thread includes: In the case where the number of the other operators is greater than 1, assign each of the other operators to each thread in the at least one fourth thread in a one-to-one manner.

7. The method according to claim 1, wherein The method is applied to an electronic device, and the electronic device includes an upper layer of the framework and a hardware device layer. The method further includes: By the following operations, provide a unified interface between the upper layer of the framework and the hardware device layer: register multiple hardware devices included in the hardware device layer in the registration manager of the electronic device.

8. An operator scheduling device, including: A determination unit, configured to determine operator information related to two or more operators included in a target task being executed; And A scheduling unit, configured to schedule the two or more operators according to the operator information, including: assigning threads to the two or more operators according to the operator information; Among them, the operator information indicates whether the operators included in the two or more operators are asynchronous operators or synchronous operators; The scheduling unit is specifically configured to: When the operator information indicates that the current operator in the operator sequence composed of the two or more operators is a synchronous operator, determine whether the subsequent operator of the current operator is an asynchronous operator or a synchronous operator; When it is determined that the subsequent operator of the current operator is a synchronous operator, then: Determine the number of subsequent operators of the current operator in the operator sequence; When the number of subsequent operators of the current operator is equal to 1, retain the subsequent operator of the current operator in the current thread; and When the number of subsequent operators of the current operator is greater than 1, retain one of the subsequent operators of the current operator in the current thread, and assign the other operators in the subsequent operators to at least one fourth thread different from the current thread; and When it is determined that the subsequent operator of the current operator is an asynchronous operator, assign the subsequent operator of the current operator to a fifth thread different from the current thread and the fourth thread.

9. An electronic device, including: Two or more processors; And A memory communicatively connected to the two or more processors; wherein, The memory stores instructions executable by the two or more processors, and the instructions are executed by the two or more processors so that the two or more processors can execute the method according to any one of claims 1-7.

10. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the method according to any one of claims 1-7.

Citation Information

Patent Citations

  • Data processing method and device, electronic equipment and computer readable storage medium

    CN111274038A

  • Task handling method and device and medium

    CN112148455A