Heterogeneous Scheduling of Sequential Computation DAGs
By splitting the DAG computing model into multiple non-interdependent child nodes, processing these child nodes in parallel, and scheduling with processing units such as multi-core CPU and GPU, the problem of resource waste in parallel processing within the node is solved, and more efficient computing is achieved.
Patent Information
- Application Number
- CN201980059038.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2018-09-11
- Filing Date
- 2019-04-28
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2039-04-28
AI Technical Summary
In the prior art, the intra-node parallel processing method of the DAG computing model leads to wasting of processing system resources, and the parallelism between different processing units cannot be effectively utilized, resulting in inefficient computing efficiency.
By splitting the DAG computing model into multiple non-interdependent child nodes, processing these child nodes in parallel, and using processing units such as multi-core CPU and GPU for scheduling, a new subDAG computing model is built to realize parallel processing between nodes.
It improves the processing efficiency of the DAG computing model, reduces processing time, makes full use of the hardware resources of multiple processing units, and achieves faster and more efficient computing.
Smart Images

Figure CN112673352B_ABST
Abstract
Description
[0001] Cross - Reference to Related Applications
[0002] This application claims priority to U.S. Provisional Patent Application No. 62 / 729,646, filed on Sep. 11, 2018, with the title "HETEROGENEOUS SCHEDULING FOR SEQUENTIAL COMPUTE DAG", the entire content of which is incorporated herein by reference in its entirety. Technical Field
[0003] The present invention generally relates to inter - node processing, and in particular embodiments, to a system and method for constructing a directed acyclic graph (DAG) computational model for inter - node parallel processing between different types of processing units. Background Art
[0004] Typically, intra - node parallelism is used to process directed acyclic graph (DAG) modeling - type computations, which are sequential. In intra - node parallel processing, multiple cores of a central processing unit (CPU), a graphics processing unit (GPU), or any other type of dedicated processor apply sequential operations to process each node of the DAG. Each computational task in the DAG modeling - type computation is associated with or mapped to a separate DAG node. In some computations, the computational tasks can be subdivided into smaller subtasks. In intra - node parallel processing, the scheduling granularity is limited within a single node, and different processing units are not used between multiple DAG nodes or within each DAG node to achieve inter - node parallel processing. Summary of the Invention
[0005] Technical advantages are generally achieved through embodiments of the present invention, which describe the construction of a directed acyclic graph (DAG) computational model for inter - node parallel processing between different types of processing units.
[0006] According to one embodiment, a method for processing a directed acyclic graph (DAG) calculation is provided. The method includes: a plurality of processors splitting the DAG calculation into a plurality of non-dependent sub-nodes in each corresponding node. The plurality of processors includes a multi-core graphics processing unit (GPU) and a multi-core central processing unit (CPU). The method further includes the plurality of processors constructing a plurality of sub-DAG calculations. Each sub-DAG calculation includes at least non-dependent sub-nodes in different nodes of the DAG calculation. The method further includes the plurality of processors processing each of the plurality of sub-DAG calculations in parallel. In one embodiment, the method further includes: the plurality of processors allocating an intermediate shared memory for the plurality of sub-DAG calculations. Optionally, in this example, or in another example, the method further includes: scheduling the processing of each of the plurality of sub-DAG calculations by the CPU or the GPU. Optionally, in any of the above examples, or in another example, the scheduling further includes: scheduling the processing of each sub-node through the cores of the GPU and the cores of the CPU according to the task type of the corresponding sub-node of the DAG calculation. Optionally, in any of the above examples, or in another example, the DAG calculation includes an image processing application, a video processing application, or a deep neural network processing application. Optionally, in any of the above examples, or in another example, if the processing of the next sub-node in the corresponding sub-DAG calculation is started, the processing of the sub-node in the corresponding sub-DAG calculation is considered completed. Optionally, in any of the above examples, or in another example, the processing of non-dependent sub-nodes in a corresponding node is independent of the processing of another non-dependent sub-node in the same corresponding node. Optionally, in any of the above examples, or in another example, the construction is manually or automatically completed by a compiler executed by the plurality of processors. Optionally, in any of the above examples, or in another example, the method further includes receiving an input for the DAG calculation. Optionally, in any of the above examples, or in another example, the method further includes: outputting output data buffers, output image buffers, output image files, or output features of the DAG calculation. Optionally, in any of the above examples, or in another example, the splitting is performed according to the task type associated with each node and the storage capabilities of the plurality of processors. Optionally, in any of the above examples, or in another example, the splitting includes splitting into uniform non-dependent sub-nodes. Optionally, in any of the above examples, or in another example, the splitting includes splitting into non-uniform non-dependent sub-nodes.Optionally, in any of the above examples, or in another example, the splitting includes covering the boundaries of non - interdependent child nodes. Optionally, in any of the above examples, or in another example, each child node is a subtask related to the corresponding node of the DAG computation. Optionally, in any of the above examples, or in another example, one or more nodes of the DAG computation are split hierarchically. Optionally, in any of the above examples, or in another example, one or more child nodes of one or more sub - DAG computations are determined based on the outputs of multiple child nodes. Optionally, in any of the above examples, or in another example, one or more child nodes of one or more sub - DAG computations are the inputs of multiple child nodes.
[0007] According to another embodiment, a computer-implemented method for processing a directed acyclic graph (DAG) calculation is provided. The method includes: multiple processors splitting the DAG calculation into multiple non-dependent sub-nodes in each corresponding node. The multiple processors include a multi-core graphics processing unit (GPU) and a multi-core central processing unit (CPU). The method further includes the multiple processors constructing multiple sub-DAG calculations. Each sub-DAG calculation includes at least non-dependent sub-nodes in different nodes of the DAG calculation. The method further includes the multiple processors processing each sub-DAG calculation in the multiple sub-DAG calculations in parallel. In one embodiment, the method further includes: the multiple processors allocating an intermediate shared memory for the multiple sub-DAG calculations. Optionally, in this example, or in another example, the method further includes: scheduling the processing of each sub-DAG calculation in the multiple sub-DAG calculations by the CPU or the GPU. Optionally, in any of the above examples, or in another example, the scheduling further includes: scheduling the processing of each sub-node through the cores of the GPU and the cores of the CPU according to the task type of the corresponding sub-node of the DAG calculation. Optionally, in any of the above examples, or in another example, the DAG calculation includes an image processing application, a video processing application, or a deep neural network processing application. Optionally, in any of the above examples, or in another example, if the processing of the next sub-node in the corresponding sub-DAG calculation is started, the processing of the sub-node in the corresponding sub-DAG calculation is considered completed. Optionally, in any of the above examples, or in another example, the processing of non-dependent sub-nodes in a corresponding node is independent of the processing of another non-dependent sub-node in the same corresponding node. Optionally, in any of the above examples, or in another example, the construction is manually or automatically completed by a compiler executed by the multiple processors. Optionally, in any of the above examples, or in another example, the method further includes receiving an input for the DAG calculation. Optionally, in any of the above examples, or in another example, the method further includes: outputting the output data buffer, output image buffer, output image file, or output features of the DAG calculation. Optionally, in any of the above examples, or in another example, the splitting is performed according to the task type associated with each node and the storage capabilities of the multiple processors. Optionally, in any of the above examples, or in another example, the splitting includes splitting into uniform non-dependent sub-nodes. Optionally, in any of the above examples, or in another example, the splitting includes splitting into non-uniform non-dependent sub-nodes.Optionally, in any of the above examples, or in another example, the splitting includes covering the boundaries of non - interdependent child nodes. Optionally, in any of the above examples, or in another example, each child node is a subtask related to the corresponding node of the DAG computation. Optionally, in any of the above examples, or in another example, one or more nodes of the DAG computation are split hierarchically. Optionally, in any of the above examples, or in another example, one or more child nodes of one or more sub - DAG computations are determined according to multiple child node outputs. Optionally, in any of the above examples, or in another example, one or more child nodes of one or more sub - DAG computations are inputs of multiple child nodes.
[0008] According to another embodiment, a non-transitory computer-readable medium is provided, and the non-transitory computer-readable medium stores computer instructions for processing a directed acyclic graph (DAG) calculation. When a plurality of processors including a first processing unit and a second processing unit execute the instructions, the following steps are performed: The plurality of processors split the DAG calculation into a plurality of non-dependent sub-nodes in each corresponding node, where the plurality of processors includes a multi-core graphics processing unit (GPU) and a multi-core central processing unit (CPU). When the instructions are executed, the following steps are performed: The plurality of processors construct a plurality of sub-DAG calculations. Each sub-DAG calculation includes at least non-dependent sub-nodes in different nodes of the DAG calculation. When the instructions are executed, the following steps are performed: The plurality of processors process each sub-DAG calculation among the plurality of sub-DAG calculations in parallel. In one embodiment, when the plurality of processors execute the instructions, the following steps are performed: The plurality of processors allocate an intermediate shared memory for the plurality of sub-DAG calculations. Optionally, in this example, or in another example, when the instructions are executed, the following steps are performed: Processing of each sub-DAG calculation among the plurality of sub-DAG calculations is scheduled by the CPU or the GPU. Optionally, in any of the above examples, or in another example, the scheduling further includes: Scheduling the processing of each sub-node through the cores of the GPU and the cores of the CPU according to the task type of the corresponding sub-node of the DAG calculation. Optionally, in any of the above examples, or in another example, when the instructions are executed, the following steps are performed: The DAG calculation includes an image processing application, a video processing application, or a deep neural network processing application. Optionally, in any of the above examples, or in another example, if the processing of the next sub-node in the corresponding sub-DAG calculation starts, the processing of the sub-node in the corresponding sub-DAG calculation is regarded as completed. Optionally, in any of the above examples, or in another example, the processing of non-dependent sub-nodes in a corresponding node is independent of the processing of another non-dependent sub-node in the same corresponding node. Optionally, in any of the above examples, or in another example, the construction is manually or automatically completed by a compiler executed by the plurality of processors. Optionally, in any of the above examples, or in another example, when the instructions are executed, the following steps are performed: Receiving an input for the DAG calculation. Optionally, in any of the above examples, or in another example, when the instructions are executed, the following steps are performed: Outputting the output data buffer, output image buffer, output image file, or output features of the DAG calculation.Optionally, in any of the above examples, or in another example, the splitting is performed according to the task type associated with each node and the storage capabilities of the plurality of processors. Optionally, in any of the above examples, or in another example, the splitting includes splitting into evenly non-dependent sub-nodes. Optionally, in any of the above examples, or in another example, the splitting includes splitting into unevenly non-dependent sub-nodes. Optionally, in any of the above examples, or in another example, the splitting includes covering the boundaries of non-dependent sub-nodes. Optionally, in any of the above examples, or in another example, each sub-node is a sub-task associated with the corresponding node of the DAG computation. Optionally, in any of the above examples, or in another example, one or more nodes of the DAG computation are split hierarchically. Optionally, in any of the above examples, or in another example, one or more sub-nodes of one or more sub-DAG computations are determined based on the outputs of multiple sub-nodes. Optionally, in any of the above examples, or in another example, one or more sub-nodes of one or more sub-DAG computations are the inputs of multiple sub-nodes.
[0009] According to one embodiment, a device for processing directed acyclic graph (DAG) calculations is provided. The device includes: a non-transitory memory including instructions; and a plurality of processors including a central processing unit (CPU) and a graphics processing unit (GPU). The plurality of processors communicate with the non-transitory memory and execute the instructions to split the DAG calculation into a plurality of non-dependent sub-nodes in each corresponding node. The plurality of processors execute the instructions to construct a plurality of sub-DAG calculations. Each sub-DAG calculation includes at least non-dependent sub-nodes in different nodes of the DAG calculation. The plurality of processors execute the instructions to process each of the plurality of sub-DAG calculations in parallel. In one example, the plurality of processors execute the instructions to allocate intermediate shared memory for the plurality of sub-DAG calculations. Optionally, in this example, or in another example, the plurality of processors execute the instructions to schedule the processing of each of the plurality of sub-DAG calculations through the CPU or the GPU. Optionally, in any of the above examples, or in another example, the scheduling further includes: scheduling the processing of each sub-node through the cores of the GPU and the cores of the CPU according to the task type of the corresponding sub-node of the DAG calculation. Optionally, in any of the above examples, or in another example, the DAG calculation includes an image processing application, a video processing application, or a deep neural network processing application. Optionally, in any of the above examples, or in another example, if the processing of the next sub-node in the corresponding sub-DAG calculation starts, the processing of the sub-node in the corresponding sub-DAG calculation is considered complete. Optionally, in any of the above examples, or in another example, the processing of non-dependent sub-nodes in a corresponding node is independent of the processing of another non-dependent sub-node in the same corresponding node. Optionally, in any of the above examples, or in another example, the construction is manually or automatically completed by a compiler executed by the plurality of processors. Optionally, in any of the above examples, or in another example, the plurality of processors execute the instructions to receive an input for the DAG calculation. Optionally, in any of the above examples, or in another example, the plurality of processors execute the instructions to output the output data buffer, output image buffer, output image file, or output features of the DAG calculation. Optionally, in any of the above examples, or in another example, the splitting is performed according to the task type associated with each node and the storage capabilities of the plurality of processors. Optionally, in any of the above examples, or in another example, the splitting includes splitting into uniform non-dependent sub-nodes.Optionally, in any of the above examples, or in another example, the splitting includes splitting into uneven and non - interdependent child nodes. Optionally, in any of the above examples, or in another example, the splitting includes covering the boundaries of non - interdependent child nodes. Optionally, in any of the above examples, or in another example, each child node is a subtask related to the corresponding node of the DAG calculation. Optionally, in any of the above examples, or in another example, one or more nodes of the DAG calculation are split hierarchically. Optionally, in any of the above examples, or in another example, one or more child nodes of one or more sub - DAG calculations are determined according to multiple child node outputs. Optionally, in any of the above examples, or in another example, one or more child nodes of one or more sub - DAG calculations are inputs of multiple child nodes.
[0010] According to another embodiment, a device for processing a directed acyclic graph (DAG) computation is provided. The device includes a non-transitory memory that includes instructions and a plurality of processors. The plurality of processors includes a first processing unit and a second processing unit, and the processor types of the first processing unit and the second processing unit are different. The plurality of processors communicate with the non-transitory memory, and the plurality of processors execute the instructions to split the DAG computation into a plurality of non-dependent sub-nodes in each corresponding node. The plurality of processors execute the instructions to construct a plurality of sub-DAG computations, where each sub-DAG computation includes at least non-dependent sub-nodes in different nodes of the DAG computation; and process each sub-DAG computation among the plurality of sub-DAG computations in parallel. In one example, the plurality of processors execute the instructions to allocate an intermediate shared memory for the plurality of sub-DAG computations. Optionally, in this example, or in another example, the plurality of processors execute the instructions to schedule the processing of each sub-DAG computation among the plurality of sub-DAG computations through the CPU or the GPU. Optionally, in any of the above examples, or in another example, the scheduling further includes: scheduling the processing of each sub-node through the cores of the GPU and the cores of the CPU according to the task type of the corresponding sub-node of the DAG computation. Optionally, in any of the above examples, or in another example, the DAG computation includes an image processing application, a video processing application, or a deep neural network processing application. Optionally, in any of the above examples, or in another example, if the processing of the next sub-node in the corresponding sub-DAG computation starts, the processing of the sub-node in the corresponding sub-DAG computation is regarded as completed. Optionally, in any of the above examples, or in another example, the processing of non-dependent sub-nodes in a corresponding node is independent of the processing of another non-dependent sub-node in the same corresponding node. Optionally, in any of the above examples, or in another example, the construction is manually or automatically completed by a compiler executed by the plurality of processors. Optionally, in any of the above examples, or in another example, the plurality of processors execute the instructions to receive an input for the DAG computation. Optionally, in any of the above examples, or in another example, the plurality of processors execute the instructions to output the output data buffer, output image buffer, output image file, or output feature of the DAG computation. Optionally, in any of the above examples, or in another example, the splitting is performed according to the task type associated with each node and the storage capabilities of the plurality of processors. Optionally, in any of the above examples, or in another example, the splitting includes splitting into uniform non-dependent sub-nodes. Optionally, in any of the above examples, or in another example, the splitting includes splitting into non-uniform non-dependent sub-nodes.Optionally, in any of the above examples, or in another example, the splitting includes covering the boundaries of non - interdependent child nodes. Optionally, in any of the above examples, or in another example, each child node is a subtask related to the corresponding node of the DAG calculation. Optionally, in any of the above examples, or in another example, the processor types of the first processing unit and the second processing unit are different and are selected from the following group of processors: central processing unit (CPU), graphics processing unit (GPU), digital signal processor (DSP), image signal processor (ISP), video processing unit (VPU), neural network processing unit (NPU), and display processing unit (DPU). Optionally, in any of the above examples, or in another example, the device further includes at least one of the following: an interconnect bus link, a shared storage unit, a storage controller, one or more storage units, a peripheral interconnect, or a combination thereof. Optionally, in any of the above examples, or in another example, one or more nodes of the DAG calculation are hierarchically split. Optionally, in any of the above examples, or in another example, one or more child nodes of one or more sub - DAG calculations are determined based on the outputs of multiple child nodes. Optionally, in any of the above examples, or in another example, one or more child nodes of one or more sub - DAG calculations are the inputs of multiple child nodes.
[0011] Brief Description of the Drawings
[0012] To more fully understand the present invention and its advantages, reference is now made to the following description taken in conjunction with the accompanying drawings, in which:
[0013] Figure 1 Schematic diagram of a processing system embodiment;
[0014] Figure 2A Exemplary directed acyclic graph (DAG) calculation model with three nodes;
[0015] Figure 2B Example of hierarchical splitting of DAG nodes;
[0016] Figure 3 Example of a DAG calculation model with two nodes and each node having multiple child nodes;
[0017] Figure 4 Flow chart of an exemplary method for splitting a DAG computation model and constructing multiple sub - DAG computation models for parallel processing between nodes;
[0018] Figure 5A Example of a DAG computation model with three nodes used in an image processing application;
[0019] Figure 5B Example of a DAG computation model with three sub - nodes used in an image processing example with multiple sub - nodes;
[0020] Figure 5C Example of constructing multiple new sub - DAG computation models based on an original DAG computation model optimized for parallel processing between nodes;
[0021] Figure 5D Exemplary data flow for memory allocation for constructing a new sub - DAG computation model;
[0022] Figure 6A Example of a DAG computation model with three nodes used in a deep neural network processing application;
[0023] Figure 6B Example of splitting the input matrix of each node into multiple input matrices;
[0024] Figure 6C Example of constructing multiple new sub - DAG computation models based on an original deep neural network DAG computation model optimized for parallel processing between nodes;
[0025] Figure 7 Example of a DAG computation model with multiple nodes used in a computer vision processing application;
[0026] Figure 8 Example of a DAG computation model with a one - to - many mapping graph model;
[0027] Figure 9 Schematic diagram of an embodiment of a wireless communication network;
[0028] Figure 10 Another schematic diagram of an embodiment of a processing system;
[0029] Figure 11 Schematic diagram of an embodiment of a transceiver. Detailed implementation manners
[0030] Many applicable inventive concepts provided by the present invention can be embodied in a variety of specific environments. Specific embodiments merely illustrate specific configurations and do not limit the scope of the claimed embodiments. Unless otherwise stated, the features of different embodiments can be combined to form other embodiments. Variations or modifications described with respect to one embodiment can also be applied to other embodiments. Further, it should be understood that various changes, substitutions, and alterations can be made herein without departing from the spirit and scope of the present disclosure as defined by the appended claims. Although aspects of the present invention are mainly described in the context of a graphics processing unit (GPU) and a central processing unit (CPU), it should also be understood that these inventive aspects can also be applied to other processing units to provide inter-node parallel processing in the calculation of a directed acyclic graph (DAG) model.
[0031] The processing of DAG calculation is generally implemented in a way that uses intra-node parallelism, that is, multiple cores of the same processing unit process each node of the DAG in an order-dependent manner. Specifically, each calculation task in the DAG calculation model is associated with or mapped to a separate node, and in some calculations, the calculation task can be subdivided into smaller subtasks. In intra-node parallel processing, the scheduling granularity is limited within a single node, and inter-node parallel processing is not implemented using different processing units between multiple nodes or within each node. For example, the subtasks associated with the first node can be processed in parallel by multiple cores of the CPU, and the subtasks associated with the second node can be processed in parallel by multiple cores of the GPU. However, the GPU processing of the second node does not start until the CPU processing scheduling of the first node is completed. Therefore, each node in intra-node parallel processing is independently and sequentially calculated by a specific processing unit. This wastes resources in the processing system.
[0032] Embodiments of the present invention provide an updated DAG calculation model for constructing and scheduling inter-node parallel processing based on the original DAG calculation model for intra-node parallel processing. Some embodiments of the present invention facilitate the implementation of parallelism between multiple DAG nodes using different processing units. In other embodiments, the parallelism using different processing units facilitates the implementation in the subtasks of different nodes of the original DAG calculation model. Specifically, each subtask previously associated with a single node can be represented as a new node in the modified DAG calculation model. The conversion of the original node into multiple new sub-nodes and the arrangement of a new DAG calculation model according to the multiple new sub-nodes can then contribute to the simultaneous use of multiple hardware resources (i.e., CPU, GPU, etc.) for the calculation of the new DAG calculation model. Therefore, the processing system can process the DAG calculation model at a faster and more efficient speed than before, for example, when executed using intra-node parallel processing. These and other details will be described in more detail below.
[0033] Figure 1 FIG. 1 shows a block diagram of an exemplary processing system 100 for performing the methods described herein, which system may be installed in a host device. As shown, processing system 100 includes central processing units (CPUs) 102 and 106, a graphics processing unit (GPU) 110, a digital signal processor (DSP) 114, an image signal processor (ISP) 118, a video processing unit (VPU) 122, a neural network processing unit (NPU) 126, a display processing unit (DPU) 130, an interconnect bus link 134, a shared storage unit 136, a memory controller 138, memory units 140 and 142, and a peripheral interconnect 144, which components may (or may not) be arranged as Figure 1 shown. Processing system 100 may include Figure 1 other components not shown in FIG. 1, such as long-term memory (e.g., non-volatile memory). In some embodiments, processing system 100 may include a subset of the various processing units. For simplicity of description, Figure 1 the number of each component is shown in FIG. 1. In various embodiments, other numbers of the same component type may be contemplated.
[0034] In some embodiments, each component of processing system 100 may be located on a single chip or circuit, e.g., in an integrated circuit (IC) of a system on a chip (SoC) type. In other embodiments, each component of processing system 100 may be located on a different chip or circuit. In one embodiment, some components of processing system 100 may be located on the same chip or circuit, while some components may be located on different chips or circuits.
[0035] CPUs 102 and 106 can be used to perform basic arithmetic, logic, input / output (I / O), and control operations of the instruction set in the processing system 100. The GPU 110 can be used to perform efficient computer graphics calculations and image processing operations of the instruction set in the processing system 100. The DSP 114 can be used to efficiently measure, filter, or compress analog signals or process digital signal processing algorithms in the processing system 100. The ISP 118 is a special type of DSP 114 that can be used to efficiently process images in the processing system 100. The VPU 122 is also a special type of DSP 114 that can be used to efficiently process video in the processing system 100. The NPU 126 can be used to process data and solve problems using neural networks in the processing system 100. The DPU 130 can be used to process data related to the display of the processing system 100. Figure 1 Other types of processing units that can be implemented using the embodiments of the present invention, not shown in the figure, are application processing units (APUs), field programmable gate arrays (FPGAs), microcontrollers, etc. Each processing unit of the processing system 100 can be architecturally optimized and designed to perform a non-limiting set of specific tasks in an efficient or accelerated manner. Figure 1 The list of processing units shown is a non-limiting example of task-specific processors, each with multiple cores. For example, the GPU 110 can be architecturally optimized to perform the same operation on a large volume of data faster and more efficiently than CPUs 102 and 106. Each of the individual processing units can independently include hardware caches 104, 108, 112, 116, 120, 124, 128, and 132, which are organized in a hierarchy of more cache levels (L1, L2, etc.). Each processing unit can also include several or hundreds of cores, which can simultaneously process thousands of threads.
[0036] The interconnect bus link 134 is a communication link or cache coherent interconnect for sending data between the individual processing units, the shared memory 136, and the peripheral interconnect 144. The interconnect bus link 134 can be a software or hardware type control bus, address bus, or data bus operating across multiple communication protocols. The interconnect bus link 134 can have various topologies, such as multi-point, daisy chain, switch, etc.
[0037] The shared memory 136 can be any component or collection of components for storing programs and / or instructions and associated input / output data and / or intermediate data for execution by any of the processing units. Each processing unit can access the shared memory 136 via the interconnect bus link 134. The shared memory 136 can be a non-transitory computer-readable medium. Non-transitory computer-readable media include various types of computer-readable media, including magnetic storage media, optical storage media, flash media, and solid-state storage media. It should be understood that software can be sold together with the processing system 100. Alternatively, the software can be obtained and loaded into the processing system 100, including obtaining the software via a physical medium or a distribution system, including, for example, obtaining the software from a server owned by the software creator or from a server not owned by the software creator but used by it. For example, the software can be stored on a server for distribution via the Internet.
[0038] The memory controller 138 is used to manage the data flow into and out of the shared memory 136. In some embodiments, the memory controller 138 can be an integrated memory controller (IMC). In some embodiments, the memory controller 138 can be an external component of the processing system 100. The memory units 140 and 142 can be of the double data rate (DDR) type or the low-power DDR (LPDDR) type of memory. The peripheral interconnect 144 can be any component or collection of components that allows the processing system 100 to communicate with other devices / components and / or users. In one embodiment, the peripheral interconnect 144 can be used to send data, control, or management messages from the processor 100 to an application installed on a host device and / or a remote device. In another example, the peripheral interconnect 144 can be used to allow a user or a user device (e.g., a personal computer (PC), etc.) to interact / communicate with the various processing units of the processing system 100.
[0039] Figure 2A An exemplary DAG computing model 180 includes three nodes N1 182, N2 184, and N3 186. The DAG computing model 180 can be, for example, a graph used in image processing, video processing, or deep neural network processing applications. Specifically, the DAG computing model 180 can be a graph model of any type of application processing that can be split into multiple layers or independent synchronous computing tasks. Figure 2AShows a DAG computing model 180 including 3 nodes; however, it should be understood that DAG computing can have more than 2 nodes. In this example, each node N1 182, N2 184, and N3 186 is associated with an independent task or computing block in the DAG computing model 180.
[0040] In a processing system having multiple processor types, a processing unit can be used to schedule or allocate each node to a specific processing unit according to the specific tasks to be completed at the node. This type of scheduling is typically to take advantage of the optimized processing inherent in different processing types. For example, in Figure 1 processing system 100, tasks can be assigned to GPU 110 to process node N1 182, tasks can be assigned to DSP 114 to process node N2 184, and tasks can be assigned to CPU 102 to process node N3 186. Each node can in turn be divided into multiple subtasks or multiple computing blocks, as further detailed below.
[0041] In a processing system where different processing types include multiple cores, each subtask or computing block in a node can be processed by an independent core of a specific processing type through the scheduling of the processing unit. In the parallel processing within the node of the DAG computing model 180, the task related to node N2 184 starts only after the task related to node N1 182 is completed. In other words, the output 188 of node N1 182 is the input of node N2 184; the output 190 of node N2 182 is the input of node N3 186; and so on for other nodes.
[0042] Generally, since each node of the DAG computing model 180 is arranged in a sequential and interdependent configuration, the total time for processing the DAG computing is the accumulation of the time for processing each node. For example, if it takes T1 time to process all the subtasks of node N1 182, T2 time to process all the subtasks of node N2 184, and T3 time to process all the subtasks of node N3 186, then the total time Ttotal for processing this computing model is T1 + T2 + T3. During time T1, the processing unit assigned to node N1 182 is in an active state, while the processing units assigned to node N2 184 and node N3 186 are in an idle state. During time T2, the processing unit assigned to node N2 184 is in an active state, while the processing units assigned to node N1 182 and node N3 186 are in an idle state. During time T3, the processing unit assigned to node N3 186 is in an active state, while the processing units assigned to node N1 182 and node N2 184 are in an idle state. Embodiments of the present invention provide a method for reconstructing the DAG computing model 180 to minimize the idle time of different processing units.
[0043] Figure 2B An example of the hierarchical splitting 181 of the DAG node 183; in the first level of the hierarchy, the node A 183 is split into child nodes A1 185, A2 187... Ak 189. In the second level of the hierarchy, the child node A2 187 is split into child nodes A2-1 191, A2-2 193... A2-L 195. The splitting of the node can continue in other hierarchical levels, such as the third level, the fourth level, the fifth level, etc. Although the child node A2 187 is shown as being split into other child nodes, it should be understood that any number of child nodes of A, such as A1-Ak 185-189, can be split into other hierarchical child nodes. The number of child nodes for each hierarchical splitting is non-limiting and can be any appropriate number related to the corresponding subtasks in the DAG computing model representing the DAG node 183.
[0044] Figure 3 An exemplary DAG computing model 200 including two nodes N1 202 and N2 204, where each node has multiple child nodes. In the context of the DAG computing model 200, the node 2 204 depends on the node 1 202. As shown, the node N1 202 includes four (4) child nodes N1-1 206, N1-2 208, N1-3 210, and N1-4 212. The node N2 204 includes sixteen (16) child nodes N2-1 214, N2-2 216... N2-15 242, and N2-16 244. Although four (4) child nodes of the node N1 202 and sixteen (16) child nodes of the node 204 are shown, the number of child nodes can be application-related, and the number of child nodes in each node is shown for ease of description.
[0045] In one example, regarding in-node parallel processing, each sub-task of node N1 202 can be processed by the cores of CPU 102, and each sub-task of node N2 204 can be processed by the cores of GPU 110. In another example, also regarding in-node parallel processing, each sub-task of node N1 202 can be processed by the cores of GPU 110, and each sub-task of node N2 204 can be processed by the cores of DSP 114. In yet another example, regarding in-node parallel processing, each sub-task of node N1 202 can be processed by the cores of ISP 118, and each sub-task of node N2 204 can be processed by the cores of DPU 130. In another example, regarding in-node parallel processing, some sub-tasks of node N1 202 can be processed by the cores of CPU 102, and some sub-tasks of node N1 202 can be processed by the cores of GPU 110. In this example, some sub-tasks of node N2 204 can be processed by the cores of DSP 114, and other sub-tasks of node N2 204 can be processed by the cores of ISP 118. It should be noted that each sub-task can be operated by different cores of a specific type of processing unit, and a specific processing unit can be selected to improve the computational efficiency based on the available processing implemented in the dedicated hardware unit.
[0046] In the implementation of in-node parallel processing for the DAG computing model 200, after each sub-task in a node is completed, the subsequent nodes of the DAG computing model 200 start processing only after all the sub-tasks of the previous node are completed. This is represented in the form of a dependency in each node to receive a complete set of outputs from the previous node. Thus, in an example where two different processing units are used to process the DAG computing model 200, the first processing unit actively processes node N1 202, while the second processing unit can be idle and wait for the first processing unit to complete the computation. Similarly, the first processing unit remains idle and the second processing unit actively processes node N2 204.
[0047] Figure 4 FIG. 250 is a flow chart of an exemplary method 250 for splitting a DAG computing model and constructing multiple sub-DAG computing models for inter-node parallel processing executable by a processing system 100. The DAG computing model has a topological order, in which each node points from an earlier node in the node sequence. In step 252, the processing system 100 identifies the ordered and acyclic set of nodes in the DAG computing model.
[0048] In step 254, the processing system 100 splits each identified node into non - dependent sub - nodes based on the task type and computational storage requirements corresponding to each sub - node and each node. The nodes can be split into uniform, non - uniform, or overlapping sub - nodes. When splitting nodes uniformly, the size of each sub - node or sub - task can be equal, while when splitting nodes non - uniformly, the size of each sub - node or sub - task can be different or unequal. When splitting nodes overlappingly, some sub - tasks can overlap with one or more other sub - tasks, or sub - tasks can have an intersection with other sub - tasks at the sub - task boundary.
[0049] For example, regarding image processing and uniform node splitting, an image can be further divided into equal smaller N - by - M (N×M) segments. For example, regarding image processing and non - uniform node splitting, an image can be further divided into unequal smaller N - by - M (N×N) segments. For example, regarding image processing and overlapping node splitting, an image can be further divided into unequal or equal but overlapping smaller N - by - M (N×N) segments.
[0050] In step 256, the processing system 100 constructs multiple sub - DAG computational models using non - dependent sub - nodes in different nodes of the original DAG computational model. It should be understood that a sub - DAG computational model has at least non - dependent sub - nodes in two different nodes, but the construction of multiple sub - DAG computational models can vary according to the computational tasks associated with the sub - nodes.
[0051] In some embodiments, each sub - DAG computational model can have a single non - dependent sub - node in each node of the original DAG computational model. In some embodiments, each sub - DAG computational model can have non - dependent sub - nodes in some nodes of the original DAG computational model. In other embodiments, some sub - DAG computational models can have non - dependent sub - nodes in each node of the original DAG computational model, while some sub - DAG computational models can have non - dependent sub - nodes in some nodes of the original DAG computational model.
[0052] The construction of the multiple sub - DAG computational models can be performed manually or automatically by a compiler. In an example of manual construction, in a DAG computational model with less than 5 nodes, the construction of multiple sub - DAG computational models can be completed using a pre - configured static mapping table. The pre - configured static mapping table can be used to map the original DAG computational model into multiple sub - DAG computational models.
[0053] As an example of an automated build or a compiler-assisted build, which is typically applicable to more complex models with multiple DAG nodes, the compiler can dynamically convert the original DAG computation model into multiple sub-DAG computation models at runtime. In some embodiments, an OFFLINE compiler can be used to precompile the conversion from the original DAG computation model to multiple sub-DAG computation models.
[0054] In step 258, the processing system 100 allocates memory for multiple sub-DAG computations using the intermediate shared memory (cache) 136. The intermediate shared memory 136 can be used as a temporary storage location for the outputs of the sub-node computations, and the outputs are used as inputs for subsequent sub-nodes of the same sub-DAG computation model. The intermediate shared memory acts as a buffer memory and reduces the read and write times associated with off-chip memory (e.g., external double data rate (DDR) type memory or cache memories such as L1, L2, etc. in the processing unit). In some embodiments, if there are no resource dependencies between the steps of splitting the DAG computation (step 254) and allocating memory (step 258), these steps can be performed simultaneously. In some embodiments, if there are no resource dependencies between the steps of building the sub-DAG computation model (step 256) and allocating memory (step 258), these steps can be performed simultaneously. In some embodiments, these steps can be performed at different times.
[0055] In step 260, the processing system 100 schedules synchronization and dynamic tasks associated with each sub-DAG computation of the multiple sub-DAG computations, for example, using the CPU 102 or 104. In some embodiments, the generated sub-DAG computation models can be different from another non-mutually dependent sub-DAG computation model. Initially, resources are allocated to the multiple sub-DAG computation models at a high level, and subsequently, the processing system 100 schedules each sub-node in each sub-DAG computation at a lower processing level associated with each sub-task.
[0056] In a DAG computation model, each node is restricted by the previous node. Similarly, each sub-node corresponding to a sub-DAG computation model is restricted by the previous sub-node. Scheduling provides the order in which each sub-task is executed in the sub-DAG computation model. In other words, scheduling provides a topological sort for the sub-tasks in the sub-DAG computation model.
[0057] The scheduling in step 260 can be inter-node and / or intra-node scheduling through one of the processing unit types. Topological sorting provides an efficient way to execute these tasks between each sub-DAG computational model and within each sub-DAG computational model based on the interdependencies of the task set and shared resources. As a result of the scheduling, the total time period for processing the original DAG computational model is reduced because there is less idle time associated with the different processing units in the processing system 100.
[0058] In step 262, the processing system 100 processes each of the multiple sub-DAG computations and compiles the relevant output files. After the inter-node parallel processing of each of the multiple sub-DAG computational models is completed, a final output equal to the final output generated by the intra-node parallel processing of the original DAG computational model is generated.
[0059] Figures 5A to 5D Illustrated is the construction of multiple sub-DAG computational models from a DAG computational model 300 using an embodiment of the present invention, as may be performed by a processing system 100 in an image processing application. An example of an image processing application that can be modeled using a DAG computational model is image blurring, with applications in video games, presentations, or high dynamic range (HDR) rendering. In these applications, image blurring or patterns can be used to reproduce the image effects of real-world cameras.
[0060] Figure 5A The DAG computational model 300 includes three nodes: a first node (node 1) 302, a second node (node 2) 304, and a third node (node 3) 306. It should be noted that other nodes can also be considered. For example, in an image processing application, each node is mapped to a specific computational task.
[0061] For example, the first node 302 may correspond to obtaining an input image, the second node 304 may correspond to converting the input image into an integral image, and the third node 306 may correspond to generating an output image based on the integral image using Gaussian filtering.
[0062] In one embodiment, the output file can be an output data buffer. In another embodiment, the output file can be an output image buffer. In another embodiment, the output can be an output image file. In some embodiments, the output file can be an output feature set of the DAG computational model. It should be understood that the specific arrangement of the specific nodes in the DAG computational model 300 is not the main subject of the present invention, and the DAG computational model 300 can be used as a general DAG computational model for describing the construction of new DAG computational models in other applications.
[0063] In Figure 5BIn this case, each node of the DAG computing model 300 can be further divided into multiple subtasks or sub-nodes. The first node 302 is divided into sub-nodes 1-1 308, sub-nodes 1-2 310, sub-nodes 1-3 312, and sub-nodes 1-4 314. The second node 304 is divided into sub-nodes 2-1 316, sub-nodes 2-2 318, sub-nodes 2-3 320, and sub-nodes 2-4 322. The third node 306 is divided into sub-nodes 3-1 324, sub-nodes 3-2 326, sub-nodes 3-3 328, and sub-nodes 3-4 330.
[0064] The division of subtasks in each task can be uniform, non-uniform, or overlapping. For example, the division 332 of the subtasks related to sub-nodes 1-3 312 and sub-nodes 1-4 314 can be of the carry_on row type, and the division 334 of the subtasks related to sub-nodes 2-3 320 and sub-nodes 2-4 322 can have an overlapping area at the boundary. In some embodiments, a DAG computing model with two adjacent sub-blocks can have a mutual dependency relationship with each other. For example, the input of sub-node 2-4 322 can be the output of sub-node 2-3 320. In these embodiments, each row in the intersection area can be a carry row, which is a position indication for calculating the carry result of adjacent sub-nodes. Among them, the overlapping area can be the intersection area between two adjacent sub-blocks, and can be a single row or multiple rows.
[0065] Each subtask can be mapped to the same or different computing subtasks related to the specific computing tasks of the corresponding node. In the parallel processing within the nodes of the DAG computing model 300, each subtask can be scheduled for different cores of the same processing unit. In this type of processing, the scheduling granularity is limited within a single DAG node. Therefore, parallelism is not implemented between DAG nodes or within DAG nodes. Since the scheduling of subsequent nodes cannot start until the scheduling of the current node is completed, the utilization rate of hardware resources is low.
[0066] Figure 5C An embodiment DAG computing model 303 is shown that includes multiple sub-DAG computing models that can be calculated by the processing system 100. The DAG computing model 303 is a modified version of the DAG computing model 300. The DAG computing model 303 includes five (5) sub-DAG computing models 352, 354, 356, 358, and 360. Although five (5) sub-DAG computing models are shown for the purpose of description, Figure 5C in this case, the total number of sub-DAG computing models can be any number greater than one (1).
[0067] In this new arrangement, computational parallelism can be achieved using inter-node parallelism, and using both inter-node parallelism and intra-node parallelism. In the new DAG computational model 303 arrangement of the child nodes, multiple hardware resources (e.g., processing units) can be used to compute the new sub-DAG computational models in parallel. In embodiments where each sub-DAG computational model is independent of another sub-DAG computational model, the total processing time is reduced from T1+T2+T3 to the longer total time of (T1+T2) or (T2+T3).
[0068] Each child node in each sub-DAG computational model is arranged and constructed to have a more optimized dependency model among the child nodes of all the nodes. This is done to improve the efficiency of the DAG computational model 300 and reduce the processing time. Each child node is processed by a different core of the processing unit. However, the arrangement of the child nodes in the sub-DAG computational model results in shorter idle times between the processing of the child nodes in the newly constructed model. As before, each processing unit is assigned to a child node based on the specific capabilities of the processing unit and the subtasks associated with the child node.
[0069] As shown, the first sub-DAG computational model 352 includes child node 2-1 316 that depends on child node 1-1 308. The second sub-DAG computational model 354 includes, in addition to child node 3-1 324 that depends on child node 2-1 316, child node 2-2 318 that depends on child node 1-2 310. The third sub-DAG computational model 356 includes, in addition to child node 3-2 326 that depends on child node 2-2 318, child node 2-3 320 that depends on child node 1-3 312. The fourth sub-DAG computational model 358 includes, in addition to child node 3-3 328 that depends on child node 2-3 320, child node 2-4 322 that depends on child node 1-4 314. Finally, the fifth sub-DAG computational model 360 includes child node 3-4 330 that depends on child node 2-4 330.
[0070] The output of the first child node of the sub-DAG computational model 352 is the input of the second child node of the sub-DAG computational model 352. Similarly, there may still be dependencies between one sub-DAG computational model and another. However, the completion time of the subtasks associated with the first sub-DAG computational model 352 is less than the completion time of the entire task associated with the DAG computational model 300. Other cores of the processing unit can be scheduled to execute other child nodes in the same or other sub-DAG computational models. Thus, the time during which the processing unit remains idle and waits for another processing unit to complete a task is significantly reduced.
[0071] Figure 5D An exemplary data flow in the memory block of the processing system 100 is shown, corresponding toFigure 5B the conversion from the DAG computing model 300 in Figure 5C to the DAG computing model 303 in. Each node of the DAG computing model 300 is divided into smaller child nodes or subtasks, which can be divided evenly or unevenly. Then, positions in memory are allocated for each sub-block of the first node 382, each sub-block of the second node 384, and each sub-block of the third node 386. In block 394, each sub-block of each node is then queued in memory, and information related to the queuing address, size, shape, order information, etc. is recorded. The splitter 395 and the scheduler 397 use the information stored in block 394 to generate a new queue for the new sub-DAG computing model. Block 398 shows an optional intermediate repository accessible from each processing unit of the processing system 100. The intermediate memory can be used to store the output results in and between the sub-DAG computing models for use by other processing units.
[0072] Figures 6A to 6C shows the construction of multiple sub-DAG computing models from the DAG computing model 450 using an embodiment of the present invention, as can be performed by the processing system 100 in a deep neural network (DNN) type application. A deep neural network is a type of machine learning that uses data representation and typically includes multiple layers: an input layer, intermediate layers (i.e., hidden layers), and an output layer. Each layer or node has a related function different from any other layer, such as image convolution, pooling, normalization, feature map generation, etc.
[0073] In the deep neural network DAG computing model 450, data flows from the input layer or the first node (node 1) 452 to the output layer or the third node (node 3) 456 without looping back. The first node 452 and the second node 454 of the deep neural network DAG computing model 450 include matrix inputs and corresponding matrix weights. The output node 456 of the deep neural network DAG computing model 450 is a normalized exponential representation using, for example, the softmax function 470. The deep neural network model has a first layer and a second layer, but other nodes can also be considered.
[0074] The first node 452 includes a first matrix input 462 and a first matrix weight 464. The second node 454 includes a second matrix input 466 and a second matrix weight 468. In a typical deep neural network application, the input matrix and the weight matrix in each node are multiplied to obtain a function output representation between 0 and 1. The deep neural network adjusts the weights and evaluates the corresponding outputs until a specific pattern is recognized.
[0075] In Figure 6BIn this case, each input matrix 462 and 466 of each node 452 and 454 is subdivided into four (4) sub - matrices. Although in this example the input matrix is subdivided into four sub - matrices, in other examples, the sub - division can be any number greater than one (1).
[0076] In typical solutions for solving deep neural networks in DAG computations using intra - node parallelism, such as those found in CAFFE or TensorFlow, each computational task associated with a node is scheduled layer - by - layer. In each layer, intra - node parallelism can be achieved by multiple cores of a specific processing unit of the processing system 100. In intra - node parallel processing, the scheduling of the second node (input2×weight2) starts only after the scheduling of the first node (input1×weight1) is completed. The completion of the first node corresponds to solving the first node (i.e., multiplying each input node by the weights of that node and completing the pattern recognition process).
[0077] Figure 6C An improved DAG computation model 455 based on the subdivided input matrix and corresponding weights is shown. The DAG computation model 455 includes four (4) sub - DAG computation models 510, 520, 530, and 540. Although for the purpose of description, four (4) sub - DAG computation models are shown Figure 6C in this case, the total number of sub - DAG computation models can be any number greater than one (1). In this new arrangement, computational parallelism can be achieved using inter - node parallelism, as well as using both inter - node parallelism and intra - node parallelism. In the new DAG computation model 455 arrangement of the sub - nodes, multiple hardware resources (e.g., processing units) can be used to compute the new sub - DAG computation models. Each subdivided matrix corresponds to a sub - task in the nodes of the DAG computation model 450. In this revised model, inter - node parallel processing can be achieved by splitting the original model into smaller sub - tasks and re - arranging the dependencies from the tasks in the DAG computation model 450 to the sub - tasks in each sub - DAG computation model.
[0078] Each sub - node in each sub - DAG computation model is arranged and constructed to have a more optimized dependency model among the sub - nodes of all the nodes. This is done to improve the efficiency of the DAG computation model 450 and reduce the processing time. Each sub - node is processed by a different core of the processing unit. However, the arrangement of the sub - nodes in the sub - DAG computation model results in shorter idle times between the processing of sub - nodes in the newly constructed model. As before, each processing unit is assigned to a sub - node according to the specific capabilities of the processing unit and the sub - tasks associated with the sub - node.
[0079] As shown in the figure, the first sub-DAG computing model 510 includes a child node 2-1 504 that depends on a child node 1-1 502, and a child node 3-1 506 that depends on the child node 2-1 504. The second sub-DAG computing model 520 includes a child node 2-2 514 that depends on a child node 1-2 512, and a child node 3-2 516 that depends on the child node 2-2 514. The third sub-DAG computing model 530 includes a child node 2-3 524 that depends on a child node 1-3 522, and a child node 3-3 526 that depends on the child node 2-3 524. Moreover, the fourth sub-DAG computing model 540 includes a child node 2-4 534 that depends on a child node 1-4 532, and a child node 3-4 536 that depends on the child node 2-4 534.
[0080] Figure 7 An exemplary DAG computing model 550 for computer vision type applications is shown. An example of a computer vision type application is an OpenVX graph. The OpenVX graph is an open and royalty-free standard method for cross-platform acceleration of computer vision applications. The OpenVX graph includes multiple steps for end-to-end image and / or video computing. Some examples of these independent steps are color conversion, channel extraction, image pyramid, optical flow, etc.
[0081] Each step of a computer vision type application (such as an OpenVX graph) can be represented by a DAG node. The DAG computing model 550 is an example of an OpenVX graph. The DAG computing model 550 includes a color conversion node 552, a channel extraction node 554, an image pyramid node 556, a pyramid node 558, an optical flow node 560, a Harris corner node 562, and a key point node 564. Understanding the specific functions of each node is not necessary for understanding the transformation of the DAG computing model 550 from a model arranged for in-node parallel processing to a model allowing inter-node parallel processing. This figure is used to show that in a typical computer vision application, computing tasks (such as YUV frame or grayscale frame generation) can be arranged in the DAG computing model.
[0082] Embodiments of the present invention provide a method of splitting each node of the DAG computing model 550 into multiple subtasks. Then, similar to the methods described previously in image processing, each subtask can be rearranged with subtasks or child nodes of other nodes of the DAG computing model 550, for example, as Figures 5A to 5C shown. The new sub-DAG computing model allows for inter-node processing both between and within each node. Therefore, the new DAG computing model allows for a faster processing time with less idle processing time for other processing units.
[0083] It should be noted that the above examples regarding image processing, deep neural networks, and video processing are non-limiting examples, and the corresponding descriptions for splitting the original DAG calculation model and constructing a new sub-DAG calculation model can be applied to any application that can be formed using the DAG calculation model.
[0084] Figure 8 Embodiment DAG calculation model 600 and corresponding construction embodiment DAG calculation model 620 are shown. The construction embodiment DAG calculation model 620 has a one-to-many mapping graph model that can be calculated by processing system 100. In general applications, such as in image processing, video processing, or deep neural network processing, each sub-DAG of the corresponding DAG calculation model can have two or more child nodes. In this arrangement, one or more child nodes can be determined based on the inputs of multiple child nodes. Also, one or more child nodes can provide inputs to multiple child nodes.
[0085] The DAG calculation model 600 shown in the figure has three nodes: node 1 602, node 2 604, and node 3 606. It should be understood that DAG calculation models with a larger number of nodes can also be considered. However, for simplicity of description, three nodes are shown.
[0086] The DAG calculation model 620 shows splitting each node in the DAG calculation model 600 into multiple child nodes and constructing multiple sub-DAG calculation models. Node 1 602 is split into child nodes 1-1 632, child nodes 1-2 634, child nodes 1-3 636, and child nodes 1-4 638. Node 2 604 is split into child nodes 2-1 640, child nodes 2-2 642, child nodes 2-3 644, and child nodes 2-4 646. Node 3 606 is split into child nodes 3-1 648, child nodes 3-2 650, child nodes 3-3 652, and child nodes 3-4 654.
[0087] Corresponding to Figure 8 An exemplary arrangement of the one-to-many mapping graph model in shows the construction of the DAG calculation model 620 and the dependencies of one or more child nodes. As shown, child node 2-1 640 depends on child node 1-1 632. Child node 2-2 642 depends on the inputs of child node 1-2 634 and child node 2-1 640. Child node 3-1 648 depends on child node 2-1 640. Child node 2-3 644 depends on child node 1-3 636. Child node 3-2 650 depends on the inputs of child node 1-3 636 and child node 2-2 642. Child nodes 2-4 646 and child node 3-3 652 both depend on the inputs of child node 1-4 638 and child node 2-3 644. Child node 3-4 654 depends on child node 2-4 646. Although Figure 8The example shows that each child node has multiple dependencies, but it should be understood that in some embodiments, the arrangement of dependencies can be different. For example, some child nodes with a single input as a dependency can have multiple dependencies. In some embodiments, the scheduling of each child node of the DAG computing model 600 can be executed by the CPU 102 / 106, GPU 110, or DSP 114 processing units of the processing system 100.
[0088] Figure 9 FIG. is a schematic diagram of a network 700 for sending data. The network 700 includes a base station 710 having a coverage area 701, a plurality of UEs 720, and a backhaul network 730. As shown, the base station 710 establishes uplink (dashed line) and / or downlink (dotted line) connections with the UEs 720, which are used to send data from the UEs 720 to the base station 710 and vice versa. The data sent through the uplink / downlink connections can include data sent between the UEs 720 and data sent to / from a remote end (not shown) through the backhaul network 730. As used herein, the term "base station" refers to any network-side device for providing wireless access to a network, such as an enhanced base station (eNodeB or eNB), gNB, transmit / receive point (TRP), macro cell, femto cell, Wi-Fi access point (AP), and other wireless devices. The base station can provide wireless access according to one or more wireless communication protocols, such as, for example, 5th Generation New Radio (5G NR), LTE, LTE-Advanced (LTE-A), High Speed Message Access (HSPA), Wi-Fi 802.11a / b / g / n / ac, etc. As used herein, the term "UE" refers to any user-side device for accessing a network by establishing a wireless connection with a base station, such as a mobile device, mobile station (STA), vehicle, and other wireless devices. In some embodiments, the network 700 can include various other wireless devices, such as repeaters, low-power nodes, etc. It can be understood that a communication system can employ multiple access nodes capable of communicating with multiple UEs, and only one base station 710 is shown for simplicity, with two UEs 720 shown.
[0089] Figure 10 FIG. shows a block diagram of another embodiment of a processing system 800 for performing the methods described herein, which system can be installed in a host device. As shown, the processing system 800 includes a processor 802, a memory 804, and interfaces 806, 808, 810, which may (or may not) be as Figure 10be arranged as shown. The processor 802 can be any component or collection of components for performing computing and / or other processing-related tasks, and the memory 804 can be any component or collection of components for storing the programming and / or instructions executed by the processor 802. In one embodiment, the memory 804 includes a non-transitory computer-readable medium. The interfaces 806, 808, 810 can be any component or collection of components that allow the processing system 800 to communicate with other devices / components and / or users. In one embodiment, one or more of the interfaces 806, 808, 810 can be used to send data, control, or management messages from the processor 802 to an application installed on a host device and / or a remote device. As another example, one or more of the interfaces 806, 808, 810 can be used to allow a user or a user device (e.g., a personal computer (PC), etc.) to interact / communicate with the processing system 800. The processing system 800 can include Figure 10 other components not shown in FIG. Figure 10 , such as long-term memory (e.g., non-volatile memory).
[0090] In some embodiments, the processing system 800 is included in a network device that accesses a telecommunications network or otherwise forms part of a telecommunications network. In one embodiment, the processing system 800 is located in a network-side device in a wireless or wired telecommunications network, such as a base station, a relay station, a scheduler, a controller, a gateway, a router, an application server, or any other device in a telecommunications network. In other embodiments, the processing system 800 is located in a user-side device that accesses a wireless or wired telecommunications network, such as a mobile station, a user equipment (UE), a personal computer (PC), a tablet, a wearable communication device (e.g., a smartwatch, etc.), a wireless-enabled vehicle, a wireless-enabled pedestrian, a wireless-enabled infrastructure element, or any other device for accessing a telecommunications network.
[0091] In some embodiments, one or more of the interfaces 806, 808, 810 connect the processing system 800 to a transceiver for sending and receiving signaling over a telecommunications network. Figure 11A block diagram of a transceiver 900 for transmitting and receiving signaling via a telecommunications network is shown. The transceiver 900 can be installed in a host device. As shown, the transceiver 900 includes a network-side interface 902, a coupler 904, a transmitter 906, a receiver 908, a signal processor 910, and a device-side interface 912. The network-side interface 902 can include any component or collection of components for transmitting or receiving signaling via a wireless or wired telecommunications network. The coupler 904 can include any component or collection of components for facilitating two-way communication via the network-side interface 902. The transmitter 906 can include any component or collection of components for converting a baseband signal into a modulated carrier signal suitable for transmission via the network-side interface 902 (e.g., an upconverter, a power amplifier, etc.). The receiver 908 can include any component or collection of components for converting a carrier signal received by the network-side interface 902 into a baseband signal (e.g., a downconverter, a low-noise amplifier, etc.). The signal processor 910 can include any component or collection of components for converting a baseband signal into a data signal suitable for communication via the device-side interface 912, or vice versa. The device-side interface 912 can include any component or collection of components for transmitting data signals between the signal processor 910 and components within the host device (e.g., a processing system 1300, a local area network (LAN) port).
[0092] The transceiver 900 can send and receive signaling via any type of communication medium. In some embodiments, the transceiver 900 sends and receives signaling via a wireless medium. In some embodiments, the transceiver 900 can be a wireless transceiver for communicating according to a radio communication protocol, such as a cellular protocol (e.g., long-term evolution (LTE), etc.), a wireless local area network (WLAN) protocol (e.g., Wi-Fi, etc.), or any other type of wireless protocol (e.g., Bluetooth, near field communication (NFC), etc.). In these embodiments, the network side interface 902 includes one or more antennas / radiation units. In some embodiments, the network side interface 902 can include a single antenna, multiple independent antennas, or a multi-antenna array for multi-layer communication such as single input multiple output (SIMO), multiple input single output (MISO), multiple input multiple output (MIMO), etc. In other embodiments, the transceiver 900 sends and receives signaling via a wired medium such as a twisted pair cable, coaxial cable, optical fiber, etc. A particular processing system and / or transceiver may use all of the components shown, or only a subset of the components, and the level of integration may vary according to the device.
[0093] Although the description has been detailed, it should be understood that various changes, substitutions, and alterations can be made without departing from the spirit and scope of the present disclosure as defined by the appended claims. Identical elements are denoted by the same reference numerals in the various figures. Additionally, the scope of the present invention is not intended to be limited to the specific embodiments described herein, and those of ordinary skill in the art will readily appreciate from the present invention that processes, machines, manufacturing processes, compositions of matter, components, methods, or steps (including those that currently exist or will be developed later) can perform substantially the same functions or achieve substantially the same effects as the corresponding embodiments described herein. Accordingly, the scope of the appended claims includes such processes, machines, products, compositions of matter, means, methods, and steps. Therefore, the description and drawings should be regarded simply as an illustration of the invention as defined by the appended claims, and should be considered to cover any and all modifications, variations, combinations, or equivalents that fall within the scope of the invention.
Claims
1. A method for processing directed acyclic graph (DAG) calculations, characterized in that, Comprising: A plurality of processors split the DAG calculation into a plurality of non - mutually - dependent sub - nodes in each corresponding node, wherein the plurality of processors includes a multi - core graphics processing unit (GPU) and a multi - core central processing unit (CPU); The plurality of processors construct a plurality of sub - DAG calculations, where each sub - DAG calculation includes at least non - mutually - dependent sub - nodes in different nodes of the DAG calculation; among the plurality of sub - DAG calculations, there are some sub - DAG calculations that include a plurality of non - mutually - dependent sub - nodes in the different nodes, and the non - mutually - dependent sub - nodes in the partial sub - DAG calculations are processed by different cores of the multi - core graphics processing unit and the multi - core central processing unit; The plurality of processors schedule the processing of each sub - DAG calculation among the plurality of sub - DAG calculations; The plurality of processors process each sub - DAG calculation among the plurality of sub - DAG calculations in parallel.
2. The method according to claim 1, wherein Further comprising: The plurality of processors allocate an intermediate shared memory for the plurality of sub - DAG calculations.
3. The method according to claim 1 or 2, characterized in that, Further comprising: The plurality of processors schedule the processing of each sub - node of each sub - DAG calculation.
4. The method according to claim 1, wherein The scheduling of the processing of each sub - DAG calculation among the plurality of sub - DAG calculations is performed according to the task type of the corresponding sub - nodes of each sub - DAG calculation.
5. The method according to claim 1, characterized in that The DAG calculation includes an image processing application, a video processing application, or a deep neural network processing application.
6. The method according to claim 1, wherein If the processing of the next sub - node in the corresponding sub - DAG calculation starts, the processing of the sub - node in the corresponding sub - DAG calculation is regarded as completed.
7. The method according to claim 1, characterized in that, The processing of non - mutually - dependent sub - nodes in a corresponding node is independent of the processing of another non - mutually - dependent sub - node in the same corresponding node.
8. The method according to claim 1, characterized in that, The construction is completed manually or automatically by a compiler executed by the plurality of processors.
9. The method according to claim 1, characterized in that Further comprising receiving an input for the DAG calculation.
10. The method according to claim 1, wherein Further comprising an output data buffer, an output image buffer, an output image file, or an output feature for outputting the DAG calculation.
11. The method according to claim 1, characterized in that, The splitting is performed according to the task type associated with each node and the storage capabilities of the plurality of processors.
12. The method according to claim 1, characterized in that, The splitting includes splitting into uniform non - mutually - dependent sub - nodes.
13. The method according to claim 1, wherein The splitting includes splitting into non - uniform non - mutually - dependent sub - nodes.
14. The method according to claim 1, wherein The splitting includes covering the boundaries of non - mutually - dependent sub - nodes.
15. The method according to claim 1, characterized in that Each sub - node is a sub - task associated with the corresponding node of the DAG calculation.
16. The method according to claim 1, wherein One or more nodes of the DAG calculation are split hierarchically.
17. The method according to claim 1, characterized in that One or more sub - nodes of one or more sub - DAG calculations are determined according to the outputs of a plurality of sub - nodes.
18. The method according to claim 1, wherein One or more sub - nodes of one or more sub - DAG calculations are the inputs of a plurality of sub - nodes.
19. A computer-implemented method for processing directed acyclic graph (DAG) calculations, characterized in that, Comprising: Multiple processors split the DAG calculation into multiple non - interdependent sub - nodes in each corresponding node, where the multiple processors include a multi - core graphics processing unit (GPU) and a multi - core central processing unit (CPU); The multiple processors construct multiple sub - DAG calculations, where each sub - DAG calculation includes at least non - interdependent sub - nodes in different nodes of the DAG calculation; among the multiple sub - DAG calculations, there are some sub - DAG calculations that include multiple non - interdependent sub - nodes in different nodes, and the non - interdependent sub - nodes in the partial sub - DAG calculations are processed by different cores of the multi - core graphics processing unit and the multi - core central processing unit; The multiple processors schedule the processing of each sub - DAG calculation among the multiple sub - DAG calculations; The multiple processors process each sub - DAG calculation among the multiple sub - DAG calculations in parallel.
20. The computer-implemented method according to claim 19, wherein It further includes: The multiple processors allocate an intermediate shared memory for the multiple sub - DAG calculations.
21. The computer-implemented method according to claim 19 or 20, wherein It further includes: The multiple processors schedule the processing of each sub - node of each sub - DAG calculation.
22. The computer-implemented method according to claim 19, wherein The scheduling of the processing of each sub - DAG calculation among the multiple sub - DAG calculations is performed according to the task types of the corresponding sub - nodes of each sub - DAG calculation.
23. The computer-implemented method according to claim 19, wherein The DAG calculation includes an image processing application, a video processing application, or a deep neural network processing application.
24. The computer-implemented method according to claim 19, characterized in that, If the processing of the next sub - node in the corresponding sub - DAG calculation starts, the processing of the sub - node in the corresponding sub - DAG calculation is regarded as completed.
25. The computer-implemented method according to claim 19, wherein The processing of non - interdependent sub - nodes in a corresponding node is independent of the processing of another non - interdependent sub - node in the same corresponding node.
26. The computer-implemented method according to claim 19, wherein The construction is completed manually or automatically by a compiler executed by the multiple processors.
27. The computer-implemented method according to claim 19, wherein It further includes receiving an input for the DAG calculation.
28. The computer-implemented method according to claim 19, wherein It further includes an output data buffer, an output image buffer, an output image file, or output features for outputting the DAG calculation.
29. The computer-implemented method according to claim 19, characterized in that, The splitting is performed according to the task types associated with each node and the storage capabilities of the multiple processors.
30. The computer-implemented method according to claim 19, wherein The splitting includes splitting into uniform non - interdependent sub - nodes.
31. The computer-implemented method according to claim 19, wherein The splitting includes splitting into non - uniform non - interdependent sub - nodes.
32. The computer-implemented method according to claim 19, wherein The splitting includes covering the boundaries of non - interdependent sub - nodes.
33. The computer-implemented method according to claim 19, wherein Each sub - node is a sub - task associated with the corresponding node of the DAG calculation.
34. The computer-implemented method according to claim 19, wherein One or more nodes of the DAG calculation are split hierarchically.
35. The computer-implemented method according to claim 19, wherein One or more sub - nodes of one or more sub - DAG calculations are determined according to the outputs of multiple sub - nodes.
36. The computer-implemented method according to claim 19, wherein, One or more sub - nodes of one or more sub - DAG calculations are the inputs of multiple sub - nodes.
37. A non-transitory computer-readable medium, characterized in that, The non - transitory computer - readable medium stores computer instructions for processing a directed acyclic graph (DAG) calculation. When multiple processors including a first processing unit and a second processing unit execute the computer instructions, the following steps are performed: Split the DAG computation into multiple non - interdependent sub - nodes in each corresponding node, where the multiple processors include a multi - core graphics processing unit (GPU) and a multi - core central processing unit (CPU); Construct multiple sub - DAG computations, where each sub - DAG computation includes at least non - interdependent sub - nodes in different nodes of the DAG computation; among the multiple sub - DAG computations, there are some sub - DAG computations that include multiple non - interdependent sub - nodes in the different nodes, and the non - interdependent sub - nodes in the part of the sub - DAG computations are processed by different cores of the multi - core graphics processing unit and the multi - core central processing unit; Schedule the processing of each sub - DAG computation among the multiple sub - DAG computations; Process each sub - DAG computation among the multiple sub - DAG computations in parallel.
38. The non-transitory computer-readable medium according to claim 37, wherein When multiple processors execute the non - transitory computer - readable medium, perform the following steps: allocate an intermediate shared memory for the multiple sub - DAG computations.
39. The non-transitory computer-readable medium according to claim 37 or 38, wherein When multiple processors execute the non - transitory computer - readable medium, perform the following steps: schedule the processing of each sub - node of each sub - DAG computation.
40. The non-transitory computer-readable medium according to claim 37, wherein The scheduling of the processing of each sub - DAG computation among the multiple sub - DAG computations is performed according to the task type of the corresponding sub - nodes of each sub - DAG computation.
41. The non-transitory computer-readable medium according to claim 37, wherein The DAG computation includes an image processing application, a video processing application, or a deep neural network processing application.
42. The non-transitory computer-readable medium according to claim 37, wherein If the processing of the next sub - node in the corresponding sub - DAG computation starts, the processing of the sub - nodes in the corresponding sub - DAG computation is considered completed.
43. The non-transitory computer-readable medium according to claim 37, wherein The processing of non - interdependent sub - nodes in a corresponding node is independent of the processing of another non - interdependent sub - node in the same corresponding node.
44. The non-transitory computer-readable medium according to claim 37, wherein The construction is completed manually or automatically by a compiler executed by the multiple processors.
45. The non-transitory computer-readable medium according to claim 37, wherein When multiple processors execute the non - transitory computer - readable medium, perform the following steps: receive an input for the DAG computation.
46. The non-transitory computer-readable medium according to claim 37, wherein When multiple processors execute the non - transitory computer - readable medium, perform the following steps: output an output data buffer, an output image buffer, an output image file, or an output feature of the DAG computation.
47. The non-transitory computer-readable medium according to claim 37, wherein The splitting is performed according to the task type associated with each node and the storage capabilities of the multiple processors.
48. The non-transitory computer-readable medium according to claim 37, wherein The splitting includes splitting into uniform non - interdependent sub - nodes.
49. The non-transitory computer-readable medium according to claim 37, wherein The splitting includes splitting into non - uniform non - interdependent sub - nodes.
50. The non-transitory computer-readable medium according to claim 37, wherein The splitting includes covering the boundaries of non - interdependent sub - nodes.
51. The non-transitory computer-readable medium according to claim 37, wherein Each sub - node is a sub - task associated with the corresponding node of the DAG computation.
52. The non-transitory computer-readable medium according to claim 37, wherein One or more nodes of the DAG computation are split hierarchically.
53. The non-transitory computer-readable medium according to claim 37, wherein One or more sub - nodes of one or more sub - DAG computations are determined according to the outputs of multiple sub - nodes.
54. The non-transitory computer-readable medium according to claim 37, wherein One or more sub - nodes of one or more sub - DAG computations are the inputs of multiple sub - nodes.
55. A device for processing directed acyclic graph (DAG) calculations, characterized in that, Include: A non - transitory memory, including instructions; A plurality of processors, including a central processing unit (CPU) and a graphics processing unit (GPU), the plurality of processors communicating with the non-transitory memory, wherein the plurality of processors execute the instructions to: Split the DAG computation into a plurality of non-dependent sub-nodes in each corresponding node; Construct a plurality of sub-DAG computations, where each sub-DAG computation includes at least non-dependent sub-nodes in different nodes of the DAG computation; wherein, among the plurality of sub-DAG computations, there are partial sub-DAG computations that include a plurality of non-dependent sub-nodes in different nodes, and the non-dependent sub-nodes in the partial sub-DAG computations are processed by different cores of a multi-core graphics processing unit and a multi-core central processing unit; Schedule the processing of each sub-DAG computation among the plurality of sub-DAG computations; Process each sub-DAG computation among the plurality of sub-DAG computations in parallel.
56. The device according to claim 55, characterized in that, The plurality of processors execute the instructions to allocate an intermediate shared memory for the plurality of sub-DAG computations.
57. The device according to claim 55 or 56, characterized in that, The plurality of processors execute the instructions to schedule the processing of each sub-node of each sub-DAG computation.
58. The device according to claim 55, wherein, The scheduling of the processing of each sub-DAG computation among the plurality of sub-DAG computations is performed according to the task type of the corresponding sub-node of each sub-DAG computation.
59. The device according to claim 55, wherein, The DAG computation includes an image processing application, a video processing application, or a deep neural network processing application.
60. The device according to claim 55, characterized in that, If the processing of the next sub-node in the corresponding sub-DAG computation starts, the processing of the sub-node in the corresponding sub-DAG computation is considered completed.
61. The device according to claim 55, characterized in that, The processing of non-dependent sub-nodes in a corresponding node is independent of the processing of another non-dependent sub-node in the same corresponding node. The device according to claim 55, wherein, The construction is completed manually or automatically by a compiler executed by the plurality of processors. The device according to claim 55, wherein, The plurality of processors execute the instructions to receive the input of the DAG computation.
64. The device according to claim 55, characterized in that, The plurality of processors execute the instructions to output the output data buffer, output image buffer, output image file, or output features of the DAG computation. The device according to claim 55, characterized in that, The splitting is performed according to the task type associated with each node and the storage capabilities of the plurality of processors.
66. The device according to claim 55, characterized in that, The splitting includes splitting into uniform non-dependent sub-nodes.
67. The device according to claim 55, characterized in that The splitting includes splitting into non-uniform non-dependent sub-nodes.
68. The device according to claim 55, characterized in that, The splitting includes covering the boundaries of non-dependent sub-nodes. The device according to claim 55, wherein Each sub-node is a sub-task associated with the corresponding node of the DAG computation. The device according to claim 55, characterized in that, One or more nodes of the DAG computation are split hierarchically.
71. The apparatus according to claim 55, characterized in that, One or more sub-nodes of one or more sub-DAG computations are determined according to the outputs of a plurality of sub-nodes.
72. The device according to claim 55, characterized in that One or more sub-nodes of one or more sub-DAG computations are the inputs of a plurality of sub-nodes.
73. A device for processing directed acyclic graph (DAG) calculations, characterized in that, Comprising: A non-transitory memory, including instructions; Multiple processors, including a first processing unit and a second processing unit, wherein the processor types of the first processing unit and the second processing unit are different, and the multiple processors communicate with the non-transitory memory, and the multiple processors execute the instructions to: Split the DAG computation into multiple non-dependent sub-nodes in each corresponding node; Construct multiple sub-DAG computations, where each sub-DAG computation includes at least non-dependent sub-nodes in different nodes of the DAG computation; among the multiple sub-DAG computations, there are some sub-DAG computations that include multiple non-dependent sub-nodes in different nodes, and the non-dependent sub-nodes in the partial sub-DAG computations are processed by different cores of a multi-core graphics processing unit and a multi-core central processing unit; Schedule the processing of each sub-DAG computation among the multiple sub-DAG computations; Process each sub-DAG computation among the multiple sub-DAG computations in parallel.
74. The apparatus according to claim 73, wherein The multiple processors execute the instructions to allocate an intermediate shared memory for the multiple sub-DAG computations. The device according to claim 73 or 74, characterized in that, The multiple processors execute the instructions to schedule the processing of each sub-node of each sub-DAG computation.
76. The device according to claim 73, characterized in that, The scheduling of the processing of each sub-DAG computation among the multiple sub-DAG computations is performed according to the task type of the corresponding sub-nodes of each sub-DAG computation.
77. The device according to claim 73, characterized in that, The DAG computation includes an image processing application, a video processing application, or a deep neural network processing application. The device according to claim 73, wherein If the processing of the next sub-node in the corresponding sub-DAG computation starts, the processing of the sub-node in the corresponding sub-DAG computation is considered completed.
79. The device according to claim 73, characterized in that, The processing of non-dependent sub-nodes in a corresponding node is independent of the processing of another non-dependent sub-node in the same corresponding node. The device according to claim 73, wherein, The construction is manually or automatically completed by a compiler executed by the multiple processors.
81. The device according to claim 73, characterized in that, The multiple processors execute the instructions to receive the input of the DAG computation. The device according to claim 73, characterized in that, The multiple processors execute the instructions to output the output data buffer, output image buffer, output image file, or output features of the DAG computation. The device according to claim 73, characterized in that, The splitting is performed according to the task type associated with each node and the storage capabilities of the multiple processors. The device according to claim 73, characterized in that, The splitting includes splitting into uniform non-dependent sub-nodes. The device according to claim 73, characterized in that The splitting includes splitting into non-uniform non-dependent sub-nodes.
86. The device according to claim 73, characterized in that, The splitting includes covering the boundaries of non-dependent sub-nodes. The device according to claim 73, characterized in that, Each sub-node is a sub-task associated with the corresponding node of the DAG computation. The device according to claim 73, characterized in that The processor types of the first processing unit and the second processing unit are different and are selected from the following group of processors: central processing unit (CPU), graphics processing unit (GPU), digital signal processor (DSP), image signal processor (ISP), video processing unit (VPU), neural network processing unit (NPU), and display processing unit (DPU). The device according to claim 73, wherein, It further includes at least one of the following: an interconnect bus link, a shared storage unit, a storage controller, one or more storage units, a peripheral interconnect, or a combination thereof. The device according to claim 73, characterized in that One or more nodes of the DAG calculation are hierarchically split.
91. The device according to claim 73, characterized in that, One or more sub-nodes of one or more sub-DAG calculations are determined according to multiple sub-node outputs. The device according to claim 73, characterized in that, One or more sub-nodes of one or more sub-DAG calculations are inputs of multiple sub-nodes.
Citation Information
Patent Citations
Graph-based application programming interface architectures with node-based destination-source mapping for enhanced image processing parallelism
US20160210720A1