Method and apparatus for implementing out-of-order pipeline execution of statically mapped workloads
Through the combination of interfaces, comparators and dispatchers, the execution order of workload nodes on computing building blocks is dynamically adjusted, which solves the problem of insufficient resource utilization of static software scheduling in heterogeneous systems, realizes efficient out-of-order pipeline execution, improves processing efficiency and reduces memory overhead.
Patent Information
- Application Number
- CN202210600897.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-08-15
- Filing Date
- 2020-06-18
- Publication Date
- 2025-09-09
- Estimated Expiration
- 2040-06-18
AI Technical Summary
Existing static software scheduling has difficulty in effectively utilizing the computing building blocks of accelerators in heterogeneous systems, resulting in large memory overhead and difficulty in dynamically adjusting the execution order of workloads, affecting processing efficiency.
Through the combination of interfaces, comparators and dispatchers, the execution order of workload nodes on computing building blocks is dynamically determined according to the availability of buffers and memories, realizing out-of-order pipeline execution of static mapping.
It improves the processing efficiency of the accelerator, reduces memory overhead, and can perform out-of-order pipeline execution independently of the overall workload graph, dynamically adjusting the execution order of the workload.
Smart Images

Figure CN114895965B_ABST
Abstract
Description
[0001] Divisional Application Instructions
[0002] This application is a divisional application of the Chinese invention patent application with application date of June 18, 2020, application number 202010559855.3, and titled "Method and device for implementing out-of-order pipeline execution of static mapping of workloads". Technical Field
[0003] The present disclosure relates generally to processing, and more particularly, to methods and apparatus for implementing statically mapped out-of-order pipeline execution of workloads. Background Art
[0004] Computer hardware manufacturers develop hardware components for use in various components of a computer platform. For example, computer hardware manufacturers develop motherboards, chipsets for motherboards, central processing units (CPUs), hard disk drives (HDDs), solid-state drives (SSDs), and other computer components. In addition, computer hardware manufacturers develop processing elements called accelerators to speed up the processing of workloads. For example, an accelerator can be a CPU, a graphics processing unit (GPU), a vision processing unit (VPU), and / or a field programmable gate array (FPGA). Summary of the Invention
[0005] According to an embodiment of the present disclosure, a device for implementing statically mapped out-of-order pipeline execution of a workload is provided, the device comprising: an interface for loading a first number of credits into a memory; a comparator for comparing the first number of credits with a threshold number of credits, wherein the threshold number of credits is associated with memory availability in a buffer; and a dispatcher for selecting a workload node in the workload to be executed at a first computing building block in one or more computing building blocks when the first number of credits meets the threshold number of credits.
[0006] According to an embodiment of the present disclosure, a computer-readable storage medium is provided, comprising instructions that, when executed, cause at least one processor to perform at least the following operations: load a first number of credits into a memory; compare the first number of credits with a threshold number of credits, wherein the threshold number of credits is associated with memory availability in a buffer; and when the first number of credits satisfies the threshold number of credits, select a workload node in the workload to be executed at a compute building block.
[0007] According to an embodiment of the present disclosure, a device for implementing out-of-order pipeline execution of static mapping of a workload is provided, comprising: an interface connection device, the interface connection device being used to perform an interface connection to load a first number of credits into a memory; a comparison device, the comparison device being used to compare the first number of credits with a threshold number of credits, wherein the threshold number of credits is associated with the memory availability in a buffer; and a dispatch device, the dispatch device being used to select a workload node in the workload to be executed at a first computing building block in one or more computing building blocks when the first number of credits meets the threshold number of credits.
[0008] According to an embodiment of the present disclosure, a method for implementing statically mapped out-of-order pipeline execution of a workload is provided, comprising: loading a first number of credits into a memory; comparing the first number of credits with a threshold number of credits, wherein the threshold number of credits is associated with memory availability in a buffer; and selecting a workload node in the workload to be executed at a first computing building block in one or more computing building blocks when the first number of credits satisfies the threshold number of credits. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] Figure 1 is a graphical illustration representing a graph of a workload executed on an accelerator of a heterogeneous system.
[0010] Figure 2 is a graphical illustration representing a graph of a workload executed on an accelerator of a heterogeneous system implementing pipelines and buffers.
[0011] Figure 3 is a block diagram illustrating an example computing system constructed in accordance with the teachings of the present disclosure.
[0012] Figure 4 is a block diagram illustrating an example computing system including an example one or more schedulers.
[0013] Figure 5 It is achievable Figure 3 and Figure 4 A block diagram of an example scheduler for one or more of the schedulers.
[0014] Figure 6 It shows Figure 5 A block diagram of an example scheduler of the buffer credit storage device further details.
[0015] Figure 7 is a graphical illustration of an example graph representing a workload executed on an accelerator of a heterogeneous system implementing pipelines and buffers.
[0016] Figure 8is a flowchart representing a process that can be implemented by machine-readable instructions that can be executed to implement Figure 5 Scheduler and / or Figure 6 The scheduler.
[0017] Figure 9 is a block diagram of an example processor platform configured to perform Figure 8 Instructions to achieve Figure 5 Scheduler and / or Figure 6 One or more instances of a scheduler.
[0018] The drawings are not to scale. Generally, the same reference numerals will be used throughout the various drawings and the accompanying written description to refer to identical or similar components. Connection references (e.g., attach, couple, connect, and engage) should be interpreted broadly and, unless otherwise indicated, may include intermediate members between sets of elements and relative movement between elements. Therefore, a connection reference does not necessarily mean that two elements are directly connected and in a fixed relationship to each other.
[0019] The descriptors "first," "second," "third," and the like are used herein to identify multiple elements or components that can be referred to separately. Unless otherwise specified or understood based on the context of their use, such descriptors are not intended to convey any meaning of priority, physical order, arrangement in a list, or temporal ordering, but are simply used as labels for separately referring to multiple elements or components to facilitate understanding of the disclosed examples. In some examples, the descriptor "first" may be used to refer to an element in a specific embodiment, while the same element may be referred to by a different descriptor, such as "second" or "third," in the claims. In such cases, it should be understood that such descriptors are used solely for ease of reference to multiple elements or components. DETAILED DESCRIPTION
[0020] Many computer hardware manufacturers develop processing elements called accelerators to speed up the processing of workloads. For example, an accelerator can be a central processing unit (CPU), a graphics processing unit (GPU), a visual processing unit (VPU), and / or a field programmable gate array (FPGA). In addition, although accelerators can handle any type of workload, they are designed to optimize specific types of workloads. For example, while CPUs and FPGAs can be designed to handle more general processing, GPUs can be designed to improve the processing of videos, games, and / or other physics-based and mathematical calculations, and VPUs can be designed to improve the processing of machine vision tasks.
[0021] In addition, some accelerators are specifically designed to improve the processing of artificial intelligence (AI) applications. Although VPU is a specific type of AI accelerator, many different AI accelerators can be used. In fact, many AI accelerators can be implemented by application-specific integrated circuits (ASICs). Such ASIC-based AI accelerators can be designed to improve the processing of tasks related to specific types of AI, such as machine learning (ML), deep learning (DL), and / or other artificial machine-driven logic (including support vector machines (SVM), neural networks (NN), recursive neural networks (RNN), convolutional neural networks (CNN), long short-term memory (LSTM), gated recursive units (GRU)).
[0022] Computer hardware manufacturers are also developing heterogeneous systems that include more than one type of processing element. For example, a computer hardware manufacturer may combine a general-purpose processing element (e.g., a CPU) with a general-purpose accelerator (e.g., an FPGA) and / or a more specialized accelerator (e.g., a GPU, a VPU, and / or other AI accelerators). Such a heterogeneous system can be implemented as a system on a chip (SoC).
[0023] When the developer expects to run a function, algorithm, program, application, and / or other code on a heterogeneous system, the developer and / or software generates a schedule for the function, algorithm, program, application, and / or other code when compiling. Once a schedule is generated, the schedule is combined with the function, algorithm, program, application, and / or other code specifications to generate an executable file (in advance (Ahead of Time) or real-time (Just in Time) paradigm). In addition, the function, algorithm, program, application, and / or other code can be represented as a graph comprising nodes, wherein the graph represents a workload, and each node represents a specific task of the workload. In addition, the connection between the different nodes in the graph represents the data input and / or output required for executing a specific node, and the vertices of the graph represent the data dependencies between the nodes of the graph.
[0024] An executable file includes a number of different executable parts, each of which can be executed by a specific processing element (e.g., a CPU, GPU, VPU, and / or FPGA). Each executable part of the executable file can also include an executable sub-part, each of which can be executed by a computational building block (CBB) of a specific processing element. Additionally or alternatively, in some examples disclosed herein, developers and / or software development software can define criteria (e.g., success criteria) for determining the successful execution of an executable file. For example, such success criteria can correspond to executing the executable file to meet and / or otherwise reach a utilization threshold of a heterogeneous system and / or a specific processing element. In other examples, the success criteria can correspond to executing the executable file within a threshold amount of time. However, when determining how to execute an executable file on a heterogeneous system and / or a specific processing element, any suitable success function can be utilized. In this way, the success criteria can be beneficial for developers, software, and / or artificial intelligence systems to generate an executable file that includes a schedule optimized to meet the success criteria.
[0025] Figure 1 is a graphical illustration of a graph 100 representing a workload executing on an accelerator of a heterogeneous system. The graph 100 includes a first workload node 102 (WN[0]), a second workload node 104 (WN[1]), a third workload node 106 (WN[2]), a fourth workload node 108 (WN[3]), and a fifth workload node 110 (WN[4]). Figure 1 In FIG, the accelerator is running the workload represented by graph 100 through static software scheduling. Static software scheduling includes determining a predefined manner for executing different workload nodes of graph 100 on the compute building blocks (CBBs) of the accelerator. For example, the static software scheduling assigns a first workload node 102 (WN[0]) to a first CBB 112, a second workload node 104 (WN[1]) to a second CBB 114, a third workload node 106 (WN[2]) to a third CBB 116, a fourth workload node 108 (WN[3]) to a fourth CBB 118, and a fifth workload node 110 (WN[4]) to a second CBB 114.
[0026] exist Figure 1 In FIG, the static software schedule outlines that the first workload node 102 (WN[0]) will be executed on the first CBB 112 in parallel with the execution of the fourth workload node 108 (WN[3]) on the fourth CBB 118. Figure 1, the fourth CBB 118 executes the fourth workload node 108 (WN[3]) faster than the first CBB 112 executes the first workload node 102 (WN[0]). Because the static software schedule depicts that the second CBB 114 will execute the second workload node 104 (WN[1]) before the second CBB 114 executes the fifth workload node 110 (WN[4]), the second CBB 114 is idle until the first CBB 112 completes execution of the first workload node 102 (WN[0]). In addition, waiting for a workload node to be fully executed before executing a subsequent workload node requires significant memory overhead because the data generated by the CBB executing the first workload node (e.g., the first workload node 102 (WN[0])) needs to be stored on the accelerator before the CBB can execute the second workload node (e.g., the second workload node 104 (WN[1])).
[0027] Figure 2 is a graphical illustration of a graph 200 representing a workload executed on an accelerator of a heterogeneous system implementing pipelines and buffers. The graph 200 includes a first workload node 102 (WN[0]), a second workload node 104 (WN[1]), a third workload node 106 (WN[2]), a fourth workload node 108 (WN[3]), and a fifth workload node 110 (WN[4]). Figure 2 In FIG. 2 , the accelerator runs the workload represented by graph 200 through static software scheduling. Figure 2 The static software schedule depicts an execution schedule for different workload nodes of the graph 200 on the CBBs of the accelerator, which implements a pipeline and includes a first buffer 202, a second buffer 204, and a third buffer 206. Furthermore, the static software schedule assigns the first workload node 102 (WN[0]) to the first CBB 112, the second workload node 104 (WN[1]) to the second CBB 114, the third workload node 106 (WN[2]) to the third CBB 116, the fourth workload node 108 (WN[3]) to the fourth CBB 118, and the fifth workload node 110 (WN[4]) to the second CBB 114. The first buffer 202 is coupled to the first CBB 112 and the second CBB 114, the second buffer 204 is coupled to the second CBB 114 and the third CBB 116, and the third buffer 206 is coupled to the fourth CBB 118 and the second CBB 114.
[0028] Buffers 202, 204, and 206 allow static software scheduling to depict that each CBB will process a portion (e.g., a tile) of a workload node within a time interval, rather than executing the entire workload node within that time interval. Similarly, static software scheduling can depict that when portions (e.g., tiles) of the workload are available, CBBs processing data generated by other CBBs (e.g., consumers) can execute those portions of the workload node. However, because the CBBs executing the workload nodes process available data and write new data to memory, in order to execute a given workload node on the CBB, a threshold amount of data must be available at runtime, and a threshold amount of space must be available in memory to write the results at runtime. While buffers reduce memory overhead over basic static software scheduling, static software scheduling using buffers is increasingly difficult to depict because static software scheduling is highly dependent on data availability and / or dependencies at runtime. Furthermore, because the load of the entire accelerator can affect the processing speed of each CBB on the accelerator, it is difficult to develop a static software schedule that effectively utilizes the CBBs of a given accelerator.
[0029] Examples disclosed herein include methods and apparatus for implementing statically mapped out-of-order pipeline execution of workloads. In contrast to static software scheduling, the examples disclosed herein do not rely on a predetermined static software schedule. Instead, the examples disclosed herein determine which workload nodes assigned to a given CBB to run based on the available data and available memory on the accelerator and / or other processing elements. In addition, each CBB tracks the amount of data available in a first buffer (expressed as a first number of credits) associated with a given workload, and tracks the amount of space available in a second buffer (expressed as a second number of credits). This allows for dynamic runtime scheduling of workload nodes on a given CBB.
[0030] For each workload node, when a first number of credits satisfies a first threshold and a second number of credits satisfies a second threshold, the CBB may execute the workload node. This allows for out-of-order pipeline execution independent of a given overall workload graph. Examples disclosed herein provide an apparatus for implementing out-of-order pipeline execution of a static mapping of a workload to one or more compute building blocks of an accelerator. The example apparatus includes an interface for loading a first number of credits into a memory; a comparator for comparing the first number of credits to a threshold number of credits, the threshold number of credits being associated with memory availability in a buffer; and a dispatcher for selecting a workload node in the workload to execute on a first compute building block of one or more compute building blocks when the first number of credits satisfies the threshold number of credits.
[0031] Figure 3 is a block diagram illustrating an example computing system 300 constructed in accordance with the teachings of the present invention. Figure 3 In the example of FIG, computing system 300 includes example system memory 302 and example heterogeneous system 304. Example heterogeneous system 304 includes example host processor 306, example first communication bus 308, example first accelerator 310a, example second accelerator 310b, and example third accelerator 310c. Each of example first accelerator 310a, example second accelerator 310b, and example third accelerator 310c includes various CBBs, some of which are common to the operation of the accelerator and some of which are specific to the operation of the corresponding accelerator.
[0032] exist Figure 3 In the example of , system memory 302 is coupled to heterogeneous system 304. System memory 302 is a memory. Figure 3 In FIG, the system memory 302 is a shared memory device between the host processor 306, the first accelerator 310a, the second accelerator 310b, and the third accelerator 310c. Figure 3 In the example of , system memory 302 is a physical storage device local to computing system 300; however, in other examples, system memory 302 can be external to computing system 300 and / or otherwise remote from computing system 300. In yet other examples, system memory 302 can be a virtual storage device. Figure 3 In some examples, system memory 302 is a permanent storage device (e.g., read-only memory (ROM), programmable ROM (PROM), erasable PROM (EPROM), electrically erasable PROM (EEPROM), etc.). In other examples, system memory 302 may be a permanent basic input / output system (BIOS) or flash memory. In yet other examples, system memory 302 may be a volatile memory.
[0033] exist Figure 3 In FIG, a heterogeneous system 304 is coupled to a system memory 302. Figure 3 In the example of , the heterogeneous system 304 processes the workload by executing the workload on the host processor 306 and / or one or more of the first accelerator 310a, the second accelerator 310b, or the third accelerator 310c. Figure 3 In the embodiment of the present invention, heterogeneous system 304 is a SoC. Alternatively, heterogeneous system 304 can be any other type of computing or hardware system.
[0034] exist Figure 3In the example of , host processor 306 is a processing element that executes instructions (e.g., machine-readable instructions) to perform, run, and / or facilitate completion of operations associated with a computer or computing device (e.g., computing system 300). Figure 3 In the example of FIG, host processor 306 is the primary processing element of heterogeneous system 304 and includes at least one core. Alternatively, host processor 306 may be a co-primary processing element (e.g., in examples using more than one CPU), while in other examples, host processor 306 may be a secondary processing element.
[0035] exist Figure 3 In the illustrated example of FIG, one or more of the first accelerator 310 a, the second accelerator 310 b, and / or the third accelerator 310 c are processing elements that can be used for computing tasks (e.g., hardware acceleration) by programs executing on the heterogeneous system 304. For example, the first accelerator 310 a is a processing element that includes processing resources designed and / or otherwise configured or constructed to improve processing speed and overall performance of machine vision tasks that process AI (e.g., a VPU).
[0036] In the examples disclosed herein, each of the host processor 306, the first accelerator 310a, the second accelerator 310b, and the third accelerator 310c communicates with other elements of the computing system 300 and / or the system memory 302. For example, the host processor 306, the first accelerator 310a, the second accelerator 310b, the third accelerator 310c, and / or the system memory 302 communicate via a first communication bus 308. In some examples disclosed herein, the host processor 306, the first accelerator 310a, the second accelerator 310b, the third accelerator 310c, and / or the system memory 302 can communicate via any suitable wired and / or wireless communication system. Furthermore, in some examples disclosed herein, each of the host processor 306, the first accelerator 310a, the second accelerator 310b, the third accelerator 310c, and / or the system memory 302 can communicate with any component external to the computing system 300 via any suitable wired and / or wireless communication system.
[0037] exist Figure 3In the example of FIG. 3 , the first accelerator 310a includes an example convolution engine 312, an example RNN engine 314, an example memory 316, an example memory management unit (MMU) 318, an example DSP 320, an example controller 322, and an example direct memory access (DMA) unit 324. In addition, each of the example convolution engine 312, the example RNN engine 314, the example DMA unit 324, the example DSP 320, and the example controller 322 includes an example first scheduler 326, an example second scheduler 328, an example third scheduler 330, an example fourth scheduler 332, and an example fifth scheduler 334, respectively. Each of the example DSP 320 and the example controller 322 additionally includes an example first kernel library 336 and an example second kernel library 338.
[0038] exist Figure 3 In the example of , the convolution engine 312 is a device configured to improve the processing of tasks associated with convolution. In addition, the convolution engine 312 improves the processing of tasks associated with analysis of visual images and / or other tasks associated with CNN. Figure 3 In the embodiment of the present invention, the RNN engine 314 is a device configured to improve the processing of tasks associated with the RNN. In addition, the RNN engine 314 improves the processing of tasks associated with analysis of unsegmented, connected handwriting recognition, speech recognition, and / or other tasks associated with the RNN.
[0039] exist Figure 3 In the example of , the memory 316 is a shared memory device between at least one of the convolution engine 312, the RNN engine 314, the MMU 318, the DSP 320, the controller 322, and the DMA unit 324. Figure 3 In the example of , the memory 316 is a physical memory device local to the first accelerator 310a; however, in other examples, the memory 316 may be external to the first accelerator 310a and / or otherwise remote from the first accelerator 310a. In other examples, the memory 316 may be a virtual memory device. Figure 3 In some examples, memory 316 is a permanent storage device (e.g., ROM, PROM, EPROM, EEPROM, etc.). In other examples, memory 316 may be a permanent BIOS or flash memory. In other examples, memory 316 may be a volatile memory.
[0040] exist Figure 3In the illustrated example of , the example MMU 318 is a device that includes references to addresses of the memory 316 and / or the system memory 302. The MMU 318 additionally translates virtual memory addresses used by one or more of the convolution engine 312, the RNN engine 314, the DSP 320, and / or the controller 322 into physical addresses in the memory 316 and / or the system memory 302.
[0041] exist Figure 3 In the example of , DSP 320 is a device that improves the processing of digital signals. For example, DSP 320 facilitates processing to measure, filter, and / or compress continuous real-world signals (e.g., data from cameras and / or other sensors related to computer vision). Figure 3 In the embodiment of the present invention, the controller 322 is implemented as a control unit of the first accelerator 310a. For example, the controller 322 directs the operation of the first accelerator 310a. In some examples, the controller 322 implements a credit manager. In addition, the controller 322 can instruct one or more of the convolution engine 312, the RNN engine 314, the memory 316, the MMU 318, and / or the DSP 320 on how to respond to machine-readable instructions received from the host processor 306.
[0042] exist Figure 3 In the example shown, the DMA unit 324 is a device that allows at least one of the convolution engine 312, the RNN engine 314, the DSP 320, and the controller 322 to access the system memory 302 independently of the host processor 306. For example, the DMA unit 324 can be implemented by one or more analog or digital circuits, logic circuits, programmable processor(s), programmable controller(s), graphics processing unit(s) (GPUs), digital signal processor(s) (DSPs), application specific integrated circuit(s) (ASICs), programmable logic device(s) (PLDs), and / or field programmable logic devices (FPLDs).
[0043] exist Figure 3In the example of FIG, each of the first scheduler 326, the second scheduler 328, the third scheduler 330, the fourth scheduler 332, and the fifth scheduler 334 is a device that determines when the convolution engine 312, the RNN engine 314, the DMA unit 324, the DSP 320, and the controller 322 are each executing a portion of a workload that has been offloaded and / or otherwise sent to the first accelerator 310a. In addition, each of the first kernel library 336 and the second kernel library 338 is a data structure that includes one or more kernels. For example, the kernels of the first kernel library 336 and the second kernel library 338 are routines compiled for high throughput on the DSP 320 and the controller 322, respectively. A kernel corresponds to, for example, an executable sub-portion of an executable file to be run on the computing system 300.
[0044] In the examples disclosed herein, each of the convolution engine 312, the RNN engine 314, the memory 316, the MMU 318, the DSP 320, the controller 322, and the DMA unit 324 communicates with other elements of the first accelerator 310a. For example, the convolution engine 312, the RNN engine 314, the memory 316, the MMU 318, the DSP 320, the controller 322, and the DMA unit 324 communicate via an example second communication bus 340. In some examples, the second communication bus 340 can be implemented by a configuration and control (CnC) infrastructure and a data fabric. In some examples disclosed herein, the convolution engine 312, the RNN engine 314, the memory 316, the MMU 318, the DSP 320, the controller 322, and the DMA unit 324 can communicate via any suitable wired and / or wireless communication system. Furthermore, in some examples disclosed herein, each of the convolution engine 312 , the RNN engine 314 , the memory 316 , the MMU 318 , the DSP 320 , the controller 322 , and the DMA unit 324 can communicate with any component external to the first accelerator 310 a via any suitable wired and / or wireless communication system.
[0045] As previously described, each of the example first accelerator 310a, the example second accelerator 310b, and the example third accelerator 310c includes various CBBs, some of which are common to the operations of the accelerators and some of which are specific to the operations of the corresponding accelerators. For example, each of the first accelerator 310a, the second accelerator 310b, and the third accelerator 310c includes a common CBB, such as a memory, an MMU, a controller, and a corresponding scheduler for each of the CBBs.
[0046] Although Figure 3In the example of FIG, the first accelerator 310a implements a VPU and includes a convolution engine 312, an RNN engine 314, and a DSP 320 (e.g., a CBB specific to the operation of the first accelerator 310a), but the second accelerator 310b and the third accelerator 310c may include additional or alternative CBBs specific to the operation of the second accelerator 310b and / or the third accelerator 310c. For example, if the second accelerator 310b implements a GPU, the CBBs specific to the operation of the second accelerator 310b may include a thread dispatcher, a graphics technology interface, and / or any other CBBs that are expected to improve the processing speed and overall performance of processing computer graphics and / or image processing. Furthermore, if the third accelerator 310c implements an FPGA, the CBBs specific to the operation of the third accelerator 310c may include one or more arithmetic logic units (ALUs) and / or any other CBBs that are expected to improve the processing speed and overall performance of processing general-purpose computations.
[0047] Although Figure 3 The heterogeneous system 304 includes a host processor 306, a first accelerator 310a, a second accelerator 310b, and a third accelerator 310c, but in some examples, the heterogeneous system 304 may include any number of processing elements (e.g., host processors and / or accelerators), including application-specific instruction set processors (ASIPs), physical processing units (PPUs), designated DSPs, image processors, coprocessors, floating point units, network processors, multi-core processors, and front-end processors.
[0048] In addition, although Figure 3 In the example of FIG3 , the convolution engine 312, the RNN engine 314, the memory 316, the MMU 318, the DSP 320, the controller 322, the DMA unit 324, the first scheduler 326, the second scheduler 328, the third scheduler 330, the fourth scheduler 332, the fifth scheduler 334, the first kernel library 336, and the second kernel library 338 are implemented on the first accelerator 310 a, but one or more of the convolution engine 312, the RNN engine 314, the memory 316, the MMU 318, the DSP 320, the controller 322, the DMA unit 324, the first scheduler 326, the second scheduler 328, the third scheduler 330, the fourth scheduler 332, the fifth scheduler 334, the first kernel library 336, and the second kernel library 338 may be implemented on the host processor 306, the second accelerator 310 b, and / or the third accelerator 310 c.
[0049] Figure 4 is a block diagram illustrating an example computing system 400 including one or more example schedulers. In some examples, computing system 400 may correspond to Figure 3The computing system 300. Figure 4 In the example of FIG, computing system 400 includes example input 402, example compiler 404, and example accelerator 406. In some examples, accelerator 406 may correspond to Figure 3 The first accelerator 310a. Figure 4 , input 402 is coupled to compiler 404. Input 402 is a workload to be executed by accelerator 406. In some examples, compiler 404 may correspond to Figure 3 host processor 306 and / or external devices.
[0050] exist Figure 4 In the example of , input 402 is, for example, a function, algorithm, program, application, and / or other code to be executed by accelerator 406. In some examples, input 402 is a graphical description of a function, algorithm, program, application, and / or other code. In additional or alternative examples, input 402 is a workload related to AI processing (e.g., deep learning and / or computer vision).
[0051] exist Figure 4 In the illustrated example of FIG, a compiler 404 is coupled to an input 402 and an accelerator 406. The compiler 404 receives the input 402 and compiles the input 402 into one or more executable files to be executed by the accelerator 406. For example, the compiler 404 is a graph compiler that receives the input 402 and assigns various workload nodes of a workload (e.g., the input 402) to various CBBs of the accelerator 406. In addition, the compiler 404 allocates memory for one or more buffers in the memory of the accelerator 406.
[0052] exist Figure 4 In the example of FIG4 , the accelerator 406 is coupled to the compiler 404 and includes an example credit manager 408, an example CnC infrastructure 410, an example data infrastructure 411, an example convolution engine 412, an example DMA unit 414, an example RNN engine 416, an example DSP 418, an example memory 420, and an example MMU 422. In addition, each of the example convolution engine 412, the example DMA unit 414, the example RNN engine 416, and the example DSP 418 includes an example first scheduler 424, an example second scheduler 426, an example third scheduler 428, and an example fourth scheduler 430, respectively. In addition, the example DSP 418 includes an example kernel library 432. In some examples, the first scheduler 424 may correspond to Figure 3 In an additional or alternative example, the second scheduler 426 may correspond to Figure 3 In another example, the third scheduler 428 may correspond to Figure 3In some examples, the fourth scheduler 430 may correspond to Figure 4 The fourth scheduler 332.
[0053] exist Figure 4 In the illustrated example, a credit manager 408 is coupled to the compiler 404 and the CnC infrastructure 410. The credit manager 408 is a device that manages credits associated with one or more of the convolution engine 412, the DMA unit 414, the RNN engine 416, and / or the DSP 418. In some examples, the credit manager 408 can be implemented by a controller as a credit manager controller. A credit represents the amount of space available in memory 420 for data associated with a workload node and / or the amount of space available in memory 420 for output from the workload node. For example, the credit manager 408 can partition the memory 420 into one or more buffers associated with each workload node for a given workload based on one or more executable files received from the compiler 404. If a workload node is configured to write data to a buffer, the workload node is a producer, while if a workload node is configured to read data from a buffer, the workload node is a consumer.
[0054] exist Figure 4 In the example of 406, credit manager 408 is additionally configured to send credits to and / or receive credits from one or more of convolution engine 412, DMA unit 414, RNN engine 416, and / or DSP 418. In some examples, credit manager 408 is implemented as a control unit of accelerator 406. For example, credit manager 408 can direct the operation of accelerator 406. Furthermore, credit manager 408 can instruct one or more of convolution engine 412, DMA unit 414, RNN engine 416, and / or DSP 418 how to respond to executable files and / or other machine-readable instructions received from compiler 404.
[0055] exist Figure 4 In the example of FIG4 , CnC infrastructure 410 is coupled to credit manager 408, convolution engine 412, DMA unit 414, RNN engine 416, and DSP 418. CnC infrastructure 410 is a network of electronic interconnects and at least one logic circuit that allows one or more of credit manager 408, convolution engine 412, DMA unit 414, RNN engine 416, and / or DSP 418 to transmit credits to and / or receive credits from one or more of credit manager 408, convolution engine 412, DMA unit 414, RNN engine 416, and / or DSP 418. In some examples, CnC infrastructure 410 may correspond to Figure 3 The second communication bus 340.
[0056] exist Figure 4 In the example of FIG4 , data infrastructure 411 is coupled to convolution engine 412, DMA unit 414, RNN engine 416, DSP 418, memory 420, and MMU 422. Data infrastructure 411 is a network of electronic interconnects and at least one logic circuit that allows one or more of credit manager 408, convolution engine 412, RNN engine 416, DSP 418, memory 420, and / or MMU 422 to transmit data to and / or receive data from one or more of credit manager 408, convolution engine 412, RNN engine 416, DSP 418, memory 420, and / or MMU 422. In some examples, data infrastructure 411 may correspond to Figure 3 The second communication bus 340.
[0057] exist Figure 4 In the example shown, convolution engine 412 is coupled to CnC infrastructure 410 and data infrastructure 411. Convolution engine 412 is a device configured to improve processing of tasks associated with convolution. In addition, convolution engine 412 improves processing of tasks associated with visual image analysis and / or other tasks associated with CNN. In some examples, convolution engine 412 may correspond to Figure 3 Convolution engine 312.
[0058] exist Figure 4 In the example shown, DMA unit 414 is coupled to CnC infrastructure 410 and data infrastructure 411. DMA unit 414 is a device that allows at least one of convolution engine 412, RNN engine 416, or DSP 418 to access memory (e.g., system memory 302) remote from accelerator 406 independent of the corresponding processor (e.g., host processor 306). In some examples, DMA unit 414 may correspond to Figure 3 DMA unit 324. For example, DMA unit 414 may be implemented by one or more analog or digital circuits, logic circuits, programmable processor(s), programmable controller(s), GPU(s), DSP(s), ASIC(s), PLD(s), and / or FPLD(s).
[0059] exist Figure 4In the example, RNN engine 416 is coupled to CnC infrastructure 410 and data infrastructure 411. RNN engine 416 is a device configured to improve the processing of tasks associated with RNN. In addition, RNN engine 416 improves the processing of tasks associated with analysis of unsegmented, connected handwriting recognition, speech recognition, and / or other tasks associated with RNN. In some examples, RNN engine 416 may correspond to Figure 3 RNN engine 314.
[0060] exist Figure 4 In the example of FIG, DSP 418 is coupled to CnC infrastructure 410 and data infrastructure 411. DSP 418 is a device that improves the processing of digital signals. For example, DSP 418 facilitates processing to measure, filter, and / or compress continuous real-world signals, such as data from cameras and / or other sensors related to computer vision. In some examples, DSP 418 may correspond to Figure 3 DSP 320.
[0061] exist Figure 4 In the example of , memory 420 is coupled to data infrastructure 411. Memory 420 is a shared memory device between at least one of convolution engine 412, DMA unit 414, RNN engine 416, and DSP 418. In some examples, memory 420 may correspond to Figure 3 Memory 420 may be divided into one or more buffers associated with one or more workload nodes of a workload associated with an executable file received by credit manager 408. Figure 4 In the example of , memory 420 is a physical memory device local to accelerator 406. However, in other examples, memory 420 may be external to accelerator 406 and / or otherwise remote therefrom. In yet other examples, memory 420 may be a virtual memory device. Figure 4 In some examples, memory 420 is a permanent memory (e.g., ROM, PROM, EPROM, EEPROM, etc.). In other examples, memory 420 may be a permanent BIOS or flash memory. In yet other examples, memory 420 may be a volatile memory.
[0062] exist Figure 4In the illustrated example of FIG4 , an example MMU 422 is coupled to the data infrastructure 411. The MMU 422 is a device that includes references to addresses of the memory 420 and / or memory remote from the accelerator 406. The MMU 422 additionally translates virtual memory addresses used by one or more of the convolution engine 412, the DMA unit 414, the RNN engine 416, and / or the DSP 418 into physical addresses in the memory 420 and / or memory remote from the accelerator 406. In some examples, the MMU 422 may correspond to Figure 3 MMU 318.
[0063] exist Figure 4 In the example of FIG, each of the first scheduler 424, the second scheduler 426, the third scheduler 428, and the fourth scheduler 430 is a device that determines when the convolution engine 412, the DMA unit 414, the RNN engine 416, and the DSP 418, respectively, execute a portion of a workload (e.g., a workload node) that has been allocated to the convolution engine 412, the DMA unit 414, the RNN engine 416, and the DSP 418, respectively, by the credit manager 408 and / or the additional CBB of the accelerator 406. Depending on the task and / or other operation of a given workload node, a workload node can be a producer or a consumer. A producer workload node produces data that is utilized by another workload node, while a consumer workload node consumes and / or otherwise processes data produced by another workload node.
[0064] exist Figure 4 In the illustrated example, kernel library 432 is a data structure that includes one or more kernels. In some examples, kernel library 432 may correspond to Figure 3 The first kernel library 336 of the kernel library 432 is a routine compiled for high throughput on the DSP 418. The kernel corresponds to an executable sub-part of an executable file to be run on the accelerator 406. Figure 4 In the example of FIG. 4 , the accelerator 406 implements a VPU and includes a credit manager 408, a CnC infrastructure 410, a data infrastructure 411, a convolution engine 412, a DMA unit 414, an RNN engine 416, a DSP 418, a memory 420, and an MMU 422. The accelerator 406 may include Figure 4 Additional CBBs or alternative CBBs to the CBBs shown in FIG.
[0065] exist Figure 4In the example of FIG. 4 , in operation, the first scheduler 424 loads credits corresponding to input buffers to and output buffers from the workload nodes for allocation to the workload nodes of the convolution engine 412. For example, an input buffer is a buffer from which the workload node is configured to read data, and an output buffer is a buffer to which the workload node is configured to write data. In some examples, the input buffer of a first workload node may be the output buffer of a second workload node. In addition, the first scheduler 424 receives and / or otherwise obtains credits from the credit manager 408.
[0066] exist Figure 4 In the example of FIG4 , in operation, the first scheduler 424 selects a workload node assigned to the convolution engine 412 and determines whether the first scheduler 424 has received a threshold amount of credits to operate on data stored in an input buffer destined for the selected workload node. For example, the first scheduler 424 compares the number of credits received from the producer workload node for the input buffer with the threshold amount of credits for the input buffer. If the first scheduler 424 has not received the threshold amount of credits, the first scheduler 424 repeats the process for another workload node assigned to the convolution engine 412.
[0067] exist Figure 4 In the example shown in FIG, in operation, if the first scheduler 424 has received a threshold number of credits to operate on data stored in an input buffer destined for a selected workload node, the first scheduler 424 determines whether the first scheduler 424 has received a threshold number of credits to write data to an output buffer for the selected workload node. For example, the first scheduler 424 compares the number of credits received from the consumer workload node for the output buffer with a threshold number of credits for the output buffer for the selected workload node. If the first scheduler 424 has not received the threshold number of credits, the first scheduler 424 repeats the process for another workload node assigned to the convolution engine 412. If the first scheduler 424 has received the threshold number of credits to write data to the output buffer, the first scheduler 424 indicates that the selected workload node is ready for execution. Subsequently, the first scheduler 424 repeats the process for additional workload nodes assigned to the convolution engine 412.
[0068] exist Figure 4In the example of FIG, in operation, after the workload nodes assigned to the convolution engine 412 have been analyzed, the first scheduler 424 schedules the workload nodes that are ready for execution. The first scheduler 424 then dispatches the workload nodes according to the schedule. After the dispatched workload nodes are executed by the convolution engine 412, the first scheduler 424 sends credits corresponding to the input buffer and / or output buffer to the credit manager 408. The first scheduler 424 determines whether there are additional workload nodes to be executed in the schedule. If there are additional workload nodes in the schedule, the first scheduler 424 causes the next workload node in the schedule to be executed on the convolution engine 412.
[0069] Figure 5 It is achievable Figure 3 and Figure 4 5. For example, scheduler 500 is an example implementation of the following scheduler: Figure 3 a first scheduler 326, a second scheduler 328, a third scheduler 330, a fourth scheduler 332, and / or a fifth scheduler 334; and / or Figure 4 a first scheduler 424, a second scheduler 426, a third scheduler 428, and / or a fourth scheduler 430; and / or Figure 6 Scheduler 600; and / or Figure 7 The first scheduler 722, the second scheduler 724, the third scheduler 726, and / or the fourth scheduler 728.
[0070] exist Figure 5 In the example of FIG5 , scheduler 500 includes an example workload interface 502, an example buffer credit storage device 504, an example credit comparator 506, an example workload node dispatcher 508, and an example communication bus 510. Scheduler 500 is a device that determines when a CBB associated with scheduler 500 executes a portion of a workload (e.g., a workload node) that has been assigned to the CBB associated with scheduler 500.
[0071] exist Figure 5In the illustrated example, workload interface 502 is a device that is configured to communicate with other devices external to scheduler 500, buffer credit storage device 504, credit comparator 506, and / or workload node dispatcher 508. For example, workload interface 502 can receive and / or otherwise obtain workload nodes to be executed by a CBB associated with scheduler 500. Additionally or alternatively, workload interface 502 can transmit and / or receive credits from other schedulers, other CBBs, and / or other devices. Furthermore, workload interface 502 can load and / or load out credits corresponding to input buffers to and / or output buffers from workload nodes into and / or out of buffer credit storage device 504.
[0072] In some examples, the example workload interface 502 implements example means for interfacing. The interfacing means is implemented by executable instructions, for example, at least Figure 8 The instructions implemented by blocks 802, 818, and 822 of . For example, Figure 8 The executable instructions of blocks 802, 818, and 822 may be executed on at least one processor, for example, Figure 9 910 and / or example accelerator 912. In other examples, the interfacing device is implemented by hardware logic, a hardware-implemented state machine, a logic circuit, and / or any other combination of hardware, software, and / or firmware.
[0073] exist Figure 5 In the example shown, the buffer credit storage device 504 is a shared storage device between at least one of the workload interface 502, the credit comparator 506, and / or the workload node dispatcher 508. The buffer credit storage device 504 is a physical storage device local to the scheduler 500; however, in other examples, the buffer credit storage device 504 may be located external to the scheduler 500 and / or otherwise remote therefrom. In other examples, the buffer credit storage device 504 may be a virtual storage device. Figure 5 In some examples, buffer credit storage device 504 is a permanent memory device (e.g., ROM, PROM, EPROM, EEPROM, etc.). In other examples, buffer credit storage device 504 may be a permanent BIOS or flash memory. In other examples, buffer credit storage device 504 may be a volatile memory.
[0074] exist Figure 5In the example of FIG5 , buffer credit storage device 504 is a memory associated with storing credits corresponding to input buffers to and / or output buffers from workload nodes associated with workload nodes assigned to a CBB associated with scheduler 500. For example, buffer credit storage device 504 can be implemented as a data structure that includes a field for each workload node assigned to a CBB associated with scheduler 500 and a field for each input buffer to and / or output buffer from a workload node associated with a workload node assigned to a CBB associated with scheduler 500.
[0075] exist Figure 5 In the illustrated example of FIG5 , buffer credit storage device 504 may additionally or alternatively store a threshold amount of credits that have been allocated to workload nodes of a CBB associated with scheduler 500 and / or corresponding to input buffers to and / or output buffers from the workload nodes. Furthermore, buffer credit storage device 504 includes fields associated with a threshold amount of credits for input buffers to and / or output buffers from each workload node.
[0076] exist Figure 5 In the example of FIG, when the workload node is a producer (e.g., the workload node generates data for use by another workload node), the threshold number of credits corresponds to a threshold amount of space in the output buffer (e.g., the space allocated in memory 420), wherein the threshold amount of space must be met before the CBB associated with scheduler 500 can execute the producer workload node. Additionally, when the workload node is a consumer (e.g., the workload node processes data generated by another workload node), the threshold number of credits corresponds to a threshold amount of data in the input buffer (e.g., the space allocated in memory 420), wherein the threshold amount of data must be met before the CBB associated with scheduler 500 can execute the consumer workload node.
[0077] In some examples, the example buffer credit storage device 504 implements an example storage device. The storage device may be configured as follows: Figure 8 For example, the executable instructions may be executed on at least one processor, for example, Figure 9 910 and / or the example accelerator 912 shown in the example of FIG. In other examples, the storage device is implemented by hardware logic, a hardware-implemented state machine, a logic circuit, and / or any other combination of hardware, software, and / or firmware.
[0078] exist Figure 5In the example shown, credit comparator 506 is a device configured to determine whether a threshold number of credits has been received corresponding to input buffers to and / or output buffers from workload nodes allocated to a CBB associated with scheduler 500. Credit comparator 506 is configured to select a workload node allocated to a CBB associated with scheduler 500.
[0079] exist Figure 5 In the example of FIG5 , the credit comparator 506 is further configured to determine whether the scheduler 500 has received a threshold amount of credits to operate on the data stored in the input buffer for the selected workload node. For example, the credit comparator 506 compares a field in the buffer credit storage device 504 associated with the number of credits received from an external device (e.g., the credit manager 408, the controller 322, etc.) with a field in the buffer credit storage device 504 associated with the threshold amount of credits for the input buffer to the selected workload node. If the scheduler 500 has not received the threshold amount of credits, the credit comparator 506 repeats the process for another workload node assigned to the CBB associated with the scheduler 500.
[0080] exist Figure 5 In the example shown in , if the scheduler 500 has received a threshold amount of credits to operate on data stored in the input buffer, the credit comparator 506 determines whether the scheduler 500 has received a threshold amount of credits to write the data to the output buffer for the selected workload node. For example, the credit comparator 506 compares a field in the buffer credit storage device 504 associated with the number of credits for the output buffer for the selected workload node received from an external device (e.g., the credit manager 408, the controller 322, etc.) with a field in the buffer credit storage device 504 associated with the threshold amount of credits for the output buffer.
[0081] exist Figure 5 In the example shown in FIG5 , if scheduler 500 has not received the threshold amount of credits, credit comparator 506 repeats the process for another workload node assigned to the CBB associated with scheduler 500. If scheduler 500 has received the threshold amount of credits to write data to the output buffer, credit comparator 506 indicates that the selected workload node is ready for execution. Subsequently, credit comparator 506 repeats the process for additional workload nodes assigned to the CBB associated with scheduler 500.
[0082] In some examples, the example credit comparator 506 implements an example comparison device. The comparison device is implemented by executable instructions, for example, at least Figure 8804, 806, 808, 810, and 812. Figure 8 The executable instructions of blocks 804, 806, 808, 810, and 812 may be executed on at least one processor, e.g., Figure 9 The example processor 910 and / or the example accelerator 912 shown in the example of FIG. In other examples, the comparison device is implemented by hardware logic, a hardware-implemented state machine, a logic circuit, and / or any other combination of hardware, software, and / or firmware.
[0083] exist Figure 5 In an example of the embodiment of the present invention, workload node dispatcher 508 is a device that schedules one or more workload nodes assigned to a CBB associated with scheduler 500 to be executed on the CBB associated with scheduler 500. For example, after the workload nodes assigned to the CBB associated with scheduler 500 have been analyzed, workload node dispatcher 508 schedules the workload nodes that are ready for execution. For example, workload node dispatcher 508 schedules the workload nodes that are ready for execution based on a scheduling algorithm such as a round robin schedule. Workload node dispatcher 508 then dispatches the workload nodes according to the schedule. In other examples, workload node dispatcher 508 can utilize any other suitable arbitration algorithm to schedule the workload nodes that are ready for execution.
[0084] exist Figure 5 In the example shown in FIG, when the assigned workload node is executed by the CBB associated with the scheduler 500, the workload interface 502 sends credits associated with the input buffer to an external device (e.g., the credit manager 408, the controller 322, etc.), where the workload interface 502 receives the credits from the external device. The workload node dispatcher 508 additionally determines whether there are additional workload nodes to be executed in the schedule. If there are additional workload nodes in the schedule, the workload node dispatcher 508 dispatches the next workload node in the schedule.
[0085] In some examples, the example workload node dispatcher 508 implements an example dispatching device. The dispatching device is implemented by executable instructions, for example, at least Figure 8 814, 816, and 820. Figure 8 The executable instructions of blocks 814, 816, and 820 may be executed on at least one processor, for example, Figure 9The example processor 910 and / or the example accelerator 912 shown in the example of FIG. In other examples, the scheduling device is implemented by hardware logic, a hardware-implemented state machine, a logic circuit, and / or any other combination of hardware, software, and / or firmware.
[0086] In the examples disclosed herein, each of the workload interface 502, the buffer credit storage device 504, the credit comparator 506, and the workload node dispatcher 508 communicates with other elements of the scheduler 500. For example, the workload interface 502, the buffer credit storage device 504, the credit comparator 506, and the workload node dispatcher 508 communicate via an example communication bus 510. In some examples disclosed herein, the workload interface 502, the buffer credit storage device 504, the credit comparator 506, and the workload node dispatcher 508 can communicate via any suitable wired and / or wireless communication system. Additionally, in some examples disclosed herein, each of the workload interface 502, the buffer credit storage device 504, the credit comparator 506, and the workload node dispatcher 508 can communicate with any component external to the scheduler 500 via any suitable wired and / or wireless communication system.
[0087] Figure 6 It shows Figure 5 6. The scheduler 600 is an example implementation of the following scheduler: Figure 3 a first scheduler 326, a second scheduler 328, a third scheduler 330, a fourth scheduler 332, and / or a fifth scheduler 334; and / or Figure 4 a first scheduler 424, a second scheduler 426, a third scheduler 428 and / or a fourth scheduler 430; and / or Figure 5 Scheduler 500; and / or Figure 7 The first scheduler 722, the second scheduler 724, the third scheduler 726, and / or the fourth scheduler 728.
[0088] exist Figure 6 In the example of FIG6 , scheduler 600 includes an example workload interface 502, an example buffer credit storage device 504, an example credit comparator 506, and an example workload node dispatcher 508. Scheduler 600 is a device that determines when a CBB associated with scheduler 600 executes a portion of a workload (e.g., a workload node) that has been assigned to the CBB associated with scheduler 600.
[0089] exist Figure 6In the example shown, workload interface 502 is coupled to one or more devices external to scheduler 600, buffer credit storage 504, and workload node dispatcher 508. Workload interface 502 is a device that is configured to communicate with other devices external to scheduler 600, buffer credit storage 504, and / or workload node dispatcher 508. For example, workload interface 502 can receive and / or otherwise obtain workload nodes to be executed by a CBB associated with scheduler 600. Additionally or alternatively, workload interface 502 can send credits to and / or receive credits from one or more devices external to scheduler 600. Furthermore, workload interface 502 can load credits corresponding to input buffers of workload nodes and / or output buffers of workload nodes into and / or out of buffer credit storage 504.
[0090] exist Figure 6 In the example shown, the buffer credit storage device 504 is a shared storage device between at least one of the workload interface 502, the credit comparator 506, and / or the workload node dispatcher 508. The buffer credit storage device 504 is a physical storage device local to the scheduler 500. However, in other examples, the buffer credit storage device 504 can be located external to the scheduler 500 and / or otherwise remote from it. In other examples, the buffer credit storage device 504 can be a virtual storage device. Figure 6 In some examples, buffer credit storage device 504 is a permanent memory device (e.g., ROM, PROM, EPROM, EEPROM, etc.). In other examples, buffer credit storage device 504 may be a permanent BIOS or flash memory. In other examples, buffer credit storage device 504 may be a volatile memory.
[0091] exist Figure 6 In the example of , the buffer credit storage device 504 is a data structure that includes rows corresponding to a first workload node WN[0], a second workload node WN[1], and an nth workload node WN[n]. The buffer credit storage device 504 also includes columns corresponding to an input buffer for a first consumer (e.g., consumer[0]), an input buffer for a lth consumer (e.g., consumer[1]), an output buffer for a first producer (e.g., producer[0]), and an output buffer for an mth producer (e.g., producer[m]). The buffer credit storage device 504 also includes columns corresponding to a threshold number of credits for the input buffer to each workload node and / or the output buffer from each workload node.
[0092] exist Figure 6 In the example shown, each of the first workload node WN[0], the second workload node WN[1], and the nth workload node WN[n] is assigned to a CBB associated with the scheduler 600. In the buffer credit storage device 504, the intersection between the rows corresponding to the first workload node WN[0], the second workload node WN[1], and the nth workload node WN[n] and the columns corresponding to the input buffer for the first consumer (e.g., consumer[0]), the input buffer for the lth consumer (e.g., consumer[1]), the output buffer for the first producer (e.g., producer[0]), and the output buffer for the mth producer (e.g., producer[m]) represents a field corresponding to the number of credits received from one or more external devices for the buffer. In addition, the column corresponding to the threshold credit number for the input buffer to each workload node and / or the output buffer from each workload node represents such a threshold credit number: after the buffer meets the threshold credit number, the CBB associated with the scheduler 600 can operate on the corresponding workload node.
[0093] exist Figure 6 In the example of FIG, a field in the buffer credit storage device 504 at the intersection between the rows corresponding to the first workload node WN[0], the second workload node WN[1], and the nth workload node WN[n] and the columns corresponding to the input buffer for the first consumer (e.g., consumer[0]) and the input buffer for the lth consumer (e.g., consumer[1]) is initialized to a value of zero by an external device (e.g., credit manager 408, controller 322, etc.). Furthermore, a field in the buffer credit storage device 504 at the intersection between the rows corresponding to the first workload node WN[0], the second workload node WN[1], and the nth workload node WN[n] and the columns corresponding to the output buffer for the first producer (e.g., producer[0]) and the output buffer for the mth producer (e.g., producer[m]) is initialized to a value corresponding to the amount of memory divided in the associated buffer by an external device (e.g., credit manager 408, controller 322, etc.). Additionally, a column corresponding to a threshold number of credits for the input buffer and / or the output buffer is initialized by an external device (eg, credit manager 408, controller 322, software executing on host processor 306, etc.).
[0094] exist Figure 6In the illustrated example of FIG, a credit comparator 506 is coupled to the buffer credit storage device 504 and the workload node dispatcher 508. The credit comparator 506 is a device configured to determine whether a threshold number of credits has been received corresponding to input buffers to and / or output buffers from workload nodes allocated to a CBB associated with the scheduler 600. Figure 6 In the example of FIG6 , workload node dispatcher 508 is coupled to workload interface 502, buffer credit storage device 504, credit comparator 506, and one or more devices external to scheduler 600. Workload node dispatcher 508 is, for example, a device that schedules one or more workload nodes assigned to a CBB associated with scheduler 600 to execute on the CBB associated with scheduler 600.
[0095] exist Figure 6 In the example of FIG. 5 , in operation, when workload interface 502 receives and / or otherwise obtains a workload node from an external device (e.g., credit manager 408, controller 322, etc.), workload interface 502 loads the workload node into a corresponding field corresponding to the workload node in buffer credit storage 504. In addition, credit comparator 506 selects a workload node to be assigned to a CBB associated with scheduler 600.
[0096] exist Figure 6 In the illustrated example, credit comparator 506 determines whether scheduler 600 has received a threshold amount of credits to operate on data stored in an input buffer for a selected workload node. For example, credit comparator 506 compares a field in buffer credit storage device 504 associated with the number of credits received from an external device (e.g., credit manager 408, controller 322, etc.) with a field in buffer credit storage device 504 associated with a threshold amount of credits for the input buffer destined for the selected workload node. The threshold amount of credits corresponds to a threshold amount of data in the input buffer (e.g., a partitioned space in memory 420) that must be met before the CBB associated with scheduler 600 can execute the consumer workload node. If scheduler 600 has not received the threshold amount of credits, credit comparator 506 repeats the process for another workload node assigned to the CBB associated with scheduler 600.
[0097] exist Figure 6In the example shown in FIG, if the scheduler 600 has received a threshold amount of credits to operate on data stored in the input buffer, the credit comparator 506 determines whether the scheduler 600 has received a threshold amount of credits to write data to the output buffer for the selected workload node. For example, the credit comparator 506 compares a field in the buffer credit storage device 504 associated with the number of credits for the output buffer for the selected workload node received from an external device (e.g., the credit manager 408, the controller 322, etc.) with a field in the buffer credit storage device 504 associated with a threshold amount of credits for the output buffer. The threshold amount of credits may correspond to a threshold amount of space in the output buffer (e.g., space partitioned in memory), wherein the threshold amount of space must be met before the CBB associated with the scheduler 600 can execute the producer workload node.
[0098] exist Figure 6 In the example shown in FIG5 , if scheduler 600 has not received the threshold amount of credits, credit comparator 506 repeats the process for another workload node assigned to the CBB associated with scheduler 600. If scheduler 600 has received the threshold amount of credits to write data to the output buffer, credit comparator 506 indicates that the selected workload node is ready for execution. Subsequently, credit comparator 506 repeats the process for additional workload nodes assigned to the CBB associated with scheduler 600.
[0099] exist Figure 6 In the example of , workload node dispatcher 508 is a device that schedules one or more workload nodes assigned to a CBB associated with scheduler 600 to execute on the CBB associated with scheduler 600. For example, after the workload nodes assigned to the CBB associated with scheduler 600 have been analyzed, workload node dispatcher 508 schedules the workload nodes that are ready for execution. For example, workload node dispatcher 508 schedules the workload nodes that are ready for execution based on a scheduling algorithm such as round-robin scheduling. Workload node dispatcher 508 then dispatches the workload nodes according to the schedule. In other examples, workload node dispatcher 508 can utilize any other suitable arbitration algorithm to schedule the workload nodes that are ready for execution.
[0100] exist Figure 6In the example shown in FIG, when the assigned workload node is executed by the CBB associated with the scheduler 600, the workload interface 502 sends the credit associated with the input buffer to an external device (e.g., the credit manager 408, the controller 322, etc.), where the workload interface 502 receives the credit from the external device. The workload node dispatcher 508 additionally determines whether there are additional workload nodes to be executed in the schedule. If there are additional workload nodes in the schedule, the workload node dispatcher 508 dispatches the next workload node in the schedule.
[0101] Figure 7 is a graphical illustration of an example graph 700 representing a workload executed on an accelerator of a heterogeneous system implementing pipelines and buffers. For example, the accelerator is the first accelerator 310a and the heterogeneous system is Figure 3 The example graph 700 includes an example first workload node 702 (WN[0]), an example second workload node 704 (WN[1]), an example third workload node 706 (WN[2]), an example fourth workload node 708 (WN[3]), and an example fifth workload node 710 (WN[4]). Figure 7 In the example of FIG700 , the accelerator is configured to execute the workload represented by the graph 700 based on a schedule from an example credit manager 712 that assigns workload nodes to various CBBs. For example, the credit manager 712 and / or another controller assigns a first workload node 702 (WN[0]) to an example first CBB 714, a second workload node 704 (WN[1]) to an example second CBB 716, a third workload node 706 (WN[2]) to an example third CBB 718, a fourth workload node 708 (WN[3]) to an example fourth CBB 720, and a fifth workload node 710 (WN[4]) to an example second CBB 716.
[0102] exist Figure 7 In the example of FIG, each of the example first CBB 714, the example second CBB 716, the example third CBB 718, and the example fourth CBB 720 includes an example first scheduler 722, an example second scheduler 724, an example third scheduler 726, and an example fourth scheduler 728. Each of the first scheduler 722, the second scheduler 724, the third scheduler 726, and the fourth scheduler 728 can be implemented by Figure 5 The scheduler 500 and / or Figure 6 The scheduler 600 is implemented.
[0103] exist Figure 7In the illustrated example, a first workload node 702 (WN[0]) and a second workload node 704 (WN[1]) are associated with an example first buffer 730. The first buffer 730 is an output buffer for the first workload node 702 (WN[0]) and an input buffer for the second workload node 704 (WN[1]). The second workload node 704 (WN[1]) and the third workload node 706 (WN[2]) are associated with an example second buffer 732. The second buffer 732 is an output buffer for the second workload node 704 (WN[1]) and an input buffer for the third workload node 706 (WN[2]). The fourth workload node 708 (WN[3]) and the fifth workload node 710 (WN[4]) are associated with an example third buffer 734. The third buffer 734 is an output buffer for the fourth workload node 708 (WN[3]) and an input buffer to the fifth workload node 710 (WN[4]). Each of the first buffer 730, the second buffer 732, and the third buffer 734 can be implemented by a circular buffer. Figure 7 In the example of , each of the first buffer 730 , the second buffer 732 , and the third buffer 734 includes five partitions of the accelerator's memory, each of which can store data tiles.
[0104] exist Figure 7 In the example shown in , since the first workload node 702 (WN[0]) is a producer workload node, the credit manager 712 initializes the first scheduler 722 with five credits for the first buffer 730. Similarly, since the second workload node 704 (WN[1]) is a producer workload node, the credit manager 712 initializes the second scheduler 724 with five credits for the second buffer 732. Furthermore, since the fourth workload node 708 (WN[3]) is a producer workload node, the credit manager 712 initializes the fourth scheduler 728 with five credits for the third buffer 734.
[0105] The five credits provided to each of the first scheduler 722, the second scheduler 724, and the fourth scheduler 728 represent the sizes of the first buffer 730, the second buffer 732, and the third buffer 734. Furthermore, since the second workload node 704 (WN[1]) is also a consumer workload node, the credit manager 712 initializes the second scheduler 724 with zero credits for the first buffer 730. Furthermore, since the third workload node 706 (WN[2]) is a consumer workload node, the credit manager 712 initializes the third scheduler 726 with zero credits for the second buffer 732. Furthermore, since the fifth workload node 710 (WN[4]) is a consumer workload node, the credit manager 712 initializes the third buffer 734 with zero credits for the second scheduler 724.
[0106] exist Figure 7 In the example of FIG, because the first scheduler 722 has received a threshold number of credits for both the input buffer to the first workload node 702 (WN[0]) and the output buffer from the first workload node 702 (WN[0]), the first scheduler 722 dispatches the first workload node 702 (WN[0]) to execute on the first CBB 714. Additionally, because the fourth scheduler 728 has received a threshold number of credits for both the input buffer to the fourth workload node 708 (WN[3]) and the output buffer from the fourth workload node 708 (WN[3]), the fourth scheduler 728 dispatches the fourth workload node 708 (WN[3]) to execute on the fourth CBB 720. While the first workload node 702 (WN[0]) is executing on the first CBB 714, the first CBB 714 transmits data to the first buffer 730. Similarly, when the fourth workload node 708 (WN[3]) executes on the fourth CBB 720 , the fourth CBB 720 transfers data to the third buffer 734 .
[0107] exist Figure 7In the example shown, as each of the first CBB 714 and the fourth CBB 720 transmits data tiles associated with the first workload node 702 (WN[0]) and the fourth workload node 708 (WN[3]), respectively, the first scheduler 722 and the fourth scheduler 728 transmit credits to the credit manager 712 for each data tile transmitted from the first CBB 714 and the fourth CBB 720 to the first buffer 730 and the third buffer 734, respectively. The credit manager 712 transmits the credits received from the first scheduler 722 to the second scheduler 724, and transmits the credits received from the fourth scheduler 728 to the second scheduler 724. When the fourth CBB 720 executes the fourth workload node 708 (WN[3]), the fourth CBB 720 generates two data tiles to be stored in the third buffer 734. Similarly, when the first CBB 714 executes the first workload node 702 (WN[0]), the first CBB 714 generates five data tiles to store in the first buffer 730.
[0108] exist Figure 7 In the example shown in FIG5 , the fourth CBB 720 executes the fourth workload node 708 (WN[3]) faster than the first CBB 714 executes the first workload node 702 (WN[0]). Although there is available memory in the second buffer 732, the second scheduler 724 selects the fifth workload node 710 (WN[4]) (rather than the second workload node 704 (WN[1])) for execution on the second CBB 716 because the data on which the fifth workload node 710 (WN[4]) depends is ready before the data on which the second workload node 704 (WN[1]) depends is ready.
[0109] exist Figure 7In the illustrated example of , as the fifth workload node 710 (WN[4]) executes on the second CBB 716 and the second CBB 716 consumes data tiles stored in the third buffer 734, the second scheduler 724 sends a credit associated with the third buffer 734 back to the credit manager 712 for each data tile from the third buffer 734 consumed by the second CBB 716. Subsequently, if a threshold amount of credits for the first buffer 730 and the second buffer 732 is met, the second scheduler 724 dispatches the second workload node 704 (WN[1]) to execute on the second CBB 716. As the second CBB 716 generates data tiles associated with the second workload node 704 (WN[1]) and outputs the data to the second buffer 732, the second scheduler 724 sends a credit associated with the second buffer 732 to the credit manager 712 for each data tile transferred from the second CBB 716 to the second buffer 732.
[0110] exist Figure 7 In the example of FIG, upon receiving credits associated with the second buffer 732 from the second scheduler 724, the credit manager 712 sends the credits associated with the second buffer 732 to the third scheduler 726. When the third scheduler 726 receives a threshold amount of credits associated with the second buffer 732, the third scheduler 726 dispatches the third workload node 706 (WN[2]) to execute on the third CBB 718. As the third CBB 718 executes the third workload node 706 (WN[2]) and the third CBB 718 consumes data tiles stored in the second buffer 732, the third scheduler 726 sends a credit associated with the second buffer 732 back to the credit manager 712 for each data tile from the second buffer 732 consumed by the third CBB 718.
[0111] In additional or alternative examples, the first CBB 714 may correspond to Figure 4 The convolution engine 412, the first scheduler 722 may correspond to Figure 4 In some examples, the second CBB 716 may correspond to Figure 4 The RNN engine 416, the second scheduler 724 may correspond to Figure 4 In a further example, the third CBB 718 may correspond to Figure 4 DMA unit 414, the third scheduler 726 may correspond to Figure 4 In some examples, the fourth CBB 720 may correspond to Figure 4 DSP 418, the fourth scheduler 728 may correspond to Figure 4The fourth scheduler 430.
[0112] Although Figure 5 and / or Figure 6 An example way to implement the following scheduler is shown in: Figure 3 a first scheduler 326, a second scheduler 328, a third scheduler 330, a fourth scheduler 332, and / or a fifth scheduler 334; and / or Figure 4 a first scheduler 424, a second scheduler 426, a third scheduler 428, and / or a fourth scheduler 430; and / or Figure 7 The first scheduler 722, the second scheduler 724, the third scheduler 726, and / or the fourth scheduler 728, but Figure 5 and / or Figure 6 One or more of the elements, processes, and / or devices shown in the example may be combined, separated, rearranged, omitted, eliminated, and / or implemented in any other manner. In addition, the example workload interface 502, the example buffer credit storage device 504, the example credit comparator 506, the example workload node dispatcher 508, the example communication bus 510, and / or more generally, Figure 5 The example scheduler 500 and / or Figure 6 The example scheduler 600 can be implemented by hardware, software, firmware, and / or any combination of hardware, software, and / or firmware. Thus, for example, the example workload interface 502, the example buffer credit storage device 504, the example credit comparator 506, the example workload node dispatcher 508, the example communication bus 510, and / or more generally, Figure 5 The example scheduler 500 and / or Figure 6 Any of the example schedulers 600 may be implemented by one or more analog or digital circuits, logic circuits, programmable processor(s), programmable controller(s), graphics processing unit(s) (GPUs), digital signal processor(s) (DSPs), application specific integrated circuit(s) (ASICs), programmable logic device(s) (PLDs), and / or field programmable logic device(s) (FPLDs). When any apparatus or system claim of this patent is read to encompass pure software and / or firmware implementations, the example workload interface 502, the example buffer credit storage device 504, the example credit comparator 506, the example workload node dispatcher 508, the example communication bus 510, and / or more generally, Figure 5 The example scheduler 500 and / or Figure 6At least one of the example schedulers 600 is expressly defined herein as comprising a non-transitory computer-readable storage device or storage disk, such as a memory, a digital versatile disk (DVD), a compact disk (CD), a Blu-ray disk, etc. (including software and / or firmware). Furthermore, Figure 5 The example scheduler 500 and / or Figure 6 The example scheduler 600 may include Figure 5 and / or Figure 6 One or more elements, processes and / or devices in addition to or in place of the elements, processes and / or devices shown, and / or Figure 5 The example scheduler 500 and / or Figure 6 The example scheduler 600 may include more than one of any or all of the elements, processes, and devices shown. As used herein, the phrase "communicating" (including variations thereof) encompasses direct communication and / or indirect communication through one or more intermediate components, and does not require direct physical (e.g., wired) communication and / or continuous communication, but additionally includes selective communication at periodic intervals, scheduled intervals, non-periodic intervals, and / or one-time events.
[0113] Figure 8 shows a representation for implementing Figure 5 The scheduler 500 and / or Figure 6 Flowchart of example hardware logic, machine readable instructions, hardware implemented state machine, and / or any combination thereof of the scheduler 600. The machine readable instructions may be one or more executable programs or (one or more) portions of executable programs for use in conjunction with, for example, Figure 9 The program is executed by a computer processor such as the processor 910 and / or accelerator 912 shown in the example processor platform 900 discussed. The program may be embodied in software stored on a non-transitory computer-readable storage medium such as a CD-ROM, floppy disk, hard drive, DVD, Blu-ray disk, or memory associated with the processor 910 and / or accelerator 912, but the entire program and / or portions thereof may alternatively be executed by a device other than the processor 910 and / or accelerator 912 and / or embodied in firmware or dedicated hardware. In addition, although the example program is referenced Figure 8 The flowchart shown in FIG is described, but the exemplary Figure 5 The scheduler 500 and / or Figure 6Many other methods of the scheduler 600 are described. For example, the order of execution of the blocks can be changed, and / or some of the described blocks can be changed, eliminated, or combined. Additionally or alternatively, any or all of the blocks can be implemented by one or more hardware circuits (e.g., discrete and / or integrated analog and / or digital circuits, FPGAs, ASICs, comparators, operational amplifiers (op-amps), logic circuits, etc.) that are structured to perform the corresponding operations without executing software or firmware.
[0114] The machine-readable instructions described herein may be stored in one or more of a compressed format, an encrypted format, a segmented format, a compiled format, an executable format, an encapsulated format, and the like. The machine-readable instructions as described herein may be stored as data (e.g., portions of instructions, code, representations of code, and the like) that may be used to create, manufacture, and / or generate machine-executable instructions. For example, the machine-readable instructions may be segmented and stored on one or more storage devices and / or computing devices (e.g., servers). The machine-readable instructions may require one or more of the following operations: installation, modification, adaptation, updating, combination, supplementation, configuration, decryption, decompression, decapsulation, distribution, redistribution, compilation, and the like so that they may be directly read, interpreted, and / or executed by a computing device and / or other machine. For example, the machine-readable instructions may be stored in multiple parts that are individually compressed, encrypted, and stored on separate computing devices, wherein the parts, when decrypted, decompressed, and combined, form an executable instruction set that implements a program such as described herein.
[0115] In another example, the machine-readable instructions may be stored in a state in which they are readable by a computer, but require the addition of a library (e.g., a dynamic link library (DLL)), a software development kit (SDK), an application programming interface (API), etc., in order to execute the instructions on a particular computing device or other device. In another example, the machine-readable instructions and / or corresponding program may need to be configured (e.g., stored settings, input data, recorded network addresses, etc.) before the machine-readable instructions and / or corresponding program can be executed in whole or in part. Therefore, the disclosed machine-readable instructions and / or corresponding program(s) are intended to cover such machine-readable instructions and / or programs regardless of the particular format or state in which the machine-readable instructions and / or programs are stored or otherwise at rest or in transmission.
[0116] The machine-readable instructions described herein may be represented by any past, present, or future instruction language, scripting language, programming language, etc. For example, the machine-readable instructions may be represented using any of the following languages: C, C++, Java, C#, Perl, Python, JavaScript, Hypertext Markup Language (HTML), Structured Query Language (SQL), Swift, etc.
[0117] As described above, the present invention may be implemented using executable instructions (e.g., computer and / or machine readable instructions) stored on a non-transitory computer and / or machine readable medium. Figure 8 In an example process, the non-transitory computer and / or machine readable medium is, for example, a hard drive, flash memory, read-only memory, compact disc, digital versatile disc, cache, random access memory, and / or any other storage device or storage disk in which information is stored for any duration (e.g., stored for an extended period of time, stored permanently, stored temporarily, stored for temporary buffering, and / or caching information). As used herein, the term non-transitory computer readable medium is expressly defined to include any type of computer readable storage device and / or storage disk and to exclude propagating signals and to exclude transmission media.
[0118] "Include" and "comprising" (and all forms and tenses thereof) are used herein as open-ended terms. Thus, whenever a claim employs any form of "include" or "comprising" (e.g., includes, comprises, having, etc.) as a preamble or within any type of claim recitation, it is understood that additional elements, terms, etc. may be present without falling outside the scope of the corresponding claim or recitation. As used herein, the phrase "at least" when used as a transition term, such as in the preamble of a claim, is open-ended (in the same manner that the terms "include" and "comprising" are open-ended). The term "and / or" when used, for example, in the form of A, B, and / or C, refers to any combination or subset of A, B, and C, such as (1) A alone, (2) B alone, (3) C alone, (4) A and B, (5) A and C, (6) B and C, and (7) A and B and C. As used herein, in the context of describing structures, components, objects, and / or things, the phrase "at least one of A and B" is intended to refer to implementations that include any of the following: (1) at least one A, (2) at least one B, and (3) at least one A and at least one B. Similarly, as used herein, in the context of describing structures, components, objects, and / or things, the phrase "at least one of A or B" is intended to refer to implementations that include any of the following: (1) at least one A, (2) at least one B, and (3) at least one A and at least one B. As used herein, in the context of describing the execution or execution of processes, instructions, actions, activities, and / or steps, the phrase "at least one of A and B" is intended to refer to implementations that include any of the following: (1) at least one A, (2) at least one B, and (3) at least one A and at least one B. Similarly, as used herein, in the context of describing the execution or execution of a process, instruction, action, activity, and / or step, the phrase "at least one of A or B" is intended to refer to an implementation that includes any of the following: (1) at least one A, (2) at least one B, and (3) at least one A and at least one B.
[0119] As used herein, singular references (e.g., "a," "an," "first," "second," etc.) do not exclude a plurality. As used herein, the term "a" or "an" entity refers to one or more of the entity. The terms "a" (or "an"), "one or more," and "at least one" are used interchangeably herein. In addition, although listed separately, multiple devices, elements, or method actions can be implemented by, for example, a single unit or processor. In addition, although individual features can be included in different examples or claims, these features can be combined, and being included in different examples or claims does not mean that the combination of features is infeasible and / or disadvantageous.
[0120] Figure 8 is a flow chart representing a process 800 that may be implemented by machine-readable instructions that may be executed to implement Figure 5 The scheduler 500 and / or Figure 6 The process 800 begins at block 802 where the workload interface 502 loads credits corresponding to input buffers and / or output buffers of workload nodes assigned to a CBB associated with the scheduler 500 and / or the scheduler 600 into the buffer credit storage 504.
[0121] exist Figure 8 In the illustrated example of FIG, process 800 continues at block 804, where credit comparator 506 selects a workload node assigned to a CBB associated with scheduler 500 and / or scheduler 600. At block 806, credit comparator 506 determines whether scheduler 500 and / or scheduler 600 has received a threshold amount of credits to operate on data stored in an input buffer for the selected workload node. For example, credit comparator 506 compares a field in an array or other data structure associated with the number of credits received from an external device (e.g., credit manager 408, controller 322, etc.) to a field in an array or other data structure associated with a threshold amount of credits for the input buffer to the selected workload node. If credit comparator 506 determines that scheduler 500 and / or scheduler 600 has not received the threshold amount of credits to operate on data stored in the input buffer for the selected workload node (block 806: NO), process 800 proceeds to block 812.
[0122] exist Figure 8In the example of FIG, if credit comparator 506 determines that scheduler 500 and / or scheduler 600 has received the threshold amount of credits to operate on the data stored in the input buffer (block 806: YES), process 800 proceeds to block 808. At block 808, credit comparator 506 determines whether scheduler 500 and / or scheduler 600 has received the threshold amount of credits to write data to the output buffer for the selected workload node. For example, credit comparator 506 compares a field in an array or other data structure associated with the number of credits for the output buffer for the selected workload node received from an external device (e.g., credit manager 408, controller 322, etc.) to a field in an array or other data structure associated with the threshold amount of credits for the output buffer. If credit comparator 506 determines that scheduler 500 and / or scheduler 600 has not received the threshold amount of credits (block 808: NO), process 800 proceeds to block 812. If credit comparator 506 determines that scheduler 500 and / or scheduler 600 has received the threshold amount of credits to write data to the output buffer (block 808 : YES), credit comparator 506 indicates that the selected workload node is ready to execute at block 810 .
[0123] exist Figure 8 In the example of FIG, at block 812, the credit comparator 506 determines whether there are additional workload nodes to be processed. If the credit comparator 506 determines that there are additional workload nodes to be processed (block 812: Yes), the credit comparator 506 selects the additional workload node, and the process 800 proceeds to block 806. If the credit comparator 506 determines that there are no additional workload nodes to be processed (block 812: No), the process 800 proceeds to block 814.
[0124] exist Figure 8 In the illustrated example, at block 814, workload node dispatcher 508 schedules workload nodes that are ready for execution. At block 816, workload node dispatcher 508 dispatches the workload nodes according to the schedule. At block 818, when the dispatched workload nodes are executed by a CBB associated with scheduler 500 and / or scheduler 600, workload interface 502 sends credits associated with the input buffer to an external device (e.g., credit manager 408, controller 322, etc.), where workload interface 502 receives the credits from the external device.
[0125] exist Figure 8In the example shown, at block 820, the workload node dispatcher 508 determines whether there are additional workload nodes in the schedule to execute. If the workload node dispatcher 508 determines that there are additional workload nodes in the schedule (block 820: yes), the process 800 proceeds to block 816. If the workload node dispatcher 508 determines that there are no additional workload nodes in the schedule (block 820: no), the process 800 proceeds to block 822.
[0126] exist Figure 8 In the example of FIG. 8 , at block 822 , workload interface 502 determines whether to continue operation. For example, the condition that causes workload interface 502 to determine to continue operation may include receiving an additional workload node. If workload interface 502 determines to continue operation (block 822 : YES), process 800 proceeds to block 802 . If workload interface 502 determines not to continue operation (block 822 : NO), process 800 terminates.
[0127] Figure 9 is constructed to execute Figure 8 Instructions to achieve Figure 5 The scheduler 500 and / or Figure 6 The processor platform 900 can be, for example, a server, a personal computer, a workstation, a self-learning machine (e.g., a neural network), a mobile device (e.g., a cell phone, a smartphone, an iPad, etc.). TM ), a personal digital assistant (PDA), an Internet appliance, a DVD player, a CD player, a digital video recorder, a Blu-ray player, a game console, a personal video recorder, a set-top box, headphones, or other wearable device, or any other type of computing device.
[0128] The processor platform 900 of the illustrated example includes a processor 910 and an accelerator 912. The processor 910 of the illustrated example is hardware. For example, the processor 910 can be implemented by one or more integrated circuits, logic circuits, microprocessors, GPUs, DSPs, or controllers from any desired series or manufacturer. The hardware processor can be a semiconductor-based (e.g., silicon-based) device. In addition, the accelerator 912 can be implemented by, for example, one or more integrated circuits, logic circuits, microprocessors, GPUs, DSPs, FPGAs, VPUs, controllers, and / or other CBBs from any desired series or manufacturer. The accelerator 912 of the illustrated example is hardware. The hardware accelerator can be a semiconductor-based (e.g., silicon-based) device. In this example, the accelerator 912 implements the example convolution engine 312, the example RNN engine 314, the example memory 316, the example MMU 318, the example DSP 320, the example controller 322, and the example DMA unit 324. In addition, each of the example convolution engine 312, the example RNN engine 314, the example DMA unit 324, the example DSP 320, and the example controller 322 includes an example first scheduler 326, an example second scheduler 328, an example third scheduler 330, an example fourth scheduler 332, and an example fifth scheduler 334, respectively. Figure 9 In the example, each of the example first scheduler 326, the example second scheduler 328, the example third scheduler 330, the example fourth scheduler 332, and the example fifth scheduler 334 includes an example workload interface 502, an example buffer credit storage device 504, an example credit comparator 506, an example workload node dispatcher 508, and / or more generally, a scheduler 500.
[0129] In additional or alternative examples, the processor 910 implements the example convolution engine 312, the example RNN engine 314, the example memory 316, the example MMU 318, the example DSP 320, the example controller 322, and the example DMA unit 324. Furthermore, in such additional or alternative examples, each of the example convolution engine 312, the example RNN engine 314, the example DMA unit 324, the example DSP 320, and the example controller 322 includes an example first scheduler 326, an example second scheduler 328, an example third scheduler 330, an example fourth scheduler 332, and an example fifth scheduler 334, respectively. In such additional or alternative examples, each of the example first scheduler 326, the example second scheduler 328, the example third scheduler 330, the example fourth scheduler 332, and the example fifth scheduler 334 includes an example workload interface 502, an example buffer credit storage device 504, an example credit comparator 506, an example workload node dispatcher 508, and / or more generally, a scheduler 500.
[0130] The processor 910 of the illustrated example includes a local memory 911 (e.g., a cache). The processor 910 of the illustrated example communicates with a main memory including a volatile memory 914 and a non-volatile memory 916 via a bus 918. In addition, the accelerator 912 of the illustrated example includes a local memory 913 (e.g., a cache). The accelerator 912 of the illustrated example communicates with a main memory including a volatile memory 914 and a non-volatile memory 916 via a bus 918. The volatile memory 914 may be a synchronous dynamic random access memory (SDRAM), a dynamic random access memory (DRAM), Dynamic Random Access Memory The non-volatile memory 916 may be implemented by flash memory and / or any other desired type of memory device. Access to the main memories 914, 916 is controlled by a memory controller.
[0131] The processor platform 900 of the illustrated example also includes an interface circuit 920. The interface circuit 920 may be implemented by any type of interface standard, such as an Ethernet interface, a universal serial bus (USB), interface, near field communication (NFC) interface, and / or PCI-express interface.
[0132] In the example shown, one or more input devices 922 are connected to the interface circuit 920. The input device(s) 922 allow a user to input data and / or commands into the processor 910 and / or accelerometer 912. The input device(s) may be implemented by, for example, an audio sensor, a microphone, a camera (still or video), a keyboard, buttons, a mouse, a touch screen, a trackpad, a trackball, an isopoint, and / or a voice recognition system.
[0133] One or more output devices 924 are also connected to the interface circuit 920 of the illustrated example. The output device 924 can be implemented, for example, by a display device (e.g., a light-emitting diode (LED), an organic light-emitting diode (OLED), a liquid crystal display (LCD), a cathode ray tube display (CRT), an in-plane switching (IPS) display, a touch screen, etc.), a tactile output device, a printer, and / or a speaker. Therefore, the interface circuit 920 of the illustrated example typically includes a graphics driver card, a graphics driver chip, and / or a graphics driver processor.
[0134] The interface circuitry 920 in the illustrated example also includes communication devices such as transmitters, receivers, transceivers, modems, residential gateways, wireless access points, and / or network interfaces to facilitate exchanging data with external machines (e.g., any kind of computing device) over the network 926. Communication can occur through, for example, an Ethernet connection, a digital subscriber line (DSL) connection, a telephone line connection, a coaxial cable system, a satellite system, a field-line wireless system, a cellular telephone system, and the like.
[0135] The processor platform 900 of the illustrated example also includes one or more mass storage devices 928 for storing software and / or data. Examples of such mass storage devices 928 include floppy disk drives, hard drives, optical disk drives, Blu-ray disk drives, redundant array of independent disks (RAID) systems, and digital versatile disk (DVD) drives.
[0136] Figure 8 The machine-executable instructions 932 may be stored in the mass storage device 928, in the volatile memory 914, in the non-volatile memory 916, and / or on a removable, non-transitory computer-readable storage medium such as a CD or DVD.
[0137] In light of the foregoing, it will be understood that example methods, apparatus, and articles have been disclosed that implement out-of-order pipeline execution of static mappings of workloads. In addition, example methods, apparatus, and articles have been disclosed that allow computational building blocks to execute workload nodes when the data on which the workload nodes depend is available and there is sufficient available memory to store the output generated by executing the workload nodes. In addition, the examples disclosed herein allow workload nodes to be executed by computational building blocks that are assigned workload nodes independently of scheduling and / or other sequencing. The disclosed methods, apparatus, and articles improve the efficiency of using computing devices by increasing the utilization of processing devices. In addition, the example methods, apparatus, and articles as disclosed herein reduce the number of computational cycles used by the processing devices to process and / or otherwise execute workloads. The disclosed methods, apparatus, and articles are accordingly directed to (one or more) improvements in computer functionality.
[0138] Disclosed herein are example methods, apparatuses, systems, and articles of manufacture for implementing statically mapped out-of-order pipeline execution of a workload. Further examples and combinations thereof include the following: Example 1 includes an apparatus comprising: an interface for loading a first number of credits into a memory; a comparator for comparing the first number of credits to a threshold number of credits, wherein the threshold number of credits is associated with memory availability in a buffer; and a dispatcher for selecting a workload node to execute at a first compute building block in one or more compute building blocks in the workload when the first number of credits satisfies the threshold number of credits.
[0139] Example 2 includes the apparatus of Example 1, wherein the interface is used to: load the first number of credits into the memory when the interface receives the first number of credits from the credit manager; and as one or more data tiles associated with the workload node are transferred from a first of the one or more compute building blocks to the buffer, transfer credits to the credit manager for each tile transferred to the buffer.
[0140] Example 3 includes the apparatus of Example 1, wherein the buffer is an output buffer associated with the workload node, the first number of credits corresponds to the output buffer, and the threshold number of credits corresponds to a threshold amount of memory in the output buffer.
[0141] Example 4 includes the apparatus of Example 1, wherein the buffer is an input buffer associated with the workload node, the first number of credits corresponds to the input buffer, and the threshold number of credits corresponds to a threshold amount of data in the input buffer.
[0142] Example 5 includes the apparatus of Example 1, wherein the buffer is a first buffer, the threshold number of credits is a first threshold number of credits, the comparator is to: compare the second number of credits to a second threshold number of credits, wherein the second threshold number of credits is associated with memory availability in the second buffer, and the dispatcher is to: select a workload node to be executed at a first compute building block of the one or more compute building blocks when (1) the first number of credits satisfies the first threshold number of credits and (2) the second number of credits satisfies the second threshold number of credits.
[0143] Example 6 includes the apparatus of Example 5, wherein the second buffer is an input buffer associated with the workload node, the second number of credits corresponds to the input buffer, and the second threshold number of credits corresponds to a threshold amount of data in the input buffer.
[0144] Example 7 includes the apparatus of Example 1, wherein the threshold number of credits is a first threshold number of credits, the workload node is a first workload node, and the dispatcher is to: schedule the first workload node and the second workload node to execute at a first computing building block in the one or more computing building blocks when (1) the first number of credits satisfies the first threshold number of credits and (2) the second number of credits satisfies the second threshold number of credits.
[0145] Example 8 includes a non-transitory computer-readable storage medium comprising instructions that, when executed, cause at least one processor to at least: load a first number of credits into a memory; compare the first number of credits to a threshold number of credits, wherein the threshold number of credits is associated with memory availability in a buffer; and select a workload node in the workload to execute at a compute building block when the first number of credits satisfies the threshold number of credits.
[0146] Example 9 includes the non-transitory computer-readable storage medium of Example 8, wherein the instruction, when executed, causes at least one processor to perform the following operations: loading the first number of credits into a memory when a first number of credits is received from a credit manager; and transmitting credits to the credit manager for each tile transmitted to the buffer as one or more data tiles associated with a workload node are transmitted from a compute building block to a buffer.
[0147] Example 10 includes the non-transitory computer-readable storage medium of Example 8, wherein the buffer is an output buffer associated with the workload node, the first number of credits corresponds to the output buffer, and the threshold number of credits corresponds to a threshold amount of memory in the output buffer.
[0148] Example 11 includes the non-transitory computer-readable storage medium of Example 8, wherein the buffer is an input buffer associated with the workload node, the first number of credits corresponds to the input buffer, and the threshold number of credits corresponds to a threshold amount of data in the input buffer.
[0149] Example 12 includes the non-transitory computer-readable storage medium of Example 8, wherein the buffer is a first buffer, the threshold number of credits is a first threshold number of credits, and wherein the instruction, when executed, causes at least one processor to perform the following operations: compare the second number of credits to a second threshold number of credits, wherein the second threshold number of credits is associated with memory availability in the second buffer; and select a workload node to execute at the compute building block when (1) the first number of credits satisfies the first threshold number of credits and (2) the second number of credits satisfies the second threshold number of credits.
[0150] Example 13 includes the non-transitory computer-readable storage medium of Example 12, wherein the second buffer is an input buffer associated with the workload node, the second number of credits corresponds to the second buffer, and the second threshold number of credits corresponds to a threshold amount of data in the input buffer.
[0151] Example 14 includes the non-transitory computer-readable storage medium of Example 8, wherein the threshold number of credits is a first threshold number of credits, the workload node is a first workload node, and the instructions, when executed, cause at least one processor to perform the following operations: when (1) the first number of credits satisfies the first threshold number of credits and (2) the second number of credits satisfies the second threshold number of credits, scheduling the first workload node and the second workload node to execute at the compute building block.
[0152] Example 15 includes an apparatus comprising: an interfacing device for interfacing to load a first number of credits into a memory; a comparing device for comparing the first number of credits to a threshold number of credits, wherein the threshold number of credits is associated with memory availability in a buffer; and a dispatching device for selecting a workload node in the workload to be executed at a first compute building block in one or more compute building blocks when the first number of credits satisfies the threshold number of credits.
[0153] Example 16 includes the apparatus of Example 15, wherein the interface connection device is used to: load the first number of credits into the memory when the interface connection device receives the first number of credits from the credit manager; and as one or more data tiles associated with the workload node are transferred from a first computing building block of the one or more computing building blocks to the buffer, transfer credits to the credit manager for each tile transferred to the buffer.
[0154] Example 17 includes the apparatus of example 15, wherein the buffer is an output buffer associated with the workload node, the first number of credits corresponds to the output buffer, and the threshold number of credits corresponds to a threshold amount of memory in the output buffer.
[0155] Example 18 includes the apparatus of example 15, wherein the buffer is an input buffer associated with the workload node, the first number of credits corresponds to the input buffer, and the threshold number of credits corresponds to a threshold amount of data in the input buffer.
[0156] Example 19 includes the apparatus of Example 15, wherein the buffer is a first buffer, the threshold number of credits is a first threshold number of credits, the comparison device is for: comparing the second number of credits with a second threshold number of credits, wherein the second threshold number of credits is associated with memory availability in the second buffer, and the dispatching device is for: selecting a workload node to be executed at a first computing building block of the one or more computing building blocks when (1) the first number of credits satisfies the first threshold number of credits and (2) the second number of credits satisfies the second threshold number of credits.
[0157] Example 20 includes the apparatus of Example 19, wherein the second buffer is an input buffer associated with the workload node, the second number of credits corresponds to the input buffer, and the second threshold number of credits corresponds to a threshold amount of data in the input buffer.
[0158] Example 21 includes the apparatus of Example 15, wherein the threshold number of credits is a first threshold number of credits, the workload node is a first workload node, and the dispatching device is used to schedule the first workload node and the second workload node to execute at a first computing building block in one or more computing building blocks when (1) the first number of credits satisfies the first threshold number of credits and (2) the second number of credits satisfies the second threshold number of credits.
[0159] Example 22 includes a method comprising: loading a first number of credits into a memory; comparing the first number of credits to a threshold number of credits, wherein the threshold number of credits is associated with memory availability in a buffer; and selecting a workload node in the workload to execute at a first compute building block in one or more compute building blocks when the first number of credits satisfies the threshold number of credits.
[0160] Example 23 includes the method of Example 22, further comprising: loading the first number of credits into a memory when the first number of credits is received from the credit manager; and as one or more data tiles associated with the workload node are transferred from a first computing building block of the one or more computing building blocks to the buffer, transferring credits to the credit manager for each tile transferred to the buffer.
[0161] Example 24 includes the method of Example 22, wherein the buffer is an output buffer associated with the workload node, the first number of credits corresponds to the output buffer, and the threshold number of credits corresponds to a threshold amount of memory in the output buffer.
[0162] Example 25 includes the method of Example 22, wherein the buffer is an input buffer associated with the workload node, the first number of credits corresponds to the input buffer, and the threshold number of credits corresponds to a threshold amount of data in the input buffer.
[0163] Although certain example methods, apparatus, and articles of manufacture have been disclosed herein, the scope of coverage of this patent is not limited thereto. On the contrary, this patent covers all methods, apparatus, and articles of manufacture fairly falling within the scope of the claims of this patent.
[0164] The following claims are incorporated into this Detailed Description, with each claim standing on its own as a separate embodiment of the disclosure.
Claims
1. A device comprising: a first computing unit comprising a first local credit manager, the first computing unit being associated with a first buffer, the first computing unit being configured to write data to the first buffer; a second computing unit comprising a second local credit manager, the second computing unit being associated with a second buffer, the second computing unit being configured to read data from the second buffer; at least one infrastructure coupled to the first computing unit and the second computing unit; and a central credit manager coupled to the at least one infrastructure, the central credit manager configured to: causing a first credit to be transmitted to the first local credit manager, the first credit corresponding to first data to be processed by the first computing unit to generate second data to be stored in the first buffer; accessing the first credit from a first local credit manager of the first computing unit; and The credit count of the second computing unit is reduced.
2. The device according to claim 1, wherein The central credit manager accesses the first credit from a first local credit manager of the first computing unit in response to the first computing unit processing the first data.
3. The device according to claim 1, wherein The central credit manager decrements the credit count of the second computing unit in response to availability of the second data at the second buffer.
4. The device according to claim 1, wherein The credit count of the second computing unit is a first credit count, and the central credit manager is configured to: Initializing a second credit count of the first computing unit; and Initializing the first credit count of the second computing unit.
5. The device according to claim 1, wherein The central credit manager causes the first credit to be transmitted to the first local credit manager based on the association of the first data with the task assigned to the first computing unit.
6. A method comprising: transmitting, by executing instructions with a processor circuit, a first credit to a first local credit manager of a first computing unit, the first credit corresponding to first data, the first data to be processed by the first computing unit to generate second data to be stored in a first buffer associated with the first computing unit, the first computing unit to write data to the first buffer; accessing the first credit from a first local credit manager of the first computing unit by executing instructions with a processor circuit; and By executing instructions with the processor circuit, a credit count of a second calculation unit including a second local credit manager is reduced, the second calculation unit being associated with the second buffer, the second calculation unit being configured to read data from the second buffer.
7. The method of claim 6, further comprising: The first credit is accessed from a first local credit manager of the first computing unit in response to the first computing unit processing the first data.
8. The method of claim 6, further comprising: A credit count of the second computing unit is reduced in response to availability of the second data at the second buffer.
9. The method of claim 6, wherein: The credit count of the second calculation unit is a first credit count, and the method further includes initializing a second credit count of the first calculation unit and the first credit count of the second calculation unit.
10. The method of claim 6, further comprising: The first credit is transmitted to the first local credit manager based on the first data being associated with a task assigned to the first computing unit.
11. A device comprising: Memory; instruction; as well as A processor circuit configured to execute the instructions to perform the method according to any one of claims 6 to 10.
12. A non-transitory computer-readable medium comprising instructions that, when executed, cause a processor circuit to perform the method of any one of claims 6 to 10.
13. An apparatus comprising means for performing the method according to any one of claims 6 to 10.
Citation Information
Patent Citations
Methods and apparatus to enable out-of-order pipelined execution of static mapping of a workload
CN112395010A
Mode dependent partial width load to wider register processors, methods, and systems
CN105453030A
Accelerator for Gather-Update-Scatter Operations
CN108228234A