Hierarchical Compilation and Execution on Machine Learning Hardware Accelerators

A hierarchical compilation and execution method optimizes machine learning computations across multi-core devices with diverse processing cores, enhancing efficiency and reducing latency by distributing jobs effectively among ARM and TPU cores.

JP7713537B2Active Publication Date: 2025-07-25GOOGLE LLC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2023571345
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-06-08
Publication Date
2025-07-25
Estimated Expiration
2041-06-08

AI Technical Summary

Technical Problem

Existing machine learning systems face inefficiencies in implementing computations, especially at the edge, due to the need for optimizing job distribution across multi-core computing devices with diverse processing cores, which are not adequately addressed by current methods.

Method used

A hierarchical compilation and execution method is employed, utilizing a multi-core computing architecture with different types of processing cores, including ARM and TPU, to efficiently distribute and execute machine learning jobs by analyzing core suitability and generating execution graphs that optimize operations across these cores.

Benefits of technology

This approach enables efficient implementation of machine learning computations by optimizing performance and power consumption through hierarchical execution graphs, facilitating pipeline parallelism, data parallelism, and model parallelism, while minimizing communication latency and memory overhead.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007713537000001
    Figure 0007713537000001
  • Figure 0007713537000002
    Figure 0007713537000002
  • Figure 0007713537000003
    Figure 0007713537000003
Patent Text Reader

Abstract

This disclosure describes systems and methods for compiling and executing machine learning inferences on an array of multi-core computing devices. Each multi-core computing device can be an application specific integrated circuit (ASIC) or a group of ASICs. In many applications, the array of computing devices changes from inference to inference and can be adjusted based on the requirements of the inference. Furthermore, each ASIC can have multiple processing cores and multiple types of processing cores. Therefore, performing optimization and scheduling at compile time greatly improves the efficiency of the array when performing inference. In some implementations, the time or effort spent on optimization during compilation can be selected, allowing users the flexibility to decide whether to spend time during compilation or during execution.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Technical Field The present disclosure generally relates to the compilation and execution of code in machine learning hardware accelerators.

Background Art

[0002] Background Machine learning systems typically undergo a period of training. Using a trained or partially trained machine learning system to perform a task generally refers to using the system for inference, i.e., processing data to perform the task. Training a machine learning system may include using the system for inference while providing training data to the system.

Summary of the Invention

Problems to be Solved by the Invention

[0003] Machine learning can be implemented on a general-purpose CPU (Central Processing Unit) and / or on dedicated machine learning hardware such as a DSP (Digital Signal Processor), GPU (Graphics Processing Unit), or TPU (Tensor Processing Unit). Some machine learning systems are implemented in the cloud, but there is an increasing need to implement machine learning systems locally or at the "edge," especially when performing inference operations.

[0004] Summary This specification generally relates to techniques for efficiently implementing machine learning and other computations, especially at the edge, e.g., in inference. In implementation, these combine a computing architecture such as a hierarchical architecture with methods adapted to that architecture to compile and execute a machine learning model.

Means for Solving the Problems

[0005] In one aspect, a method for distributing jobs executable in an array of multi-core computing devices is described. The method includes receiving a plurality of jobs to be executed in the array of multi-core computing devices, where each multi-core computing device comprises a plurality of different types of processing cores, and allocating each particular job of the plurality of jobs to be executed by one of the plurality of different types of processing cores.

[0006] Allocating may include analyzing a particular job to determine which of the plurality of different types of processing cores is suitable for execution of the particular job, and allocating the particular job to a core type based on the analysis. Analyzing which core is suitable for a particular job may include, for example, using a model to evaluate one or more metrics of core suitability for the job and / or processing data representing the job, which may be a machine learning model for deterministically or probabilistically allocating jobs to cores.

[0007] The method may further include compiling each job of the plurality of jobs into an individually executable file, and generating an execution graph representing a mapping of the individually executable files to a particular one of the plurality of different types of processing cores. In an implementation, the execution graph identifies dependencies between the individually executable files, particularly to distribute the executable jobs across the multi-core computing devices, more specifically to the processing cores. For example, in an implementation, the dependencies within the execution graph define the sequence in which the jobs are to be executed.

[0008] This method can include executing individually executable files. This can include receiving an execution graph and using the graph, and in particular the dependencies it identifies, to assign jobs within the execution graph to a plurality of multi-core computing devices within an array of multi-core computing devices. This can also include each multi-core computing device executing the assigned jobs, returning the output of the executed jobs to shared memory, and combining the returned outputs to generate an execution graph return.

[0009] Analyzing which cores are suitable for a particular job can include heuristic analysis. For example, the metrics for core suitability for a job can include heuristic metrics. The depth of analysis for each particular job can be selected based on user input prior to compile time. The depth of analysis can be represented, for example, by the computing resources or time allowed for the analysis to determine core suitability for a job.

[0010] Different types of processing cores can include a first type of core and a second type of core. The first type of core can be an ARM processor (core), i.e., a core with a RISC (Reduced Instruction Set Computing) architecture. Such an architecture can be characterized by a load / store architecture with single-cycle memory access instructions. The second type of core is a TPU (Tensor Processing Unit) or a TPU tile processor (core). Such a TPU core can be characterized by hardware configured to implement one or more of the following. The following are tensor operations on tensors of three or more dimensions, matrix-matrix multiplication, integer matrix operations, systolic arrays, and activation units that implement neural network activation functions.

[0011] In an implementation, the execution graph is hierarchical. Thus, the execution graph can include (hierarchical) subgraphs arranged in layers, for example, at least four layers. The layers can include: i) a TPU layer including executable files executed on a TPU core type; ii) a chip-level (physical integrated circuit level) layer including one or more subgraphs of the TPU layer and executable files executed on an ARM core type; iii) a multi-chip layer including two or more chip-level subgraphs; and iv) a host-level layer including a subgraph of the multi-chip layer and one or more subgraphs configured to be executed on a third type of core. The third type of core can be a CPU, for example, the CPU of a host device. A subgraph of a layer can constitute a part of the execution graph, more specifically, a part of the execution graph executed in a lower layer. Similar to the execution graph, it can define dependencies between jobs and the order in which jobs need to be executed.

[0012] In some implementations, a mechanism can be provided to coordinate and order operations between an ARM core and a multi-core computing device, such as a TPU core of an ASIC. For example, this can include code (an "interpreter") executed on the ARM core, for example, in firmware, that schedules operations of either the ARM core or the TPU core, for example, for low latency. This can be used to schedule jobs at runtime and assign them to processing cores.

[0013] In some implementations, the execution graph can include another constant buffer, that is, a memory area allocated for storing constants. In such cases, the constant buffer does not necessarily have to be part of the execution graph itself. Instead, one or more "off-chip buffers" can be associated with the graph at runtime. This helps keep the footprint of the graph memory small.

[0014] In another aspect, a method for compiling jobs executable for execution in an array of multi-core computing devices is described. In an implementation, the array of multi-core computing devices is combined with hardware such as a host system. In an implementation, the hardware, such as the host system, includes processing cores of a first core type. Each multi-core computing device of the array of multi-core computing devices includes processing cores of a second core type and a third core type.

[0015] In an implementation, the method includes receiving a machine learning model to be used for inference, analyzing the machine learning model to determine a plurality of jobs to be executed, and generating an execution graph representing each of the plurality of jobs to be executed and, for example, dependencies between the plurality of jobs to be executed as described above.

[0016] In an implementation, the method further includes calling a multi-chip level compiler to generate an execution graph, such as a mapped execution graph. The mapped execution graph can represent a mapping of individually executable files to specific ones of a plurality of different types of processing cores. In an implementation, this includes the multi-chip level compiler identifying one or more first jobs among the plurality of jobs to be executed by the first core type and compiling the one or more first jobs into an executable file to be executed by the first core type. In an implementation, the first jobs are not compatible with the multi-core computing devices within the array. A job may not be compatible with the multi-core computing devices if one or more operations of the job cannot be executed on the multi-core computing devices or if the job is not suitable for execution on the multi-core computing devices (where suitability can be determined as described above).

[0017] In an implementation, the method further includes splitting the remaining jobs of the execution graph into a plurality of first subgraphs, assigning each first subgraph to a specific multi-core computing device of an array of multi-core computing devices, and invoking a single-chip level compiler for each first subgraph.

[0018] In an implementation, the method includes a single-chip level compiler that identifies one or more chip-level jobs from a first subgraph executed by a second core type, compiles each of the one or more chip-level jobs from the first subgraph into an executable file executed by the second core type, splits the remaining jobs of the first subgraph into a plurality of second subgraphs, and assigns each of the plurality of second subgraphs to a third core type. In an implementation, the method further includes invoking a core-level compiler for each of the plurality of second subgraphs, and the core-level compiler compiles each of the second subgraphs into an executable file executed by the third core type.

[0019] In an implementation, the mapped execution graph is for distributing the compiled executable jobs, for example, as described above. Thus, this method includes using the mapped execution graph for distributing executable jobs, for example, as described above.

[0020] The first, second, and third core types may respectively correspond to the third type of core, the first type of core, and the second type of core described above. For example, the first core type may be a CPU (of the host system). The second core type may be an ARM (RISC) core. The third core type may be a TPU (tile) core. Each multi-core computing device of the array may include an application-specific integrated circuit (ASIC) with a TPU.

[0021] In an implementation, identifying one or more first jobs and / or one or more chip-level jobs is performed based on a heuristic analysis of a plurality of jobs to be executed. The heuristic analysis may be an analysis that includes determining a heuristic metric for each of the jobs. In an implementation, based on, for example, computing resources or time allowed for the analysis, the depth of the heuristic analysis for each specific job is selected based on user input prior to compile time.

[0022] In an implementation, the method further includes receiving, by an array of multi-core computing devices, a mapped execution graph including a first job, and a plurality of first subgraphs including one or more chip-level jobs and a plurality of second subgraphs, and assigning the first job, the chip-level job, and the plurality of remaining jobs within the second subgraph to associated cores within the array of multi-core computing devices. In an implementation, the method further includes executing, by each core of the multi-core computing device, the assigned job, returning the output of the executed job to shared memory, and combining the returned outputs to generate a return of the execution graph.

[0023] In some implementations, a second core type, such as an ARM or RISC core, can be assigned control flow operations that span jobs on multiple chips, facilitating single-chip operations or processes, such as beam search operations or processes, followed by multi-chip operations or processes. This method may then include determining (e.g., using a compiler) references to control flow operations executed by the second core type in a sequence graph that combines a plurality of first (single-chip level) subgraphs, which may be on another chip (ASIC), such as a master chip (ASIC). The sequence graph may be processed, for example, at runtime by an interpreter, for example, so that the second core type of the master chip controls the execution of multi-chip operations or processes spanning multiple chips.

[0024] The methods and method features according to the above aspects can be combined. [Advantages of the Invention]

[0025] Various implementations provide one or more of the following advantages. The implementation of these methods provides a hierarchical compiler that generates hierarchical executable files that can span the host CPU and the firmware within the multi-core computing device (ASIC) at runtime. The described hierarchical compiler method, in combination with a hierarchical architecture that includes different types of hardware resources, facilitates an efficient implementation of machine learning and other computations. This is because, in the implementation, various types of hardware resources, including ARM cores and TPUs, are exposed to the compiler. For example, the compiler can be compiled into firmware that runs on these resources. For example, given a machine learning model or other computation, the method can analyze the model or computation and determine an optimal way to execute the model / computation using the various hardware units exposed to the compiler, e.g., from the perspective of performance and power consumption. Also, operations can be split between different hardware units to limit communication between the CPU and the ASIC and optimize the generated executable code for high performance and low power consumption. Further, the graph of operations can be split so that operations that are not well-suited for the TPU core are executed on the ARM core in a low-latency manner.

[0026] In some implementations, the lowest level of the hierarchical architecture is a level of only TPUs, followed by a single-chip (ASIC) level that includes an ARM core and one or more TPU cores, optionally followed by a multi-chip (multi-ASIC) level, and optionally followed by a host (CPU) level. The executable files generated at a certain level of the compiler can be embedded in executable files generated at a higher level, etc., and the executable files at a particular level become the "contract" between the compiler and runtime at that level. With this approach, since the compiler can compile at the single-chip level and the multi-chip level, efficient operations are further facilitated, e.g., to implement pipeline parallelism, data parallelism, and / or model parallelism, by calling the single-chip level compiler for each chip to compile the subgraph that runs on that chip. The multi-chip level code can be executed on the firmware.

[0027] In such a hierarchical approach, an executable file at the single-chip level can include operations on multiple different types of cores, such as TPU cores and ARM cores. This facilitates the execution of operations such as beam search and sorting. These operations are facilitated by the availability of ARM cores in addition to the TPU. Also, at the single-chip level, this approach enables the parallel execution of mixed operations of the TPU and ARM cores while performing data transfer streaming and synchronization. This synchronization can be represented through the execution graph.

[0028] In implementation, including a host CPU within the hierarchy can further facilitate efficient operations. For example, this enables sharing of buffers between the host and the ASIC, avoiding costly memory copy operations. This can facilitate fine-grained synchronization. Also, it becomes easier for the host CPU to consume the data generated by the ASIC. Such partitioning and scheduling are facilitated by the described graph-based job mapping and execution.

[0029] Details of one or more implementations of the subject matter of this specification are set forth in the accompanying drawings and the following description. Other features, aspects, and advantages of the subject matter will become apparent from the description, the drawings, and the claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0030]

Figure 1

Figure 2

Figure 3A

Figure 3B

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

DETAILED DESCRIPTION OF THE INVENTION

[0031] Like reference numerals and names in the various drawings indicate like elements. DETAILED DESCRIPTION The present disclosure describes systems and methods for compiling and executing machine learning inferences in an array of multi-core computing devices. Each multi-core computing device can be a dedicated integrated circuit (ASIC) or a group of ASICs. In many applications, the array of computing devices changes for each inference and can be adjusted based on the requirements of the inference. Further, each ASIC can have multiple processing cores and multiple types of processing cores. Thus, performing optimization and scheduling during compilation can significantly improve the efficiency of the array during inference execution. In some implementations, since the time or effort spent on optimization during compilation can be selected, the user can flexibly decide whether to spend time during compilation or during execution.

[0032] FIG. 1 shows an example of a system architecture of a machine learning hardware accelerator that compiles and executes a graph executable file. The hardware accelerator 100 includes a host system 102 that not only instructs and coordinates operations but also provides an interface between the user and the accelerator 100. The host system 102 interacts with an array of ASICs 108. Each ASIC 108 includes a plurality of core types and is configured to perform most operations during machine learning inference.

[0033] The host system 102 includes one or more central processing units, i.e., CPUs 104. The CPUs 104 can provide processing to the host and execute certain control or logistics operations. In some implementations, the CPUs 104 can execute some processes during inference. Generally, the CPUs 104 execute instructions, manipulate data, and perform operations of the host system 102. Each CPU 104 can have a single or multiple cores, and each core is available for hosting and executing individual processing threads. Further, the number, type, and specific CPUs 104 used to perform the operations described herein can be dynamically determined based on the requirements, interactions, and number of operations associated with the host system 102.

[0034] The host system 102 also includes a memory 106. The memory 106 of the host system 102 can represent a single memory or multiple memories. The memory 106 can include any memory or database module, can take the form of volatile or non-volatile memory, which includes, but is not limited to, magnetic media, optical media, random access memory (RAM), read-only memory (ROM), removable media, or any other suitable local or remote memory component. The memory 106 can store various objects or data, including an execution graph, a machine learning model, management settings, a cache, an application, backup data, and any other suitable information related to the host system 102, which includes any parameters, variables, algorithms, instructions, rules, constraints, or references thereto. Although shown within the host system 102, the memory 106, or any portion thereof that includes some or all of the specific components shown, may in some cases be located remotely from the host system 102, and in some cases, this includes as a cloud application or repository, or as a separate cloud application or repository if the host system 102 itself is a cloud-based system. In some examples, the data stored in the memory 106 is accessible, for example, via the network 120 and can be obtained by the function of a specific application or the hardware accelerator 100.

[0035] Generally, the host system 102 executes Possible jobs (described in more detail below) while distributing it across the array of ASICs 108 to execute high-level applications and provide a "front end" to the user.

[0036] The ASIC 108 within the array includes a host interface 110, a core processor 112, an array of tiles 116 that can be the main computing unit of the ASIC 108, as well as a peer-to-peer interface 114 and a shared memory 118. The core processor 112 can be a processor that executes operations and controls the ASIC 108 and can include, for example, ARC, Alpha, Am29000, ARM, Atmel AVR, Blackfin, i860, i960, M88000, MIPS, PA-RISC, Power ISA, RISC-V, SuperH, SPARC, or other processing architectures.

[0037] The shared memory 118 can be a memory that is accessed by the tiles 116, the core processor 112, and across multiple ASICs 108 via a high-speed network 122. The shared memory 118 can include any memory or database module and can take the form of volatile or non-volatile memory, including, but not limited to, magnetic media, optical media, random access memory (RAM), read-only memory (ROM), removable media, or any other suitable local or remote memory component. The shared memory 118 can store various objects or data, management settings, caches, applications, backup data, a repository for storing dynamic information, and any other suitable information related to the hardware accelerator 100, including any parameters, variables, algorithms, instructions, rules, constraints, or references for inference. The shared memory 118 includes a shared address space used by each of the multiple tiles 116 within the ASIC 108.

[0038] The host interface 110 is used to coordinate and manage the communication between the ASIC 108 and the host system 102. Generally, the host interface 110 comprises logic encoded in software and / or hardware in an appropriate combination and operable to communicate with the host system 102 and other components. More specifically, the interface 110 can comprise software that supports one or more communication protocols associated with communication such that the network 110 and / or the hardware of the interface are operable to communicate physical signals both inside and outside the illustrated accelerator 100. Further, the interface 110 can enable the ASIC 108 to communicate with the host system and / or network 120 to perform the operations described herein.

[0039] The peer-to-peer interface 114 can be made similar to the host interface 110, except that it provides and manages communication from ASIC 108 to ASIC 108. In this way, the ASICs 108 can distribute compute jobs between them and share return or intermediate parameters. Peer-to-peer communication can minimize the load on the host system 102 and its associated CPU 104 and provide a scalable solution. This enables a system 100 with any number of ASICs 108 that is not limited by the host system 102 or CPU 104.

[0040] The ASIC can include a core processor 112 such as an Advanced RISC Machine (ARM) core. The core processor 112 can process the control and management of jobs and tasks distributed among the tiles 116. The core processor 112 performs computational operations and manages the ASIC 108. Further, some operations during inference can be executed more efficiently or quickly on the core processor 112. The core processor 112 instructs and commands the tiles 116 to perform calculations. It maintains one or more contexts that define the information necessary to execute the inference process. Each context can include, but is not limited to, instructions, activation data, parameters, hardware state, computational operands, and results. This data can be stored in tile memory or shared memory 118. In some implementations, the core processor 112 operates on ARC, Alpha, Am29000, ARM, Atmel AVR, Blackfin, i860, i960, M88000, MIPS, PA-RISC, Power ISA, RISC-V, SuperH, SPARC, or other processing architectures.

[0041] The tile 116 can be a custom computing core configured to execute inference. Each tile 116 can include memory and can receive inputs and outputs that can be shared among the tiles or between the core processor 112 and the tile 116. Each tile 116 can access the shared memory 118 via a high-speed network 122 in addition to its own memory (e.g., SRAM). The tiles 116 are described in more detail below with reference to FIGS. 6 and 7.

[0042] FIG. 2 is a diagram showing an example of an execution graph 200. The execution graph 200 includes a plurality of operations 202A to 202J that are executed by a machine learning hardware accelerator. The arrows in FIG. 2 represent the dependency relationships between the operations. For example, operation 202C depends on the outputs of operations 202A and 202B. Note that the illustrated execution graph is simplified for clarity, and an actual execution graph can be composed of thousands of operations and dependency relationships. This initial graph can be constructed based on a trained machine learning model to be executed. For example, an MLIR file can be provided and compiled, or partially compiled to generate a list of operations and dependency relationships to construct the execution graph 200.

[0043] The execution graph 200 can generally describe the operations that need to occur to perform an inference. The operations can be basic-level calculations (e.g., AND operation, OR operation, or XOR operation) or more advanced-level calculations such as comparisons and averages. The computational costs of all operations are not equal, and some operations are executed faster or more efficiently on specific core types. For example, an operation that requires a series of sequential calculations may be more suitable for an ARM-type processor or the like (e.g., the core processor 112 in FIG. 1). In another example, a group of parallel operations that share a single input may be most suitable for a parallel processor such as a GPU or TPU (e.g., the tile 116 in FIG. 1).

[0044] FIG. 3A is a diagram showing an example of an execution graph with partitioning and assignment. For example, when compiling the graph for a specific inference on a specific hardware accelerator to improve the efficiency of the machine learning hardware accelerator when executing the execution graph, the graph can be further processed. For example, when compiling this graph, each partition 306 can be assigned to a specific ASIC, and each operation 302 can be assigned to a specific core for execution.

[0045] The execution graph can be executed in a distributed environment, and the way the operations are split among various processing units can affect the efficiency and speed of the inference being executed. The hardware configuration is determined at compile time. For example, a host system can use 10 multi-core devices (e.g., Google TPU) to execute a particular inference. Each multi-core device can have multiple processing cores and multiple core types. For example, a multi-core device can have an array of processing "tiles" (e.g., tile 116 described with respect to FIG. 1) and one or more core processors. Once the hardware configuration for executing the inference is known, the execution graph 300 can be further defined and optimized to execute on the known hardware configuration.

[0046] The execution graph can be split into various subgraphs. The split 306 can be chosen to separate groups of relatively independent operations. In some implementations, the split can represent a checkpoint where the hardware accelerator synchronizes parameters before continuing execution, i.e., a synchronization point of the inference. In some implementations, to maximize parallel computing, each part of the execution graph 300 after splitting can be split among the processing devices of the hardware accelerator.

[0047] Each operation within the execution graph 300 can be evaluated, and a preferred processing core can be selected for that operation. For example, operations 302A - 302H are more suitable for execution on the TPU tile, and thus should be preferentially executed by the TPU tile. Operations 304A - 304C are more suitable for execution on the core processor, and thus can be preferentially executed by the ASIC's core processor (e.g., core processor 112 of FIG. 1). In some implementations, the preferred core is not necessarily the core on which the operation is executed. For example, in the case of a series of operations where the preferred core types alternate, it may be optimal to execute all operations on a single core to minimize communication traffic within the hardware accelerator. Further, although only two preferred core types (operations 302 and 304) are shown, three or more preferred types are contemplated by this disclosure. For example, some operations may be optimal for the CPU of the host system (e.g., CPU 104 of FIG. 1) and thus should be executed by the host system.

[0048] In some implementations, certain operations require a specific core type. For example, some operations may only be executable by an ARM core (or other specific RISC core), and these operations cannot be properly executed on the TPU tile. For example, in some implementations, gather, scatter, or beam search operations may not be executable by the TPU tile. Operations with hard core type requirements can be appropriately assigned to the appropriate core. Many operations can be executed on either core, but one type is more efficient. Further, combinations or groups of operations may be more suitable for a particular core type. A heuristic analysis can be performed to analyze the execution graph 300, or operations 302 and 304, to determine which core type is preferred for each operation. The heuristic analysis may include an evaluation of the number of tiles used (e.g., an attempt to maximize the number of tiles used). In some implementations, the heuristic analysis calculates the time delay or software overhead of each core type.

[0049] Once a preferred core type is determined and the execution graph 300 is split into partitions, operations can be assigned to specific hardware of the machine learning hardware accelerator. Generally, the assignment of operations to specific hardware components can be performed based on the preferred core type of the operation, the expected communication traffic, and the available hardware. For example, a general machine learning hardware accelerator may have more TPU tiles than a core processor, so operations can be preferentially assigned to the TPU tiles. Further, the assignment can be completed hierarchically. For example, the host system can distribute large groups of operations among the available ASICs and individually assign operations within the host system to specific tiles / processors. In some implementations, the host system may only need to assign the entire execution graph 300 to a single ASIC and can distribute a portion of the graph in a peer-to-peer manner across an array of ASICs. This will be further described with reference to FIG. 3B.

[0050] In some implementations, the amount of optimization performed on the execution graph 300 during compilation is adjustable. For example, the user can specify a specific time spent on optimizing and analyzing the execution graph 300 before inference begins.

[0051] FIG. 3B is a diagram showing an example of an execution graph with additional upper-level assignments. As shown in FIG. 3A, after a preferred core type is determined, individual operations are assigned to specific cores. Once the required core type is determined, for each partition, a specific ASIC can be selected by the upper-level compiler. For example, in the second partition, operations 302D and 302E are assigned to tiles A and B of ASIC#1 of the hardware accelerator, respectively. On the other hand, operations 302C and 304B are assigned to tile A and the ARM core of ASIC#2.

[0052] Due to their hierarchical nature, each layer of the assignment needs to be performed only by the relevant components of the machine learning hardware accelerator. For example, a host system (e.g., host system 102) can provide the execution graph 300 to ASIC#1 of the hardware accelerator. Then, ASIC#1 can offload operations 302C and 304B to ASIC#2 and assign operations 302D and 302E to tiles A and B, respectively. On the other hand, ASIC#2 can receive the assigned operations and distribute those operations among appropriate computing cores (e.g., tiles, or ARM cores) within ASIC#2.

[0053] Figure 4 is a flowchart illustrating an exemplary process for distributing executable jobs that can be executed in an array of multi-core computing devices. Process 400 can be executed by a machine learning hardware accelerator (e.g., machine learning hardware accelerator 100 described with respect to FIG. 1) or a portion thereof.

[0054] At 402, a plurality of executable jobs are received in an array of multi-core computing devices. The array of multi-core computing devices can be similar to the machine learning hardware accelerator 100 described with reference to FIG. 1. In some implementations, the plurality of executable jobs can be received as a list of trained machine learning models or model characteristics.

[0055] At 404, each job of a plurality of jobs is assigned to a specific core type of a multi-core computing device. Optionally, an analysis can be performed to determine the core type that is most optimal for the execution of each job. Examples of core types can include, but are not limited to, ARM cores (or other RISC cores), CPUs, GPUs, and TPUs. The analysis can be performed based on executable jobs, user input, as well as heuristic analysis of hardware requirements and availability. The heuristic analysis can determine which job or group of jobs can be executed most efficiently on which core type. The user can input, including parameters that define the heuristic analysis to be performed, the time spent on the analysis, the priority of the analysis, and the like. In some implementations, the user input can include the desired depth of analysis, which can, for example, describe the number of calculations to be performed per job in order to determine to which core type a job should be assigned. The hardware requirements can include specific hardware limitations that are necessary to execute a particular job on a particular core type. For example, a TPU tile may not be able to execute a communication routing job that involves sending the return of a tensor object to another ASIC. This job may need to be executed by the core processor of the ASIC itself. Further, the available hardware can provide information for the analysis for job assignment. For example, a hardware accelerator may be able to utilize more of the first type of processing core than the second type of processing core. In this example, based on the additional relative availability of the first type of core, jobs can be preferentially assigned to the first type of core.

[0056] At 406, each job is compiled into an individually executable file according to its core type assignment. These individually executable files can be consumed by their assigned core type with one or more inputs and configured to generate one or more outputs. For example, a job assigned to a TPU tile is compiled into a TPU-executable file. Similarly, a job assigned to an ARM is compiled into an ARM-executable file.

[0057] At 408, an execution graph is generated that represents the mapping of a particular type of individually executable file to a processing core. The execution graph can identify dependencies between individually executable files. In some implementations, the execution graph is a node and edge graph, where each node represents an executable file and additional metadata or information, and each edge represents a dependency between two nodes. The execution graph can be similar to the execution graph 200 or 300 described with respect to FIGS. 2, 3A, and 3B. The execution graph can further be hierarchical and include one or more subgraphs similar to the execution graph 800 described in more detail below with reference to FIG. 8. Each node in the execution graph can include one or more executable files, and these executable files can be distributed across the entire machine learning hardware accelerator.

[0058] FIG. 5 is a flowchart illustrating an exemplary process for compiling jobs executable in an array of multi-core computing devices. Process 500 can be executed by a machine learning hardware accelerator (e.g., the machine learning hardware accelerator 100 described with respect to FIG. 1) or a portion thereof.

[0059] At 502, the machine learning hardware accelerator receives a machine learning model for performing inference. The machine learning model can define parameters such as the weights and connections between neurons in a neural network, and the number of layers / neurons in each layer. The received machine learning model can also include one or more inputs provided to the neural network for performing inference, as well as the operations to be performed and the specific inputs, outputs, and parameters used in the operations.

[0060] At 504, the received machine learning model is analyzed to determine a plurality of jobs to be executed. The plurality of jobs can include jobs that depend on the results from other jobs, as well as other computations to be completed for communication between systems and performing inference. In some implementations, the jobs are evaluated to determine on which core type they should be preferentially executed. For example, a heuristic analysis similar to that described above can be performed to identify the core type on which a job will perform optimally.

[0061] At 506, an execution graph representing the plurality of jobs and the dependencies between the plurality of jobs is generated. In some implementations, the execution graph is a node and edge diagram similar to execution graph 200 or 300 described with reference to FIGS. 2 and 3. In some cases, this is completed similar to 408 above.

[0062] At 508, a multi-chip compiler is invoked to compile the execution graph into a mapped execution graph. The mapped execution graph is a compiled graph where all jobs are assigned to appropriate core types. The mapped execution graph includes the compiled executable files necessary to be executed on a multi-core machine learning accelerator. Generally, the multi-chip compiler divides the graph into subgraphs to be processed by lower-level compilers. Some jobs that need to be executed at the topmost level are immediately compiled, and the remaining jobs are further divided into subgraphs and compiled by respective compilers. Although described as a three-layer hierarchy of multi-chip level, single-chip level, and core level, within the scope of the present disclosure, more or fewer layers may be considered.

[0063] At 510, one or more first jobs among a plurality of jobs that are not compatible with the multi-core computing device are identified. In other words, these jobs must be executed at a high level (e.g., by a host CPU such as CPU 104 described with reference to FIG. 1). For example, start of execution, stop, job return, etc. The host CPU may constitute the first core type.

[0064] At 512, the remaining jobs of the execution graph are divided into a plurality of first subgraphs representing the single-chip level and assigned to multi-core computing devices within an array of multi-core computing devices. The array of multi-core computing devices can be an ASIC similar to ASIC 108 described in more detail with reference to FIG. 1 and hereinafter with reference to FIGS. 6 and 7.

[0065] At 514, a chip-level compiler is invoked to compile each of the first subgraphs. Each upper-level subgraph includes an executable file assigned to the current layer or an executable file to be further divided into lower-level subgraphs.

[0066] At 516, one or more chip-level jobs that are only executable at the chip level, not suitable at the chip level, or preferentially executed at the chip level are identified. These jobs are, for example, traffic shaping jobs or synchronization jobs between cores of a multi-core computing device. Next, these jobs are compiled to be executed by a second core type (e.g., an ARM core controller of a multi-core computing device).

[0067] At 518, the remaining jobs of the first subgraph are split into a plurality of second subgraphs and assigned to a third core type (e.g., a TPU tile). At 520, a core-level compiler is called to process each of the second subgraphs. At 522, the core-level compiler compiles each of the second subgraphs into one or more executable files to be executed by a third core type (e.g., a TPU tile).

[0068] At 524, the resulting mapped execution graph is returned, resulting in an execution graph that includes executable files and subgraphs. Each subgraph itself includes an executable file and potential additional subgraphs. Each graph and subgraph can specify the core type to be executed in an array of multi-core computing devices.

[0069] FIG. 6 shows a block diagram of an ASIC used as a machine learning hardware accelerator as an exemplary computing system 600 for accelerating tensor calculations related to a deep neural network (DNN). The system 600 can be, for example, the ASIC 108 described with reference to FIG. 1. The system 600 generally includes a controller 602, a host interface 608, an input / output (I / O) link 610, a plurality of tiles including a first tile set 612 and a second tile set 614, a classifier portion 616, and a data bus identified by a bus map 618 (shown for clarity but not included in the system 600). The tiles of tile set 612 and tile set 614 may be the same as or different from the tiles 116 described with reference to FIG. 1. The controller 602 generally includes a data memory 604, an instruction memory 606, and at least one processor configured to execute one or more instructions encoded on a computer-readable storage medium. The instruction memory 606 can store one or more machine-readable instructions executable by one or more processors of the controller 602. The data memory 604 can store and later access various data related to calculations occurring within the system 600 and can be any of various data storage media.

[0070] Controller 602 is configured to execute one or more instructions related to tensor calculations within system 600, including the instructions stored in instruction memory 606. In some implementations, data memory 604 and instruction memory 606 are one or more volatile memory units. In some other implementations, data memory 604 and instruction memory 606 are one or more non-volatile memory units. Data memory 604 and instruction memory 606 may be another form of computer-readable medium, which may be, for example, a floppy (registered trademark) disk device, a hard disk device, an optical disk device, or a tape device, a flash memory or other similar solid-state memory device, or an array of devices including devices within a storage area network or other configurations. In various implementations, controller 602 may be referred to as, or also called, core manager 602.

[0071] As shown, host interface 608 is coupled to I / O link 610, controller 602, and classifier portion 616. Host interface 608 receives instructions and data parameters from I / O link 610 and provides the instructions and parameters to controller 602. Generally, instructions can be provided to one or more devices within system 600 through instruction bus 624 (described later), and parameters can be provided to one or more devices within system 600 through ring bus 628 (described later). In some implementations, instructions are first received by controller 602 from host interface 618 and stored in instruction memory 606 for later execution by controller 602.

[0072] Classifier portion 616 is similarly coupled to controller 602 and tile 7 of the second tile set 614. In some implementations, classifier portion 616 is implemented as a separate tile within system 600. In an alternative implementation, classifier portion 616 is arranged or positioned within controller 602 as a sub-circuit or sub-device of controller 602. Classifier portion 616 is generally configured to perform one or more functions on the accumulated pre-activation values received as the output of a fully connected layer. The fully connected layer may be split across the tiles within tile sets 612 and 614. Thus, each tile is configured to generate a subset of the pre-activation values (i.e., linear outputs) that can be stored in the memory unit of the tile. Classification result bus 620 provides a data path from classifier portion 616 to controller 602. Data including post-function values (i.e., results) is provided from classifier portion 616 to controller 602 via classification result bus 620.

[0073] Bus map 618 indicates a data bus that provides one or more interconnected data communication paths between the tiles of the first tile set 612 and the tiles of the second tile set 614. Bus map 618 provides a legend for identifying classification result bus 620, CSR / master bus 622, instruction bus 624, mesh bus 626, and ring bus 628, as shown in FIG. 6. Generally, a tile is a core component within the accelerator architecture of system 600 and is the focus of tensor calculations performed within the system. Each tile is an individual computing unit that coordinates with other tiles within the system to accelerate calculations across one or more layers of a multi-layer neural network. Tiles within tile sets 612, 614 can share the execution of tensor calculations associated with a given instruction, but the individual computing units are self-contained computing components configured to execute a subset of the tensor calculations independently of other corresponding tiles within tile sets 612, 614.

[0074] The CSR bus 622 is a single-master multi-slave bus that enables the controller 602 to send one or more instructions for setting the program configuration and reading status registers associated with one or more tiles. The CSR bus 622 can be connected in a single daisy-chain configuration having one master bus segment and multiple slave bus segments. As shown in FIG. 6, the CSR bus 622 provides a communication coupling via a bus data path that connects the tiles within tile sets 612, 614 and the controller 602 within the ring to the host interface 610. In some implementations, the host interface 610 is the single master of the CSR bus ring, and the entire CSR bus address space is memory-mapped to the memory space within the host interface 610.

[0075] The CSR bus 622 can be used by the host interface 610 to perform one or more operations, such as programming the memory buffer pointer within the controller 602 to enable the controller 602 to start fetching instructions from the instruction memory 606, updating / programming various tile settings (e.g., coefficient tables for polynomial approximation calculations) that remain static during one or more calculations, and / or loading / reloading firmware into the classifier portion 616. In one example, the reloading of firmware can include a new function applied to the linear output (i.e., pre-activation value). Thus, all slaves accessible via the CSR bus 622 will have a distinct node identifier (node ID) associated with and identifying the slave. The node ID is part of the instruction address and is used, examined, or otherwise checked by the CSR slaves (i.e., the controller 602, tiles 612, 614, and classifier 616) to determine whether a CSR packet is addressed to the slave. Device Including. In one example, the reloading of firmware can include a new function applied to the linear output (i.e., pre-activation value). Thus, all slaves accessible via the CSR bus 622 will have a distinct node identifier (node ID) associated with and identifying the slave. The node ID is part of the instruction address and is used, examined, or otherwise checked by the CSR slaves (i.e., the controller 602, tiles 612, 614, and classifier 616) to determine whether a CSR packet is addressed to the slave.

[0076] In some implementations, one or more instructions may be sent by the host interface 602 via the controller 602. The instructions can be, for example, 32 bits wide, and the first 7 bits contain header information indicating the instruction address / destination to receive and execute the instruction. The first 7 bits of the header may contain a data parameter representing a specific node ID. Thus, a slave on the CSR bus ring (e.g., each tile) can examine the header of the instruction to determine whether the request by the master (host interface 610) is addressed to the tile examining the header. If the node ID of the header does not indicate that the destination is the examining tile, the examining tile copies the input CSR instruction packet to the CSR bus input connected to the next tile for examination by the next tile.

[0077] The instruction bus 624 starts from the controller 602 and, similar to the CSR bus 622, also provides a communication connection via a bus data path that connects the tiles in the tile sets 612, 614 in the ring back to the controller 602. In one implementation, the controller 602 broadcasts one or more instructions via the instruction bus 624. The instructions broadcast by the controller 602 can be different from the instructions provided via the CSR bus 622. However, the way a tile receives and / or consumes or executes an instruction received via the bus 624 can be similar to the process for executing an instruction received via the CSR bus 622.

[0078] In one example, the header of the instruction (i.e., the bitmap) indicates to the receiving tile that the receiving tile needs to consume a specific instruction based on the bitmap associated with the instruction. The bitmap can have a specific width defined with respect to bits. Instructions are typically transferred from one tile to the next based on the parameters of the instruction. In one implementation, the width of the instruction bus 624 can be configured to be smaller than the size / width of the instruction. Thus, in such a configuration, the transmission of the instruction is done over several cycles, and the bus stop of the instruction bus 624 has a decoder that places the instruction received at the tile into the appropriate target instruction buffer associated with that tile.

[0079] As further described below, the tiles within the tile sets 612, 614 are generally configured to support two broad categories of instructions. The two broad categories can also be referred to as instruction types. The instruction types include TensorOp instructions and Direct Memory Access (DMAOp) instructions. In some implementations, the DMAOp instructions have one or more specializations that are permitted to execute simultaneously. The one or more specializations can be referred to as DMAOp instruction subtypes or opcodes. In some cases, all the unique and / or valid DMAOp instruction type / subtype tuples will have individual instruction buffers within a particular tile.

[0080] At a particular tile of the tiles 612, 614, the bus stop associated with the instruction bus 624 examines the header bitmap to determine the instruction type / subtype. The instruction can be received by the tile and then written to the instruction buffer of the tile before execution of the instruction by the tile. The instruction buffer of the tile where the instruction is written can be determined by the type and subtype indicator / field of the instruction. The instruction buffer can include a first-in, first-out (FIFO) control scheme that prioritizes the consumption of one or more related instructions. Thus, in this FIFO control scheme, instructions of the same type / subtype are always executed in the order in which the instructions arrived at the instruction bus.

[0081] The different instruction buffers within a tile are the TensorOp instruction buffer and the DMAOp instruction buffer. As shown above, the instruction types include TensorOp instructions and DMAOp instructions. With respect to DMAOp instructions, the instruction subtypes (indicating the location of the "destination" buffer) include the following. The following are 1) the mesh receive instruction buffer, 2) the mesh transmit instruction buffer, 3) the narrow-wide DMA instruction buffer, 4) the wide / narrow DMA instruction buffer, and 5) the ring bus DMA instruction buffer. These buffer locations will be described in more detail below with reference to FIG. 7. The wide and narrow designations are used throughout the specification and generally refer to the approximate size of the width (bits / bytes) of one or more memory units. As used herein, "narrow" can refer to one or more memory units each having a size or width of less than 16 bits, and "wide" can refer to one or more memory units each having a size or width of less than 64 bits.

[0082] The mesh bus 626 provides a data communication path different from the CSR bus 622, the instruction bus 624, and the ring bus 628 (described below). As shown in FIG. 6, the mesh bus 626 provides a communication path that couples or connects each tile to its corresponding adjacent tile in both the X and Y dimensions. In various implementations, the mesh bus 626 can be used to transfer input activation amounts between one or more narrow memory units within adjacent tiles. As shown, the mesh bus 626 does not permit direct transfer of input activation data to non-adjacent tiles.

[0083] In various implementations, the mesh bus 626 and the various tiles connected via the mesh bus 626 may have the following configuration. The four corner tiles of the mesh have two transmit ports and two receive ports. The four edge tiles of the mesh have three receive ports and three transmit ports. All tiles other than the edges and corners have four receive ports and four transmit ports. In general, considering an example of an NxN tile layout, an edge tile is a tile that has only three adjacent tiles, while a corner tile is a tile that has two adjacent tiles. Regarding the data flow method via the mesh bus 626, generally, all input activations arriving via the mesh bus 626 for a particular tile must be committed to one or more narrow memory units of the tile. Further, in the case of a tile configuration with less than four receive ports, the DMAOp instruction may write a zero value to a location within the narrow memory of the tile instead of waiting for data on a non-existent input port. Similarly, in the case of a tile configuration with less than four transmit ports, the DMAOp instruction does not perform a read of the narrow memory associated with the transfer of a non-existent port and a write to the port.

[0084] In some implementations, the location or address of the narrow memory unit to which a particular input activation is written or read is generated by a Tensor Traversal Unit (hereinafter "TTU") based on the receive / transmit DMAOp provided via the mesh bus 626. The receive DMAOp and the transmit DMAOp can be executed simultaneously, and any necessary synchronization will be managed through a synchronization flag control method managed by the controller 602. The TTU will be described in more detail below with reference to FIG. 7.

[0085] The ring bus 628 starts from the controller 602 and, like the CSR bus 622 and the instruction bus 624, also provides a communication connection via a bus data path that connects the tiles 612, 614 within the ring back to the controller 602. In various implementations, the ring bus 628 generally connects or couples full-width memory units (described in more detail below with reference to FIG. 7) within all the tiles 612, 614. Thus, the payload width of the ring bus 628 corresponds to the width of the wide memory units arranged within each tile of the tile set 612, 614. As discussed above, the ring bus 628 also includes a bitmap header indicating the tiles that need to consume the payload data including the instructions or parameters communicated via the ring bus 628.

[0086] Regarding the data (i.e., payload) received at a particular tile via the ring bus 628, in response to the receipt of the information, each tile zeros (i.e., clears) the position data indicated by the bitmap header specific to the receiving tile before transferring the data to another tile. Thus, if there is no remaining bit set data in the header bitmap indicating the particular tile that received the payload, the transfer of the payload to another tile stops. The payload data typically refers to activations and weights used by one or more tiles during tensor calculations executed based on the execution of deeply nested loops.

[0087] In some implementations, the controller 602 may be described as being part of the ring bus 628. In one example, for DMAOp instructions executed within a particular tile, the controller 602 can be used to pop data / payload from the ring bus stop and transfer the payload to the ring bus stop of the next tile within the ring. The controller 602 can also cause the payload data to be committed to one or more wide memory units of the tile if such an action is required by an instruction in the bitmap header. The address of one or more wide memory units to which data needs to be written can be generated by a DMAOp instruction within a particular tile.

[0088] In various implementations, each tile of the tile sets 612, 614 can be either a producer of payload data or a consumer of payload data. If a tile is a producer of payload data, the tile reads data from one or more of its wide memory units and multicasts the data via the ring bus 628 for consumption by one or more other tiles. If a tile is a consumer of payload data, the tile receives the data and writes it to one or more wide memory units within the tile and transfers the payload data for consumption at one or more other tiles. With respect to the movement of payload data via the ring bus 628, typically there is only one producer / master of data present on the ring bus 628 at any given time. The execution order of DMAOp instructions in all tiles (e.g., FIFO control scheme) ensures that there is only one producer / master of data present on the ring bus 628 at a given time.

[0089] In some implementations, the controller 602 uses a synchronous flag control architecture to ensure that only one producer / master of payload data exists on the ring bus 628 at a given time. In one example, each write to the ring output by a tile triggers an increment of the corresponding synchronous flag count. The controller 602 can inspect the payload data to determine the number of data chunks or segments that include the payload. Next, the controller 602 monitors tile execution to ensure that the expected number of data segments are transferred and / or consumed by the tiles before another tile is executed in master mode.

[0090] If a local multicast group without overlapping regions is connected via the ring bus 628, an exception occurs to ensure that only one producer / master of data exists on the ring bus 628 at any given time. For example, tile 0 (master) can multicast (i.e., generate data) to the tiles within the tile 0 - tile 3 group, while tile 4 (master) can do the same for the tiles within the tile 4 - tile 7 group. An important requirement of this dual-master multicast methodology is to prevent different multicast groups from being able to reference each other's data packets, as packet duplication can occur and cause one or more data calculation errors.

[0091] As shown in FIG. 6, the controller 602 provides a communication data path that couples or connects the tiles within the tile sets 612, 614 to the I / O 610 and includes several core functions. The core functions of the controller 602 generally include supplying one or more I / O input activations to the tiles within the tile sets 612, 614, supplying one or more input activations and parameters received from the I / O 610 to the tiles, supplying one or more instructions received from the I / O 610 to the tiles, transmitting I / O output activations to the host interface 608, and functioning as a ring stop for the CSR bus 622 and the ring bus 628. As will be described in more detail below, the first tile set 612 and the second tile set 614 each include a plurality of tiles used to perform one or more tensor calculations that are executed based on a deep loop nest composed of an inner loop and an outer loop.

[0092] The system 600 generally operates as follows. The host interface 608 provides the controller 602 with one or more instructions that define the direct memory access operations (DMAOps) to be performed for a given calculation. The descriptors associated with the instructions supplied to the controller 602 will include the information required by the controller to facilitate large dot product calculations associated with multi-dimensional data arrays (tensors). Generally, the controller 602 receives from the host interface 608 the input activations, tile instructions, and model parameters (i.e., weights) for performing tensor calculations for a given layer of the neural network. Next, the controller 602 can multicast the instructions to the tiles 612, 614 in the data flow manner defined by the instructions. As discussed above, the tiles that consume the instructions can then initiate the broadcast of new / subsequent instructions to other tiles based on the bitmap data within the instruction header.

[0093] Regarding the data flow, the input activations and parameters are sent to the tiles of tile sets 612, 614 via the ring bus 628. Each of tiles 612, 614 stores a subset of the input activations necessary to compute a subset of the output activations assigned to that particular tile. The input activations are moved from wide memory to narrow memory by a DMAOp instruction for the tile. The computation within the tile starts when the necessary input activations, parameters / weights, and compute instructions (TTU operations, memory addresses, etc.) become available within the tile. The computation performed within the tile ends when the MAC operator within the tile (described later) has completed all dot product operations defined by the instruction set, and a pre-activation function is applied to the result of the multiplication operation (i.e., the output activation).

[0094] The result of one or more tensor computations includes writing the output activations of the computational layer to the narrow memory unit of the tile that executes the computation. In a particular tensor computation, the output edge activations are transferred to adjacent tiles via the mesh bus 626. When the computation spans multiple layers, it is necessary to transfer the output edge activations to adjacent tiles to compute the output activations of subsequent layers. When the computation of all layers is complete, the DMAOp moves the final activation to the classifier tile 616 via the ring bus 628. The controller 602 then reads the final activation from the classifier tile 616 and executes a DMAOp to move the final activation to the host interface 608. In some implementations, the classifier portion 616 executes the computation of the output layer (i.e., the last layer) of the NN. In other implementations, the output layer of the NN is one of a classification layer, a regression layer, or another layer type generally associated with a neural network.

[0095] FIG. 7 shows an example of a neural network (NN) computing tile 700 that can be used in the ASIC 106 as described with reference to FIG. 1. Generally, the exemplary tile 700 may correspond to any of the tiles within the first tile set 612 and the second tile set 614 described above with reference to FIG. 6. In various implementations, the computing tile 700 may also be referred to or called the computing unit 700. Each computing tile 700 is a built-in computing unit configured to execute instructions independently of other corresponding tiles within the tile sets 612, 614. As briefly described above, each computing tile 700 executes two types of instructions: TensorOp instructions and DMAOp instructions. Generally, since each instruction type includes computational operations associated with deep loop nests, each instruction type is typically executed over multiple time epochs to ensure completion of all loop iterations.

[0096] As will be described in more detail below, different instruction types are executed by independent control units within the computing tile 700 that synchronize data through synchronization flag control managed within the computing tile 700. The synchronization flag control manages the simultaneity between the execution of different instruction types within the computing tile 700. Each computational operation associated with each instruction type is executed in a strict issue order (i.e., first-in, first-out). For the two instruction types, TensorO p and DMAOp, there is no order guarantee between these different instruction types, and each type is treated as a separate control thread by the computing tile 700.

[0097] Regarding the data flow structure, the compute tile 700 generally includes a data path 702 and a data path 705 that each provide a communication path for the data flow entering and exiting the compute tile 700. As described above, the system 600 includes three different data bus structures laid out in a ring configuration, namely the CSR bus 622, the instruction bus 624, and the ring bus 628. Referring to FIG. 7, the data path 705 corresponds to the instruction bus 624, while the data path 702 generally corresponds to one of the CSR bus 622 and the ring bus 628. As shown, the data path 702 includes a ring output 703 that provides an output path for data exiting the compute tile 700 and a ring input 704 that provides an input path for data entering the compute tile 700.

[0098] The compute tile 700 further includes a TensorOp control 706 that includes a TensorOp tensor traversal unit (TTU) 726 and a DMAOp control 708 that includes a DMAOp TTU 728. The TensorOp control 706 generally manages writes to and reads from the TensorOp TTU register 732 and manages the traversal operations executed by the TensorOp TTU 726. Similarly, the DMAOp control 708 generally manages writes to and reads from the DMAOp TTU register 734 and manages the traversal operations executed by the DMAOp TTU 728. The TTU register 732 includes an instruction buffer for storing one or more instructions including the operations executed by the TensorOp TTU 726 when an instruction by the TensorOp control 706 is executed. Similarly, the TTU register 734 includes an instruction buffer for storing one or more instructions including the operations executed by the TTU 708 when an instruction by the DMAOp control 708 is executed. As further described below, the TTU is generally used by the compute tile 700 to traverse array elements of one or more tensors residing in the narrow memory 710 and the wide memory 712.

[0099] In some implementations, specific instructions executed by compute tile 700 arrive at the tile via data path 705 (i.e., part of instruction bus 624). Compute tile 700 examines the header bitmap to determine the instruction type (TensorOp or DMAOp) and instruction subtype (read operation or write operation). Instructions received by compute tile 700 are then written to a specific instruction buffer according to the instruction type. Generally, instructions are received and stored (i.e., written to a buffer) before execution of the instructions by components of compute tile 700. As shown in FIG. 7, the instruction buffers (i.e., TensorOp TTU register 732 and DMAOp TTU register 734) may each include a first-in first-out (FIFO) control scheme that prioritizes consumption (execution) of one or more related instructions.

[0100] As briefly described above, a tensor is a multi-dimensional geometric object, examples of which include matrices and data arrays. Algorithms that include deeply nested loops are executed by compute tile 700 and can perform tensor computations by iterating through one or more nested loops to traverse an N-dimensional tensor. In one example of the computation process, each loop of the loop nest may be responsible for traversing a particular dimension of the N-dimensional tensor. As described herein, TensorOp control 706 generally manages one or more tensor operations that drive a sequence in which the dimensional elements of a particular tensor structure are traversed and accessed to complete computations defined by deeply nested loops.

[0101] The compute tile 700 further includes a narrow memory 710 and a wide memory 712. The narrow and wide designations generally refer to the size (bits / bytes) of the width of the memory units of the narrow memory 710 and the wide memory 712. In some implementations, the narrow memory 710 includes memory units each having a size or width of less than 16 bits, and the wide memory 712 includes memory units each having a size or width of less than 32 bits. Generally, the compute tile 700 receives input activations via the data path 705, and the DMA control 708 performs an operation to write the input activations to the narrow memory 710. Similarly, the compute tile 700 receives parameters (weights) via the data path 702, and the DMA control 708 performs an operation to write the parameters to the wide memory 712. In some implementations, the narrow memory 710 can include a memory arbiter typically used in shared memory systems, which determines, for each memory cycle, which control device (e.g., TensorOp control 706 or DMAOp control 708) is permitted to access that shared memory unit of the narrow memory 710.

[0102] The computing tile 700 further includes an input activation bus 716 and a MAC array 714 including a plurality of cells each including a MAC operator 715 and a sum register 720. Generally, the MAC array 714 uses the MAC operators 715 and sum registers 720 across the plurality of cells to perform tensor calculations including arithmetic operations related to dot product calculations. The input activation bus 716 provides a data path through which input activations are provided one by one by the narrow memory 710 for each respective access by each MAC operator 715 of the MAC array 714. Thus, based on one-by-one broadcasting of the input activations, a single MAC operator 715 of a particular cell will receive each input activation. The arithmetic operations performed by the MAC operators of the MAC array 714 generally include multiplying the input activations provided by the narrow memory 710 by the parameters accessed from the wide memory 712 to generate a single output activation value.

[0103] During the arithmetic operations, partial sums are accumulated and can be stored in the corresponding, for example, sum register 720, or written to the wide memory 712 and re-accessed by a particular cell of the MAC array 714 to complete subsequent multiplication operations. The tensor calculation can be described as having a first part and a second part. The first part is completed when the output activation is generated by the multiplication operation. For example, it is performed by completing the multiplication of the input activation and the parameters that generate the output activation. The second part includes applying a non-linear function to the output activation, and the second part is completed when the output activation is written to the narrow memory 710 after the application of the function.

[0104] The compute tile 700 further includes an output activation bus 718, a non-linear unit (NLU) 722 with an output activation pipeline 724, an NLU control 738, and a reference map 730 indicating core attributes of components within the compute tile 700. Although the reference map 730 is shown for clarity, it is not included in the compute tile 700. Core attributes include whether a particular component is a unit, a storage device, an operator, a control device, or a data path. Generally, when the first part of the tensor calculation is complete, the output activation is provided from the MAC array 714 to the NLU 722 via the output activation bus 718. After arriving at the NLU 722, data specifying the activation function received via the activation pipeline 724 is applied to the output activation, and the output activation is written to the narrow memory 710. In some implementations, the output activation bus 718 includes at least one pipelined shift register 736, and completing the second part of the tensor calculation includes shifting the output activation towards the narrow memory 710 using the shift register 736 of the activation bus 718.

[0105] For example, for the dot product calculation of two multi-dimensional data arrays, for a single compute tile 700, the MAC array 714 provides a robust single instruction multiple data (SIMD) function. SIMD generally means that all parallel units (multiple MAC operators 715) share the same instruction (based on a deep loop nest), but each MAC operator 715 executes the instruction on different data elements. In one basic example, to add the arrays [1, 2, 3, 4] and [5, 6, 7, 8] element by element to obtain the array [6, 8, 10, 12] in one cycle, usually four arithmetic units are required to execute the operation for each element. By using SIMD, the four units can share the same instruction (e.g., "add") and perform parallel calculations. Thus, the system 600 and the compute tile 700 provide accelerated and parallel processing of tensor calculations that are enhanced compared to conventional methods.

[0106] In one example, as described in more detail below, a single instruction can be provided by the controller 602 to the plurality of compute tiles 700 for consumption by the plurality of MAC arrays 714 (see tile sets 612, 614 of FIG. 6). Generally, a neural network layer can include a plurality of output neurons, and the output neurons can be partitioned such that tensor computations associated with a subset of the output neurons can be assigned to specific tiles of the tile sets 612, 614. Each tile of the tile sets 612, 614 can then perform the associated tensor computations for different groups of neurons of a given layer. Thus, the compute tiles 700 can provide at least two forms of parallel processing. 1) One form includes partitioning output activations (corresponding to a subset of the output neurons) across the plurality of tiles of the tile sets 612, 614. 2) Another form includes simultaneous computation (by a single instruction) of multiple subsets of the output neurons based on the partitioning across the tiles of the tile sets 612, 614.

[0107] FIG. 8 shows an example of a hierarchical execution graph. The illustrated hierarchical execution graph 800 shows a higher-level graph that may be similar to the execution graph 200 or 300 described with respect to FIGS. 2 and 3. In some implementations, the execution graph 800 is generated by a process similar to the process 400 or process 500 described with reference to FIGS. 4 and 5.

[0108] The root graph 802 describes the entire execution graph 800. This includes instructions at the execution level, multi-chip level instructions, single-chip level instructions, a list and parameters defining the hardware layout to be executed, a list of constant buffers that can define variables for the memory space for inference, and a list of subgraphs 804.

[0109] Each sub-graph 804 included in the root graph 802 includes a list 808 of tensors and a list 806 of operations. The sub-graph 804 further includes a list of inputs and outputs, and an index defining tensor parameters and storage locations. Further, the sub-graph 804 can include additional sub-graphs (not shown) that can provide additional layers to the hierarchical execution graph 800.

[0110] The operations 806 included within the sub-graph 804 can be compiled executable files having specific type definitions and additional data specific to the operation. The operations 806 can include metadata specifying a preferred core type, or other parameters related to its execution. Similar to the sub-graph, the operations can include a list of inputs and outputs that can include an index identifying the locations of various tensors or other data required for the operation to be performed. In some implementations, a tensor is a multi-dimensional array of elements, and all elements are of a single known data type.

[0111] The tensors 808 define the data that is potentially captured and processed by the various sub-graphs 804 and operations 808, and further potentially by the root graph 802. Each tensor 808 can have a predefined dimension or shape, as well as a predefined variable type. In some implementations, the tensors are stored in shared memory and can be accessed from multiple cores or computing devices.

[0112] The foregoing description has been provided in relation to one or more particular implementations. Various modifications, changes, and substitutions of the disclosed implementations can be made without departing from the scope of the present disclosure. Accordingly, the present disclosure is not intended to be limited to only the described or illustrated implementations, but the broadest scope consistent with the principles and features disclosed herein should be given.

[0113] Although this specification contains many details of particular implementations, these should not be construed as limitations on the scope of any invention or of what may be claimed, but rather as descriptions of features that may be specific to particular embodiments of a particular invention. The particular features described herein in the context of separate embodiments may also be implemented in combination within a single embodiment. Conversely, the various features described in the context of a single embodiment may also be implemented separately in multiple embodiments or in any suitable subcombination. Further, even if features are described above as acting in a particular combination and are initially claimed as such, one or more features from the claimed combination may in some cases be excluded from the combination, and the claimed combination may be directed to a subcombination or variation of a subcombination.

[0114] Similarly, although the operations are shown in the drawings in a particular order, this should not be understood as requiring that such operations be performed in the particular order shown or in a sequential order, or that all of the illustrated operations be performed, to achieve desirable results. In certain circumstances, multitasking and parallel processing may be advantageous. Further, the separation of the various system modules and components in the embodiments described above should not be understood as requiring such separation in all embodiments, and it should be understood that the program components and systems described may generally be integrated into a single software product or packaged into multiple software products.

[0115] Particular embodiments of the subject matter have been described. Other embodiments are within the scope of the following claims. For example, although a bus line has been described as "controllable", not all bus lines need to have the same level of control. For example, the degree of controllability may vary if some bus lines can only be controlled when they are restricted with respect to the number of tiles that are the source or destination of data. In another example, some bus lines can be specialized to provide data along a single direction such as north, east, west, or south as described herein. In some cases, desirable results can be achieved even if the actions recited in the claims are performed in a different order. As an example, the processes shown in the accompanying figures do not necessarily require the particular order, or series of orders, shown to achieve desirable results. In certain implementations, multitasking and parallel processing may be advantageous.

Claims

1. A method for distributing jobs executable in an array of multi-core computing devices, comprising: receiving a plurality of jobs to be executed in the array of multi-core computing devices, each multi-core computing device including a plurality of different types of processing cores; the method further comprising: allocating each particular job of the plurality of jobs to be executed by one of the plurality of different types of processing cores, the allocating comprising: analyzing the particular job to determine which of the plurality of different types of processing cores is suitable for execution of the particular job; and allocating the particular job to a core type based on the analysis; the method further comprising: compiling each job of the plurality of jobs into an individually executable file; and generating an execution graph representing a mapping of the individually executable files to a particular one of the plurality of different types of processing cores, the execution graph identifying dependencies between the individually executable files.

2. The method of claim 1, further comprising executing the individually executable files, the executing comprising: receiving the execution graph; allocating jobs within the execution graph to a plurality of multi-core computing devices within the array of multi-core computing devices; executing the allocated jobs by each multi-core computing device; returning the output of the executed jobs by each multi-core computing device to a shared memory; and combining the returned outputs to generate a return of the execution graph.

3. The method of claim 1 or claim 2, wherein analyzing the particular job is accomplished using heuristic analysis.

4. The method of claim 1, claim 2, or claim 3, wherein the depth of analysis of each particular job is selected based on user input prior to compile time.

5. The method according to any one of claims 1 to 4, wherein the plurality of different types of processing cores includes a first core type and a second core type.

6. The method according to claim 5, wherein the first core type is a core processor and the second core type is a TPU tile processor.

7. The execution graph has a hierarchical nature including subgraphs arranged in at least four layers, and the at least four layers are a TPU layer including an executable file executed by the second core type, a chip-level layer including one or more subgraphs of the TPU layer and an executable file executed by the first core type, a multi-chip layer including two or more chip-level subgraphs, and a host-level layer including a subgraph of the multi-chip layer and one or more subgraphs configured to be executed by a third type of core, the method according to claim 5 or claim 6.

8. The method according to claim 7, wherein the third type of core is a CPU of a host device.

9. A computer program storing instructions, which, when executed by one or more processors, cause the one or more processors to execute the method according to any one of claims 1 to 8.

10. One or more computers, and a computer-readable storage device coupled to the one or more computers and storing instructions, which, when executed by the one or more computers, cause the one or more computers to execute the method according to any one of claims 1 to 8.

11. A method for compiling an executable job for execution in an array of multi-core computing devices in combination with hardware including a processing core of a first core type, wherein each multi-core computing device in the array of multi-core computing devices includes processing cores of a second core type and a third core type, and the method includes receiving a machine learning model used for inference, analyzing the machine learning model to determine a plurality of jobs to be executed, generating an execution graph representing each of the plurality of jobs to be executed and the dependencies between the plurality of jobs to be executed, and calling a multi-chip level compiler to generate a mapped execution graph, wherein the multi-chip level compiler Identify one or more first jobs among the plurality of jobs executed by the first core type, Compile the one or more first jobs into an executable file to be executed by the first core type, where the first jobs are incompatible with the multi-core computing devices in the array, The multi-chip level compiler further, Split the remaining jobs of the execution graph into a plurality of first subgraphs, Assign each first subgraph to a specific multi-core computing device of the array of multi-core computing devices, For each first subgraph, call a single-chip level compiler, where the single-chip level compiler, Identify one or more chip-level jobs to be executed by the second core type from the first subgraph, Compile each of the one or more chip-level jobs from the first subgraph into an executable file to be executed by the second core type, Split the remaining jobs of the first subgraph into a plurality of second subgraphs, Assign each of the plurality of second subgraphs to the third core type, For each of the plurality of second subgraphs, call a core-level compiler, where the core-level compiler, Compile each of the second subgraphs into an executable file to be executed by the third core type, a method. **Claim 12** The method according to claim 11, wherein the first core type is a host system CPU, the second core type is a core processor, and the third core type is a TPU tile core. **Claim 13** The method according to claim 11 or claim 12, wherein each multi-core computing device of the array of multi-core computing devices is an application-specific integrated circuit (ASIC) including a TPU. **Claim 14** The method according to claim 11, claim 12, or claim 13, wherein identifying one or more first jobs and identifying one or more chip-level jobs are performed based on a heuristic analysis of the plurality of jobs to be executed. **Claim 15** The method according to claim 14, wherein the depth of the heuristic analysis for each specific job is selected based on user input prior to the compilation time. **Claim 16** Receiving, by the array of the multi-core computing devices, the mapped execution graph including the first job, and the plurality of first sub-graphs including the one or more chip-level jobs and the plurality of second sub-graphs, Assigning jobs within the plurality of second sub-graphs to associated cores within the array of the multi-core computing devices, Executing, by each core of the multi-core computing devices, the assigned jobs, Returning, by each multi-core computing device, the output of the executed jobs to a shared memory, Combining the returned outputs to generate a return of the execution graph, the method according to any one of claims 11 to 15. **Claim 17** A computer program storing instructions which, when executed by one or more processors, cause the one or more processors to execute the method according to any one of claims 11 to 16. **Claim 18** A system, One or more computers, and A computer-readable storage device coupled to the one or more computers and storing instructions which, when executed by the one or more computers, cause the one or more computers to execute the method according to any one of claims 11 to 16.

Citation Information

Patent Citations

  • Compile-time scheduling

    US11003429B1

  • Control of scheduling dependencies by a neural network compiler

    US20190391796A1

  • Computation graph mapping in heterogeneous computer system

    US20200342286A1

  • Heterogeneous scheduling for sequential compute dag

    WO2020052241A1