Propagation delay reduction

By adjusting the operational scheduling in the machine learning accelerator, the calculation and propagation delay problems between tiles are solved, performance is improved and efficient resource utilization is achieved.

CN114026543BActive Publication Date: 2025-07-25GOOGLE LLC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202080047574.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-08-22
Filing Date
2020-08-20
Publication Date
2025-07-25
Estimated Expiration
2040-08-20

AI Technical Summary

Technical Problem

There are problems with computational delay and propagation delay in machine learning accelerators, especially performance bottlenecks caused by data transfer and computational dependencies between tiles.

Method used

By adjusting operational scheduling, especially changing the switching of row and column priorities of matrix operations, the calculation and data transmission order between tiles are optimized to hide propagation delays and improve utilization.

Benefits of technology

Reduces computation and propagation latency in machine learning accelerators, improves performance without the need for expensive or complex hardware improvements, achieving nearly 100% utilization in a single tile case.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114026543B_ABST
    Figure CN114026543B_ABST
Patent Text Reader

Abstract

Methods, systems, and apparatus for scheduling operations to reduce propagation latency between tiles of an accelerator, including computer programs encoded on a computer storage medium. One method includes: receiving a request to generate a schedule for a first layer of a program to be executed by an accelerator configured to perform matrix operations at least partially in parallel, where the program defines a plurality of layers including the first layer, and each layer of the program defines a matrix operation to be performed using a corresponding value matrix. Allocating a plurality of initial blocks of the schedule according to an initial allocation direction. Starting to switch the allocation direction at a specific period such that blocks processed after the selected specific period are processed along a different second dimension of a first matrix. Then allocating all remaining unallocated blocks according to the switched allocation direction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to machine learning accelerators. Background Art

[0002] A machine learning accelerator is an application specific integrated circuit (ASIC) designed to perform highly parallel synchronous operations. Parallelism is achieved by integrating many different independent processing elements that can execute concurrently.

[0003] Such a device is well suited for accelerating inference through neural networks. A neural network is a machine learning model that uses multiple layers of operations to predict one or more outputs from one or more inputs. A neural network typically includes one or more hidden layers between an input layer and an output layer. The output of each layer is used as the input to another layer in the network (e.g., the next hidden layer or the output layer).

[0004] Typically, the computational operations required for each layer can be implemented by performing matrix multiplication. Typically, one of the matrices is a vector, e.g., matrix-vector multiplication. A machine learning accelerator thus allows the multiplication and addition of matrix multiplication to be performed with high parallelism.

[0005] However, due to the dependencies between the layers of a neural network, there is an inherent latency in these computational mechanisms. The latency occurs because the output of one layer becomes the input of the next layer. Thus, the layers of a neural network typically must be executed sequentially rather than in parallel. In other words, typically the last computational operation of one layer must be completed before the first computation of the next layer begins.

[0006] Two types of latency typically occur in a machine learning accelerator that uses multiple tiles assigned to different respective layers. First, computational latency occurs because the chip components wait for input data when they are actually available for performing computations. Second, propagation latency occurs because the output of one layer computed by one tile needs to be propagated to the input of another layer computed by a second tile. Computational latency can be improved by manufacturing a larger device with more computational elements. However, propagation latency tends to increase as the device gets larger because the distance the data needs to travel between the tiles also gets larger. Summary of the Invention

[0007] This specification describes how a system generates a schedule for a machine learning accelerator that reduces the computational latency and propagation latency between tiles in the machine learning accelerator.

[0008] Specific embodiments of the subject matter described in this specification may be implemented to achieve one or more of the following advantages. The computational latency and propagation latency of a machine learning accelerator can be reduced by modifying the scheduling of operations. This results in improved performance without the need for expensive or complex hardware changes. The performance improvement of the scheduling techniques described below also provides a computational advantage when there is only one tile. In this case, even though there are inherent computational dependencies, some scheduling can achieve utilization close to 100%.

[0009] Details of one or more embodiments of the subject matter of this specification are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages of the subject matter will become apparent from the description, the drawings, and the claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] Figure 1A Shows how changing the schedule reduces the latency between two layers of a neural network.

[0011] Figure 1B Shows the schedule assignment for a single tile.

[0012] Figure 2 Is a flowchart of an example process for generating a schedule that reduces the latency between tiles of an accelerator.

[0013] Figure 3A Shows performing row-major order and then switching to column-major order.

[0014] Figure 3B Shows performing row-major order with a row limit.

[0015] Figure 4 Shows diagonal scheduling.

[0016] Figure 5 Is a schematic diagram showing an example of dedicated logic circuitry.

[0017] Figure 6 Shows an example of a tile used in an ASIC chip.

[0018] Like reference numerals and names in the various figures indicate like elements. DETAILED DESCRIPTION

[0019] This specification describes techniques for scheduling tile operations to reduce the propagation latency between tiles of a multi-tile accelerator (e.g., a machine learning accelerator).

[0020] In this specification, a tile is a device having an array of computing units that can perform computations on a portion of a matrix. Thus, a tile is any suitable accelerator configured to perform matrix-vector multiplication as a fixed-size block. Each unit may include circuitry that allows the unit to perform mathematical or other computations. In a typical scenario, a tile receives an input vector, multiplies the input vector by a weight matrix using the computing array, and generates an output vector.

[0021] In this specification, a schedule is a temporal sequence of the portions of the matrix on which a particular tile should operate. In this specification, such a discrete portion of the matrix will also be referred to as a block. Thus, a schedule specifies the ordering of the blocks for a particular tile.

[0022] Each time a tile operates on a different block of the matrix can be referred to as an iteration of the schedule. If the matrix fits entirely within the computing array of a tile, all matrix operations can be performed without any schedule. However, when the matrix is larger than the computing array, the system can generate a schedule that specifies in which order the different blocks of the matrix should be processed. For convenience, the operations of a schedule in this specification will be referred to as being assigned to particular identifiable clock cycles. However, these clock cycles do not need to correspond to actual hardware clock cycles, and the same techniques can be used to assign computations to time periods that include multiple hardware clock cycles.

[0023] Figure 1A Shows how changing the schedule reduces the latency between two layers of a neural network. Figure 1A The left-hand side of [FIGURE] shows a simple and straightforward schedule in which two tiles are used to perform the operations of two neural network layers. However, this simple and straightforward schedule has a latency that can be reduced by using the Figure 1A enhanced schedule on the right-hand side of [FIGURE].

[0024] The first layer 102 has a first weight matrix M1 110. The operations of the first layer 102 include receiving an input vector V1 115 and multiplying the input vector 115 by the first weight matrix 110 to generate an output vector V2 117.

[0025] In this example, the first weight matrix 110 is larger than the computing array of the first tile assigned to perform the operations of the first layer 102. The first weight matrix 110 is twice the width and twice the height of the computing array of the first tile. Thus, the operations of the first layer must be performed in multiple blocks over multiple clock cycles according to a particular schedule.

[0026] In Figure 1AIn the example, the first scheduling 106 assigns a row-major scheduling to the operations of the first layer 102, which means that the first tile assigned to the first layer 102 will operate on two iterations on the upper half of the first matrix 110 and then on two iterations on the lower half of the first matrix 110. In Figure 1A it, the clock cycle assignments are shown on the corresponding matrix blocks. Thus, for the first matrix 110 according to the first scheduling, the first tile will process the upper half of the matrix on cycles 0 and 1 in sequence, and the lower half of the matrix on cycles 2 and 3.

[0027] The output vector 117 of the first layer 102 is then generated by summing the partial results of the respective iterations. Thus, the first half of the output vector 117 includes summing the partial results from cycles 0 and 2. The second half of the output vector 117 includes summing the partial results from cycles 1 and 3.

[0028] The output vector 117 is then propagated through the communication hardware to the second tile assigned to perform the matrix operations of the second layer 104 with the second weight matrix M2 120. In this example, it is assumed that the propagation delay of the accelerator is two clock cycles.

[0029] In this figure, the second layer 104 also has a row-major scheduling according to the first scheduling 106.

[0030] The first tile and the second tile assigned to the first layer 102 and the second layer 104 respectively can execute operations concurrently. However, certain data dependencies are naturally introduced between the layers, and the propagation delay introduces a latency that affects when the operations of the second layer 104 can start.

[0031] Specifically, the upper left block of the second matrix 120 cannot be executed until cycles 0 and 2 have both been executed by the first layer 102. Thus, after cycle 2 of the first layer has been executed, cycles 3 and 4 will be spent propagating the left half of the output vector 117 to the second tile computing the second layer 104. Thus, the earliest time point at which the result of the second layer can be computed is at cycle 5.

[0032] For the same reason, the lower left block of the second matrix 120 of the second layer 104 cannot be executed until cycles 1 and 3 have both been executed on the first layer 102 and until the data has been propagated, which incurs a propagation latency of two cycles. Since cycle 6 has been assigned to the upper right block, the first scheduling 106 assigns the lower left part of the second matrix 120 to be processed starting from cycle 7.

[0033] Thus, Figure 1A shows how the first scheduling 106 results in a total execution time of 8 cycles.

[0034] The second scheduler 108 adjusts the execution order of the first layer 102. The second scheduler 108 assigns column-major ordering to the first layer 102 instead of row-major ordering.

[0035] In other words, the first layer can first operate on the upper left part of the first matrix 110 in cycle 0, and then operate on the lower left part of the first matrix 110 in cycle 1.

[0036] Note that at this time, the operation of the second layer 104 can immediately start processing the upper left block of the second matrix 120. Therefore, after two cycles of propagation delay in cycles 2 and 3, the upper left block of the second matrix 120 can already be processed in cycle 4, and the upper right block of the second matrix 120 can be processed in cycle 5.

[0037] This rearrangement of the row / column ordering of the operations of the first layer 102 reduces the total execution time of the two layers to 7 cycles. In fact, by changing the row / column ordering in the first layer 102, the system can hide a full cycle of propagation delay between the two tiles assigned to operate on the first and second layers. Although this is a simple example, the time savings is still 12.5% for a single pass through layers 102 and 104.

[0038] This technique can be generalized and refined into the problem of choosing two values: (1) a specific cycle M at which to perform the switch in the assignment direction, and (2) a specific cycle T at which to process the "lower left block" of the matrix i . In this specification, the "lower left" block of a matrix refers to the last block of the matrix that needs to be processed before the subsequent layer can start processing the output generated by that layer. Therefore, the "lower left" block can be any corner block of the matrix, or any edge block using the last arriving part of a row or column from the previous layer, depending on the specific arrangement in the schedule.

[0039] For an accelerator with a propagation delay of N cycles between layer n-1 and layer n, and a propagation delay of C cycles between layer n and layer n+1, the system can mitigate the propagation delay by scheduling the lower left block of the matrix in layer n to be processed for at least N cycles from the start of the layer and at least C cycles from the end of the layer.

[0040] Therefore, the enhanced scheduler switches in the assignment direction after the selected cycle M. Generally, M specifies a cycle at or before a specific cycle T i . At cycle M, the scheduler can switch from assigning blocks in row-major order to column-major order, or vice versa. This is because at cycle Ti After that, the tile continues to receive data sufficient to generate further output for the next layer. The techniques described below further describe how to change the scheduling row / column allocation direction to mitigate latency for matrices of any size.

[0041] The same switch in the allocation direction can also reduce latency in a machine learning accelerator with only one tile and little or no propagation latency. For example, assume the device includes only a single tile responsible for computing the results of two layers.

[0042] Figure 1B Shows the scheduling allocation for a single tile that has nine compute elements that process 4×4 matrices on each of two layers.

[0043] The first schedule 107 shows a basic row-major ordering. One problem that can occur is that some compute elements may have nothing to do because they are waiting for the results of other computations to complete.

[0044] At cycle 0, all nine compute elements are successfully put to work on the first two rows of M1 111 and the first element of the third row of M1111. But at cycle 1 of the first schedule 107, only seven of the nine compute elements can be given work. This is because when using row-major scheduling, the top left corner of the second layer cannot be computed until the bottom right corner of the first layer has been processed. Thus, the first result of the second layer 104 cannot be computed until one cycle later.

[0045] Instead, consider a second schedule 109 that uses a switch in the allocation direction. That is, after allocating the first row of the allocation matrix 111, the system can switch to column-major allocation. Thus, the bottom left block of matrix 111 is computed at cycle 0 instead of cycle 1. Then, the operations for the second layer can start immediately at cycle 1 because the bottom left block has been processed at cycle 0.

[0046] As a result, cycle 1 in the second schedule with the switch in the allocation direction can achieve 100% utilization because some elements of the compute array can start working on the second layer operations without waiting for the first layer operations to complete. The same technique can improve utilization across the layers of a neural network.

[0047] Figure 2 Is a flowchart of an example process for generating a schedule that reduces latency for an accelerator. For convenience, the process will be described as being performed by a system of one or more computers located in one or more locations and appropriately programmed in accordance with this specification.

[0048] The system receives a request (210) to generate a schedule for a first layer having a first matrix. The first layer can be one of a plurality of layers defined by an input program that specifies the operations to be performed for each layer. In a device having multiple tiles, each layer can be assigned to a corresponding tile of the device having multiple tiles. Each layer can have its own matrix. For example, the input program can specify the operations of a neural network architecture.

[0049] The system assigns a plurality of initial blocks of the schedule according to an initial assignment direction in a first dimension (220). The assignment direction specifies the first dimension of the matrix along which the iterations of the schedule should be performed. For example, the assignment direction can initially specify row-major ordering or column-major ordering.

[0050] The system selects a period (230) for the bottom-left block. As described above, T i represents the period on which the bottom-left block of the matrix will be executed. Also as described above, the selection of T i and a particular type of schedule can also determine M, which is the period at which the assignment direction switches.

[0051] Generally, regardless of how T i is selected, the latency of T i cycles can be hidden between layer i - 1 and layer i, and the latency of W i x H i - T i cycles can be hidden between layer i and layer i + 1. In other words, the system can select T i in order to trade off between hiding the latency of the i - 1 to i transition and the i to i + 1 transition.

[0052] Some matrices may be large enough such that the propagation latency can be completely hidden. Assume that L i represents the total end-layer latency at the end of layer i, which includes any end computations or activation functions as well as the propagation latency. In order to hide all the latency of layer i, the following inequality must hold:

[0053] W i x H i ≥ L i-1 + L i

[0054] where W i is the matrix width in terms of blocks, and H i is the matrix height in terms of blocks. The block size can be determined by the tile hardware.

[0055] When the condition holds, the system can select T i to be L i-1 .​

[0056] In other words, the system can schedule the blocks such that the bottom left block is executed as soon as possible after the previous layer has finished generating the output required to process the block.

[0057] However, not all matrices are large enough to completely hide the latency between layers. In these cases, the scheduling can introduce idle cycles in order to force waiting for the results to be ready. If there are S i idle cycles after layer i, then the following inequality holds for all valid schedules of layer i:

[0058] W i x H i ≥ max(L i-1 – S i-1 , 0) + max(L i – S i , 0)

[0059] If this inequality holds for a valid schedule, the system can allocate T i as follows:

[0060] T i = max(L i-1 - S i-1 , 0)

[0061] When using this arrangement for idle cycles, the system also programmatically selects the number of idle cycles through each layer in order to minimize the total latency introduced by the idle cycles. To this end, the system can perform an optimization process to select an integer number of idle cycles S k for each layer k such that the following inequality holds:

[0062] W i x H i - max(L i – S i , 0) ≥ 0

[0063] and

[0064] S i-1 ≥ L i-1 + max(L i – S i , 0) - W i x H i

[0065] The system switches the allocation direction such that the blocks processed after a particular block are processed sequentially along the second dimension (240). The choice of M (switching period) depends on the type of scheduling used. An example of choosing M is described in more detail below with reference to Figure 3A - Figure 3B more detailed

[0066] The system allocates all remaining unallocated blocks (250) according to the switched allocation direction. In other words, the system can allocate all unscheduled blocks in the order sorted according to the second dimension.

[0067] Figure 3A - Figure 4 An example schedule using the switched allocation direction is shown. In Figure 3A - Figure 3B it, the numbered arrows represent the blocks of the circuit line designated to be executed in a specific order.

[0068] Figure 3A It shows the execution of row-major order and then switching to column-major order. In other words, the system allocates blocks along the top row for the first processing, then along the second row for the next processing, and so on.

[0069] In this example, cycle M occurs somewhere in the middle of the fourth row of the blocks. Therefore, the system switches in the allocation direction and starts to allocate blocks in column-major order. The system can do this so that the lower left corner of the scheduling matrix is executed at the selected cycle T i In other words, the system calculates the row-major order until the number of unprocessed rows is equal to the difference between the current cycle and T i between.

[0070] Figure 3A The schedule shown results in most of the calculations being spent in the column-major phase. This tends to deliver the output at a very uniform rate and leaves some idle cycles at the end of each column. For example, this may be advantageous for the case when the output of each layer requires additional processing (such as in the case of LSTM).

[0071] Figure 3B It shows the execution of the priority order using row limits. In this example, the row-major phase only processes a limited number of blocks before moving to the next row. In this example schedule, the initial row includes more blocks than the subsequent rows. In some embodiments, the system calculates the row limit by calculating the value N = (T i / H i - 1), where H i is the number of blocks in each column of the matrix. Then, the system can use the upper limit of N for the initial row and the lower limit of N for the subsequent rows.

[0072] Therefore, the period of the lower left block T i in this example is given by two N values and the number of rows in the matrix. In other words, if there are 8 rows in the matrix, floor(N) = 3, and ceiling(N) = 4, then T i = 5x4 + 3x3 - (3 - 1) = 27. The switching cycle M in this case is given by M = 5x4 + 3x3 = 29.

[0073] Figure 3B The scheduling in [reference] eliminates the latency when processing the first few columns and reduces the memory requirements. However, Figure 3B the scheduling in [reference] may be more complex to implement.

[0074] Figure 4 Diagonal scheduling is shown. As shown, during row-major order, each row receives a decreasing number of blocks defined by the slope of the diagonal. In this example, the system selects T by calculating the number of blocks required to fill the upper-left diagonal, and the system can select M = T i , and the system can select M = T i .

[0075] Diagonal scheduling has symmetry between the row-major phase and the column-major phase, but has the disadvantages of the above two schedulings.

[0076] Figure 5 is a schematic diagram showing an example of a dedicated logic circuit (specifically, ASIC 500). ASIC 500 includes a plurality of synchronous processors, which are referred to as tiles for simplicity. For example, ASIC 500 includes tile 502, and one or more of the tiles 502 include dedicated circuits configured to perform synchronous calculations (such as multiplication and addition operations). Specifically, each tile 502 may include a computational array of cells, where each cell is configured to perform a mathematical operation (for example, see the exemplary tile 200 shown in Figure 6 and described herein). In some embodiments, the tiles 502 are arranged in a grid pattern, where the tiles 502 are arranged along a first dimension 501 (e.g., rows) and along a second dimension 503 (e.g., columns). For example, in the example shown in Figure 5 , the tile 502 is divided into four different parts (510a, 510b, 510c, 510d), and each part contains 288 tiles arranged in a grid of 18 tiles longitudinally by 16 tiles transversely. In some embodiments, Figure 5 the ASIC 500 shown in [[reference]] can be understood as including a single systolic unit array that is subdivided / arranged into separate tiles, where each tile includes a subset / subarray of cells, local memory, and bus lines (for example, see Figure 6 ).

[0077] The ASIC 500 also includes a vector processing unit 504. The vector processing unit 504 includes circuitry configured to receive an output from the tile 502 and calculate a vector calculation output value based on the output received from the tile 502. For example, in some embodiments, the vector processing unit 504 includes circuitry (e.g., a multiplication circuit, an adder circuit, a shifter, and / or a memory) configured to perform an accumulation operation on the output received from the tile 502. Alternatively or additionally, the vector processing unit 504 includes circuitry configured to apply a non-linear function to the output of the tile 502. Alternatively or additionally, the vector processing unit 504 generates a normalization value, a pooling value, or both. The vector calculation output of the vector processing unit may be stored in one or more tiles. For example, the vector calculation output may be stored in a memory uniquely associated with the tile 502. Alternatively or additionally, the vector calculation output of the vector processing unit 504 may be transmitted to circuitry external to the ASIC 500, e.g., as an output of a calculation. In some embodiments, the vector processing unit 504 is partitioned such that each segment includes circuitry configured to receive an output from a corresponding set of tiles 502 and calculate a vector calculation output based on the received output. For example, in Figure 5 the example shown, the vector processing unit 504 includes two rows spanning along a first dimension 501, each row including 32 segments 506 arranged in 32 columns. Each segment 506 includes circuitry (e.g., a multiplication circuit, an adder circuit, a shifter, and / or a memory) configured to perform a vector calculation based on the output (e.g., a cumulative sum) from the corresponding column of the tile 502, as explained herein. As Figure 5 shown, the vector processing unit 504 may be located in the middle of the grid of tiles 502. Other positional arrangements of the vector processing unit 504 are possible.

[0078] The ASIC 500 also includes a communication interface 508 (e.g., interfaces 508a, 508b). The communication interface 508 includes one or more sets of a serializer / deserializer (SerDes) interface and a general-purpose input / output (GPIO) interface. The SerDes interface is configured to receive instructions (e.g., instructions for operating the controllable bus lines described below) and / or input data for the ASIC 500, and output data from the ASIC 500 to external circuitry. For example, the SerDes interface may be configured to transmit instructions and / or input data at a data rate of 32 Gbps, 56 Gbps, or any suitable data rate through the set of SerDes interfaces included in the communication interface 508. The GPIO interface is configured to provide an interface for debugging and / or booting. For example, when the ASIC 500 is powered on, the ASIC 500 may run a boot program. If the program fails, an administrator may use the GPIO interface to debug the source of the failure.

[0079] The ASIC 500 also includes a plurality of controllable bus lines configured to transfer data between the communication interface 508, the vector processing unit 504, and the plurality of tiles 502 (see, e.g., Figure 6 ). The controllable bus lines include, for example, wiring extending along a first dimension 501 (e.g., rows) of the grid and a second dimension 503 (e.g., columns) of the grid. A first subset of the controllable bus lines extending along the first dimension 501 may be configured to transfer data in a first direction (e.g., to the Figure 5 right side of). A second subset of the controllable bus lines extending along the first dimension 501 may be configured to transfer data in a second direction (e.g., to the Figure 5 left side of). A first subset of the controllable bus lines extending along the second dimension 503 may be configured to transfer data in a third direction (e.g., to the Figure 5 top of). A second subset of the controllable bus lines extending along the second dimension 503 may be configured to transfer data in a fourth direction (e.g., to the Figure 5 bottom of).

[0080] Each controllable bus line includes a plurality of transmitter elements, such as flip-flops, for transferring data along the line in accordance with a clock signal. Transferring data through the controllable bus lines may include shifting the data from a first transmitter element of the controllable bus line to a second adjacent transmitter element of the controllable bus line at each clock cycle. In some embodiments, the data is transferred through the controllable bus lines at the rising or falling edge of the clock cycle. For example, at a first clock cycle, the data present on a first transmitter element (e.g., flip-flop) of the controllable bus line may be transferred to a second transmitter element (e.g., flip-flop) of the controllable bus line at a second clock cycle. In some embodiments, the transmitter elements may be periodically spaced apart from each other at a fixed distance. For example, in some cases, each controllable bus line includes a plurality of transmitter elements, where each transmitter element is located within or near a corresponding tile 502.

[0081] Each controllable bus line also includes a plurality of multiplexers and / or demultiplexers. The multiplexer / demultiplexer of the controllable bus line is configured to transfer data between the bus line and components of the ASIC chip 500. For example, the multiplexer / demultiplexer of the controllable bus line can be configured to transfer data to and / or from the tile 502, transfer data to and / or from the vector processing unit 504, or transfer data to and / or from the communication interface 508. Transferring data between the tile 502, the vector processing unit 504, and the communication interface can include sending control signals to the multiplexer based on the desired data transfer to occur. The control signals can be stored in registers directly coupled to the multiplexer and / or demultiplexer. Then, the value of the control signal can determine, for example, what data is transferred from a source (e.g., a memory within the tile 502 or the vector processing unit 504) to the controllable bus line, or alternatively, what data is transferred from the controllable bus line to a sink (e.g., a memory within the tile 502 or the vector processing unit 504).

[0082] The controllable bus line is configured to be controlled at a local level such that each tile, vector processing unit, and / or communication interface includes its own set of control elements for manipulating the controllable bus line passing through that tile, vector processing unit, and / or communication interface. For example, each tile, 1D vector processing unit, and communication interface can include a corresponding set of transmitter elements, multiplexers, and / or demultiplexers for controlling data transfer to and from the tile, 1D vector processing unit, and communication interface.

[0083] To minimize the latency associated with the operation of the ASIC 500, the tile 502 and the vector processing unit 504 can be positioned to reduce the distance that data travels between various components. In a particular embodiment, both the tile 502 and the communication interface 508 can be divided into multiple parts, where the tile parts and the communication interface parts are arranged such that the maximum distance that data travels between the tile and the communication interface is reduced. For example, in some embodiments, a first group of tiles 502 can be arranged in a first part on a first side of the communication interface 508, and a second group of tiles 502 can be arranged in a second part on a second side of the communication interface. As a result, compared to a configuration where all tiles 502 are arranged in a single part on one side of the communication interface, the distance from the communication interface to the farthest tile can be halved.

[0084] Alternatively, the tiles can be arranged in a different number of parts (such as four parts). For example, in Figure 5In the example shown, multiple tiles 502 of the ASIC 500 are arranged in multiple sections (510a, 510b, 510c, 510d). Each section includes a similar number of tiles 502 arranged in a grid pattern (e.g., each section may include 256 tiles arranged in 16 rows and 16 columns). The communication interface 508 is also divided into multiple sections: a first communication interface 508a and a second communication interface 508b are arranged on either side of the section of tiles 502. The first communication interface 508a can be coupled to two tile sections 510a, 510c on the left side of the ASIC chip 500 via controllable bus lines. The second communication interface 508b can be coupled to two tile sections 510b, 510d on the right side of the ASIC chip 500 via controllable bus lines. As a result, compared to an arrangement where only a single communication interface is available, the maximum distance that data travels to and / or from the communication interface 508 (and thus also the latency associated with data propagation) can be halved. Other coupling arrangements of the tiles 502 and the communication interface 508 may also reduce data latency. The coupling arrangement of the tiles 502 and the communication interface 508 can be programmed by providing control signals to the transmitter elements and multiplexers of the controllable bus lines.

[0085] In some embodiments, one or more tiles 502 are configured to initiate read and write operations regarding the controllable bus lines and / or other tiles within the ASIC 500 (referred to herein as "control tiles"). The remaining tiles within the ASIC 500 can be configured to perform computations (e.g., compute layer inferences) based on input data. In some embodiments, the control tiles include the same components and configurations as other tiles within the ASIC 500. The control tiles can be added as one or more additional tiles, one or more additional rows, or one or more additional columns of the ASIC 500. For example, for a symmetric grid of tiles 502 (where each tile 502 is configured to perform computations on input data), one or more additional rows of control tiles can be included to handle the read and write operations for the tiles 502 that perform computations on input data. For example, each section includes 18 rows of tiles, where the last two rows of tiles can include control tiles. In some embodiments, providing separate control tiles increases the amount of available memory in the other tiles used for performing computations. However, separate tiles dedicated to providing control as described herein are not required, and in some cases, no separate control tiles are provided. Instead, each tile can store instructions in its local memory for initiating read and write operations for that tile.

[0086] In addition, although Figure 5Each of the sections shown includes tiles arranged in 18 rows by 16 columns, but the number of tiles 502 and their arrangement within a section can be different. For example, in some cases, a section can include an equal number of rows and columns.

[0087] In addition, although shown in Figure 5 as divided into four sections, the tiles 502 can be divided into other different groupings. For example, in some embodiments, the tiles 502 are grouped into two different sections, such as a first section above the vector processing unit 504 (e.g., closer to Figure 5 the top of the page shown) and a second section below the vector processing unit 504 (e.g., closer to Figure 5 the bottom of the page shown). In such an arrangement, each section can contain, for example, 576 tiles arranged in a grid of 18 tiles longitudinally (along direction 503) by 32 tiles transversely (along direction 501). A section can contain other total numbers of tiles and can be arranged in arrays of different sizes. In some cases, the division between sections is depicted by the hardware characteristics of the ASIC 500. For example, as shown in Figure 5 , sections 510a, 510b can be separated from sections 510c, 510d by the vector processing unit 504.

[0088] Latency can also be reduced by centering the vector processing unit 504 relative to the tile sections. In some embodiments, the first half of the tiles 502 is arranged on a first side of the vector processing unit 504, and the second half of the tiles 502 is arranged on a second side of the vector processing unit 504.

[0089] For example, in the ASIC chip 500 shown in Figure 5 , the vector processing unit 504 includes two sections (e.g., two rows), each section including a plurality of segments 506 that match the number of columns of the tiles 502. Each segment 506 can be positioned and configured to receive outputs, such as cumulative sums, from the corresponding columns of the tiles 502 within a section of the tiles. In Figure 5In the example shown, tile portions 510a, 510b located on the first side of vector processing unit 504 (e.g., above vector processing unit 504) can be coupled to the top row of segment 506 via controllable bus lines. Tile portions 510c, 510d located on the second side of vector processing unit 504 (e.g., below vector processing unit 504) can be coupled to the bottom row of segment 506 via controllable bus lines. Additionally, each tile 502 within the first half above processing unit 504 can be located at the same distance from vector processing unit 504 as the corresponding tile 502 within the second half below processing unit 504, such that there is no difference in the total latency between the two halves. For example, a tile 502 in row i of the first portion 510a (where variable i corresponds to the row position) can be located at the same distance from vector processing unit 504 as a tile 502 in row m - 1 - i of the second portion of tiles (e.g., portion 510c) (where m represents the total number of rows in each portion and assuming the rows increment in the same direction in both portions).

[0090] Configuring the tile portions in this manner can halve the distance that data travels to and / or from vector processing unit 504 (and thus the latency associated with data propagation) compared to an arrangement where vector processing unit 504 is located at the distal end (e.g., bottom) of all tiles 502. For example, the latency associated with receiving a cumulative sum from portion 510a through a column of tiles 502 can be half the latency associated with receiving a cumulative sum from portions 510a and 510c through a column of tiles 502. The coupling arrangement of tiles 502 and vector processing unit 504 can be programmed by providing control signals to the transmitter elements and multiplexers of the controllable bus lines.

[0091] During operation of ASIC chip 500, activation inputs can be shifted between tiles. For example, activation inputs can be shifted along the first dimension 501. Additionally, the outputs from computations performed by tiles 502 (e.g., the outputs from computations performed by the computation arrays within tiles 502) can be shifted between tiles along the second dimension 503.

[0092] In some embodiments, the controllable bus lines may be physically hardwired such that data skips over tiles 502, thereby reducing the latency associated with the operation of the ASIC chip 500. For example, the output of a computation performed by a first tile 502 may be shifted along the second dimension 503 of the grid to a second tile 502 located at least one tile away from the first tile 502, thereby skipping the intervening tiles. In another example, an activation input from a first tile 502 may be shifted along the first dimension 501 of the grid to a second tile 502 located at least one tile away from the first tile 502, thereby skipping the intervening tiles. By skipping at least one tile when shifting activation inputs or output data, the total data path length can be reduced such that the data is transferred more quickly (e.g., without storing the data at the skipped tiles using clock cycles), and the latency is reduced.

[0093] In an example embodiment, each tile 502 within each column of section 510a may be configured to transfer output data along the second dimension 503 towards the vector processing unit 504 via the controllable bus lines. The tiles 502 within each column may also be configured to transfer data towards the vector processing unit 504 by skipping the next adjacent tile (e.g., via physical hardwiring of the controllable bus lines between the tiles). That is, the tile 502 at position (i,j)=(0,0) in the first section 510a (where variable i corresponds to the row position and variable j corresponds to the column position) may be hardwired to transfer the output data to the tile 502 at position (i,j)=(2,0); similarly, the tile 502 at position (i,j)=(2,0) in the first section 510a may be hardwired to transfer the output data to the tile 502 at position (i,j)=(4,0), and so on. The last tile not skipped (e.g., the tile 502 located at position (i,j)=(16,0)) transfers the output data to the vector processing unit 504. For a section with 18 rows of tiles, such as Figure 5 the example shown, tile skipping ensures that all tiles within the section are at most 9 "tile hops" away from the vector processing unit 504, thereby improving the performance of the ASIC chip 500 by halving the data path length and the resulting data latency.

[0094] In another example embodiment, each tile 502 within each row of sections 510a, 510c and within each row of sections 510b, 510d can be configured to pass activation inputs along a first dimension 501 via controllable bus lines. For example, some tiles within sections 510a, 510b, 510c, 510d can be configured to pass activation inputs towards the center of the grid 500 or towards the communication interface 508. The tiles 502 within each row can also be configured to skip adjacent tiles, for example, by hard-wiring controllable bus lines between the tiles. For example, the tile 502 at position (i,j) = (0,0) in the first section 510a (where the variable i corresponds to the row position and the variable j corresponds to the column position) can be configured to pass the activation input to the tile 502 at position (i,j) = (0,2); similarly, the tile 502 at position (i,j) = (0,2) in the first section 510a can be configured to pass the activation input to the tile 502 at position (i,j) = (0,4), and so on. In some cases, the last tile not skipped (e.g., the tile 502 at position (i,j) = (0,14)) does not pass the activation input to another tile.

[0095] Similarly, the skipped tiles can pass the activation input in the opposite direction. For example, the tile 502 at position (i,j) = (0,15) in the first section 510a (where the variable i corresponds to the row position and the variable j corresponds to the column position) can be configured to pass the activation input to the tile 502 at position (i,j) = (0,13); similarly, the tile 502 at position (i,j) = (0,13) in the first section 510a can be configured to pass the activation input to the tile 502 at position (i,j) = (0,11), and so on. In some cases, the last tile not skipped (e.g., the tile 502 at position (i,j) = (0,1)) does not pass the activation input to another tile. By skipping tiles, in some embodiments, the performance of the ASIC chip 500 can be improved by halving the data path length and the resulting data latency.

[0096] As explained herein, in some embodiments, one or more tiles 502 are dedicated to storing control information. That is, the tiles 502 dedicated to storing control information do not participate in performing calculations on input data such as weight inputs and activation inputs. The control information can include, for example, control data for configuring controllable bus lines during operation of the ASIC chip 500 such that data can be moved around the ASIC chip 500. The control data can be provided to the controllable bus lines in the form of control signals for controlling the transmitter elements and multiplexers of the controllable bus lines. The control data specifies whether a particular transmitter element of the controllable bus line will pass data to the next transmitter element of the controllable bus line such that data is transferred between tiles according to a predetermined schedule. The control data additionally specifies whether data is transferred from or to the bus line. For example, the control data can include control signals that direct the multiplexer to transfer data from the bus line to a memory and / or other circuitry within the tile. In another example, the control data can include control signals that direct the multiplexer to transfer data from a memory and / or circuitry within the tile to the bus line. In another example, the control data can include control signals that direct the multiplexer to transfer data between the bus line and the communication interface 508 and / or between the bus line and the vector processing unit 504. Alternatively, as disclosed herein, dedicated control tiles are not used. Instead, in such cases, the local memory of each tile stores the control information for that particular tile.

[0097] Figure 6 An example of a tile 600 used in the ASIC chip 500 is shown. Each tile 600 includes a local memory 602 and a computing array 604 coupled to the memory 602. The local memory 602 includes physical memory located near the computing array 604. The computing array 604 includes a plurality of cells 606. Each cell 606 of the computing array 604 includes circuitry configured to perform calculations (e.g., multiply and accumulate operations) based on data inputs (such as activation inputs and weight inputs) to the cell 606. Each cell can perform calculations (e.g., multiply and accumulate operations) on a one-cycle clock signal. The computing array 604 can have more rows than columns, more columns than rows, or an equal number of columns and rows. For example, in Figure 6 the example shown, the computing array 604 includes 64 cells arranged in 8 rows and 8 columns. Other computing array sizes are possible, such as computing arrays having 16 cells, 32 cells, 128 cells, or 256 cells, etc. Each tile can include the same number of cells and / or the same size computing array. Then, the total number of operations that can be performed in parallel for the ASIC chip depends on the total number of tiles within the chip having the same size computing array. For example, for Figure 5The ASIC chip 500 shown, which contains approximately 1150 tiles, means that approximately 72,000 calculations can be performed in parallel per cycle. Examples of clock speeds that can be used include, but are not limited to, 225 MHz, 500 MHz, 750 MHz, 1 GHz, 1.25 GHz, 1.5 GHz, 1.75 GHz, or 2 GHz. As Figure 6 shown, the computational array 604 of each individual tile is a subset of the larger systolic array of the tile.

[0098] The memory 602 included in the tile 600 can include, for example, random access memory (RAM), such as SRAM. Each memory 602 can be configured to store 1 / n of the total memory associated with Figure 5 the n tiles 502 of the ASIC chip shown. The memory 602 can be provided as a single chip or multiple chips. For example, Figure 6 the memory 602 shown in is provided as four single-port SRAMs, each SRAM coupled to the computational array 604. Alternatively, the memory 602 can be provided as two single-port SRAMs or eight single-port SRAMs and other configurations. After error correction coding, the combined capacity of the memory can be, but is not limited to, for example, 16 kB, 32 kB, 64 kB, or 128 kB. In some embodiments, by locally providing the physical memory 602 to the computational array, the wiring density of the ASIC 500 can be greatly reduced. In an alternative configuration where the memory is centralized within the ASIC 500, as opposed to the local provision described herein, wiring may be required for each bit of memory bandwidth. The total amount of wiring required to cover each tile of the ASIC 500 would far exceed the available space within the ASIC 100. Instead, by providing dedicated memory for each tile, the total amount required to span the area of the ASIC 500 can be significantly reduced.

[0099] The tile 600 also includes controllable bus lines. The controllable bus lines can be classified into multiple different groups. For example, the controllable bus lines can include a first group of general controllable bus lines 610, which are configured to transfer data between tiles in each major direction. That is, the first group of controllable bus lines 610 can include: bus line 610a, which is configured to transfer data along the first dimension 501 of the grid of tiles towards a first direction (referred to as "east" in Figure 6 ); bus line 610b, which is configured to transfer data along the first dimension 101 of the grid of tiles towards a second direction (referred to as "west" in Figure 6 ), where the second direction is opposite to the first direction; bus line 610c, which is configured to transfer data along the second dimension 103 of the grid of tiles towards a third direction (referred to as Figure 6transmit data along a second dimension 103 of the grid of tiles towards a fourth direction (referred to as "north" in Figure 6 what is referred to as "south" in

[0100] where the fourth direction is opposite to the third direction. The general bus line 610 may be configured to carry control data, activation input data, data from and / or to the communication interface, data from and / or to the vector processing unit, and data to be stored and / or used by the tile 600 (e.g., weight input). The tile 600 may include one or more control elements 621 (e.g., flip - flops and multiplexers) for controlling the controllable bus lines and thus routing data to and / or from the tile 600 and / or the memory 602. Figure 6 The controllable bus lines may also include a second set of controllable bus lines, herein referred to as the compute array partial sum bus lines 620. The compute array partial sum bus lines 620 may be configured to carry data outputs from computations performed by the compute array 604. For example, the bus lines 620 may be configured to carry partial sum data obtained from rows in the compute array 604, as Figure 5 shown. In this case, the number of bus lines 620 will match the number of rows in the array 604. For example, for an 8×8 compute array, there will be 8 partial sum bus lines 620, each bus line coupled to the output of a corresponding row in the compute array 604. The compute array output bus lines 620 may also be configured to couple to another tile within the ASIC chip, e.g., as an input to the compute array of another tile within the ASIC chip. For example, the array partial sum bus lines 620 of tile 600 may be configured to receive an input (e.g., partial sum 620a) from the compute array of a second tile located at least one tile away from tile 600. The output of the compute array 604 is then added to the partial sum line 620 to produce a new partial sum 620b, and the partial sum 620b may be output from tile 600. Then, the partial sum 620b may be passed to another tile, or alternatively, passed to the vector processing unit. For example, each bus line 620 may be coupled to a corresponding segment (such as

[0101] segment 506 in Figure 5 As explained with reference to Figure 5For further explanation, the controllable bus line may include circuitry such as a multiplexer configured to allow data to be transmitted between communication interfaces of different tiles, vector processing units, and ASIC chips. The multiplexer can be located anywhere there is a data source or a data slot. For example, in some embodiments, as Figure 6 shown, the control circuit 621 (such as a multiplexer) may be located at the intersection of the controllable bus lines (e.g., the intersection of general bus lines 610a and 610d, the intersection of general bus lines 610a and 610c, the intersection of general bus lines 610b and 610d, and / or the intersection of general bus lines 610b and 610c). The multiplexer at the bus line intersection can be configured to transmit data between the bus lines at the intersection. Accordingly, by appropriate operation of the multiplexer, the direction in which data travels on the controllable bus line can be changed. For example, data traveling along the first dimension 101 on the general bus line 610a can be transmitted to the general bus line 610d so that the data instead travels along the second dimension 103. In some embodiments, the multiplexer may be located near the memory 602 of the tile 600 so that data can be transmitted to and / or from the memory 602.

[0102] Embodiments of the subject matter and the functional operations described in this specification can be implemented in digital electronic circuitry, in tangibly embodied computer software or firmware, in computer hardware including the structures disclosed in this specification and their structural equivalents, or in one or more of their combinations. Embodiments of the subject matter described in this specification can be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a tangible non-transitory storage medium for execution by, or to control the operation of, a data processing apparatus. A computer storage medium may be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or a combination of one or more of them. Alternatively or additionally, the program instructions may be encoded on an artificially generated propagated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal, that is generated to encode information for transmission to a suitable receiver for execution by the data processing apparatus.

[0103] The term "data processing apparatus" refers to data processing hardware and encompasses a variety of devices, equipment, and machines for processing data, including, for example, programmable processors, computers, or multiple processors or computers. The apparatus may also be or further include dedicated logic circuitry, such as, for example, an FPGA (Field Programmable Gate Array) or an ASIC (Application Specific Integrated Circuit). In addition to the hardware, the apparatus may optionally include code that creates an execution environment for a computer program, such as, for example, code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them.

[0104] A computer program (which may also be referred to as or described as a program, software, software application, applet, module, software module, script, or code) can be written in any form of programming language, including compiled or interpreted languages, or declarative or procedural languages, and it can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A program can, but need not, correspond to a file in a file system. A program can be stored in a part of a file that holds other programs or data (such as one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple cooperating files (such as files that store one or more modules, subroutines, or portions of code). A computer program can be deployed to execute on one computer or on multiple computers that are located at one site or distributed across multiple sites and interconnected by a data communication network.

[0105] For a system of one or more computers, being configured to perform particular operations or actions means that the system has software, firmware, hardware, or a combination of them installed on it, which, in operation, cause the system to perform those operations or actions. For one or more computer programs, being configured to perform particular operations or actions means that the one or more programs include instructions that, when executed by a data processing apparatus, cause the apparatus to perform those operations or actions.

[0106] As used in this specification, an "engine" or "software engine" refers to a software-implemented input / output system that provides an output different from the input. An engine can be an encoded functional block, such as a library, platform, software development kit ("SDK"), or object. Each engine can be implemented on any suitable type of computing device, such as, for example, a server, mobile phone, tablet computer, laptop computer, music player, e-book reader, laptop or desktop computer, PDA, smart phone, or other fixed or portable device that includes one or more processors and a computer-readable medium. Additionally, two or more engines can be implemented on the same computing device or on different computing devices.

[0107] The processes and logical flows described in this specification can be performed by one or more programmable computers executing one or more computer programs to perform functions by operating on input data and generating output. These processes and logical flows can also be performed by special-purpose logic circuitry, such as an FPGA or ASIC, or by a combination of special-purpose logic circuitry and one or more programmed computers.

[0108] Computers suitable for executing computer programs can be based on general or special-purpose microprocessors or both, or any other type of central processing unit. Generally, the central processing unit will receive instructions and data from a read-only memory or a random access memory or both. The basic elements of a computer are a central processing unit for executing or running instructions and one or more memory devices for storing instructions and data. The central processing unit and the memory can be supplemented by, or incorporated in, special-purpose logic circuitry. Generally, a computer will also include, or be operatively coupled to, one or more mass storage devices (such as magnetic disks, magneto-optical disks, or optical disks) for storing data, to receive data from it, or to send data to it, or both. However, a computer need not have such devices. In addition, a computer can be embedded in another device, such as a mobile phone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a global positioning system (GPS) receiver, or a portable storage device (such as a universal serial bus (USB) flash drive), to name just a few.

[0109] Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, for example: semiconductor memory devices, such as EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal hard disks or removable disks; magneto-optical disks; and optical CD-ROM and DVD-ROM disks.

[0110] To provide for interaction with a user, embodiments of the subject matter described in this specification may be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user, and a keyboard and a pointing device (e.g., a mouse, a trackball, or a pressure-sensitive display or other surface) by which the user may provide input to the computer. Other types of devices may also be used to provide for interaction with the user; for example, feedback provided to the user may be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback; and input from the user may be received in any form, including voice, speech, or tactile input. Additionally, the computer may interact with the user by sending documents to and receiving documents from the devices the user uses; for example, by sending a web page to a web browser on a user device in response to a request received from the web browser. Further, the computer may interact with the user by sending text messages or other forms of messages to a personal device (e.g., a smart phone), running a messaging application, and receiving a response message from the user in exchange.

[0111] Embodiments of the subject matter described in this specification may be implemented in a computing system that includes a backend component (e.g., as a data server), or includes a middleware component (e.g., an application server), or includes a frontend component (e.g., a client computer having a graphical user interface, a web browser, or an application through which a user may interact with an implementation of the subject matter described in this specification), or any combination of one or more such backend, middleware, or frontend components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN) and a wide area network (WAN), such as the Internet.

[0112] The computing system may include a client and a server. The client and the server are typically remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on respective computers and having a client-server relationship to each other. In some embodiments, the server sends data (e.g., an HTML page) to a user device, e.g., to display data to and receive user input from a user interacting with the device acting as the client. Data generated at the user device (e.g., the result of a user interaction) may be received at the server from the device.

[0113] In addition to the above embodiments, the following embodiments are also innovative:

[0114] Embodiment 1 is a method that includes:

[0115] Receive a request to generate a schedule for a first layer of a program to be executed by an accelerator, the accelerator being configured to perform matrix operations at least partially in parallel, where the program definition includes multiple layers including the first layer, and each layer of the program defines a matrix operation to be performed using a corresponding value matrix;

[0116] Allocate a plurality of initial blocks of the schedule according to an initial allocation direction, where the initial allocation direction specifies a first dimension of a first matrix for the first layer, and the plurality of initial blocks are to be executed along the first dimension;

[0117] Select a specific cycle to process the last block of the matrix required before a subsequent layer can start processing;

[0118] Switch the allocation direction such that the blocks processed after the selected specific cycle are processed along a different second dimension of the first matrix; and

[0119] Allocate all remaining unallocated blocks according to the switched allocation direction.

[0120] Example 2 is the method according to Example 1, where selecting the specific cycle includes:

[0121] Calculate the propagation delay of the previous layer; and

[0122] Allocate the specific cycle based on the propagation delay of the previous layer.

[0123] Example 3 is the method according to any one of Examples 1-2, where selecting the specific cycle includes:

[0124] Calculate the propagation delay of the previous layer;

[0125] Calculate the number of idle cycles of the previous layer; and

[0126] Select the maximum value between the propagation delay of the previous layer and the number of idle cycles of the previous layer.

[0127] Example 4 is the method according to any one of Examples 1-3, where the schedule allocates a plurality of initial blocks in row-major order, and where all remaining unallocated blocks are allocated in column-major order.

[0128] Example 5 is the method according to Example 4, further including selecting a cycle to switch the allocation direction, including: selecting a cycle where the number of unscheduled rows is equal to the difference between the current cycle and the selected specific cycle.

[0129] Example 6 is the method according to Example 4, where the schedule allocates a plurality of initial blocks only along partial rows of the matrix.

[0130] Example 7 is the method according to Example 6, wherein the scheduling allocates a plurality of initial partial rows and a plurality of subsequent partial rows, and the subsequent partial rows are smaller than the initial partial rows.

[0131] Example 8 is the method according to Example 7, wherein the initial partial rows have a length given by ceiling(N), and the subsequent partial rows have a length given by floor(N), where N is given by dividing the selected period by the block height of the matrix on the previous layer.

[0132] Example 9 is the method according to Example 4, wherein the scheduling allocates the initial blocks in row-major order to fill the space defined by the diagonal in the matrix.

[0133] Example 10 is the method according to Example 9, wherein the switching of the allocation direction occurs at a specific selected period.

[0134] Example 11 is the method according to any one of Examples 1-10, wherein the accelerator has a plurality of tiles, and each layer will be calculated by a corresponding tile among the plurality of tiles.

[0135] Example 12 is the method according to any one of Examples 1-10, wherein the accelerator has a single tile to perform the operations of two layers.

[0136] Example 13 is a system, comprising: one or more computers and one or more storage devices storing instructions that, when executed by the one or more computers, are operable to cause the one or more computers to perform the method according to any one of Examples 1 to 12.

[0137] Example 14 is a computer storage medium encoded with a computer program, the program comprising instructions that, when executed by a data processing device, are operable to cause the data processing device to perform the method according to any one of Examples 1 to 12.

[0138] Although this specification contains many specific implementation details, these should not be construed as limitations on the scope of any invention or on the scope of what is claimed, but rather as descriptions of features that may be specific to particular embodiments of a particular invention. Certain features described in the context of separate embodiments in this specification can also be implemented in combination in a single embodiment. Conversely, the various features described in the context of a single embodiment can also be implemented separately in multiple embodiments or in any suitable sub-combination. Additionally, although features may be described as acting in certain combinations and even initially claimed as such, in some cases, one or more features from a claimed combination can be deleted from that combination, and the claimed combination can be directed to a sub-combination or a variant of a sub-combination.

[0139] Similarly, although the operations are described in a particular order in the figures, this should not be understood as requiring that the operations be performed in the particular order shown or in sequential order, or that all of the illustrated operations be performed to obtain the desired result. In some cases, multitasking and parallel processing may be advantageous. Additionally, the separation of various system modules and components in the above-described embodiments should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated in a single software product or packaged into multiple software products.

[0140] Particular embodiments of the subject matter have been described. Other embodiments are within the scope of the following claims. For example, the acts recited in the claims can be performed in a different order and still obtain the desired result. As one example, the processes described in the figures do not necessarily need the particular order or sequential order shown to obtain the desired result. In some cases, multitasking and parallel processing may be advantageous.

Claims

1. A computer-implemented method, comprising: Receiving a request to generate a schedule for a first layer of a program to be executed by an accelerator, the accelerator being configured to perform matrix operations at least partially in parallel, wherein the program defines a plurality of layers including the first layer, and each layer of the program defines a matrix operation to be performed using a corresponding value matrix; Allocating a plurality of initial blocks of the schedule according to an initial allocation direction, wherein the initial allocation direction specifies a first dimension of a first matrix for the first layer, and the plurality of initial blocks are to be executed along the first dimension; Selecting a specific period to process a specific block of the matrix, wherein the specific block needs to be processed before a subsequent layer can start processing; Switching the allocation direction such that blocks processed after the selected specific period are processed along a different second dimension of the first matrix; and Allocating all remaining unallocated blocks according to the switched allocation direction; wherein selecting the specific period includes: Calculating the propagation delay of the previous layer; and Allocating the specific period based on the propagation delay of the previous layer.

2. The method according to claim 1, wherein Selecting the specific period includes: Calculating the propagation delay of the previous layer; Calculating the number of idle cycles of the previous layer; and Selecting the maximum value between the propagation delay of the previous layer and the number of idle cycles of the previous layer.

3. The method according to claim 1, wherein The schedule allocates the plurality of initial blocks in row-major order, and wherein all remaining unallocated blocks are allocated in column-major order.

4. The method according to claim 3 further comprises: Selecting the period to switch the allocation direction includes: selecting a period in which the number of unscheduled rows is equal to the difference between the current period and the selected specific period.

5. The method according to claim 3, wherein, The schedule allocates the plurality of initial blocks only along a portion of the rows of the matrix.

6. The method according to claim 5, wherein, The schedule allocates a plurality of initial partial rows and a plurality of subsequent partial rows, wherein the subsequent partial rows are smaller than the initial partial rows.

7. The method according to claim 6, wherein, The initial partial rows have a length given by ceiling(N), and the subsequent partial rows have a length given by floor(N), where N is given by dividing the selected period by the block height of the matrix on the previous layer.

8. The method according to claim 3, wherein The schedule allocates the initial blocks in row-major order to fill the space defined by the diagonal in the matrix.

9. The method according to claim 8, wherein Switching the allocation direction occurs at the selected specific period.

10. The method according to claim 1, wherein, The accelerator has a plurality of tiles, and each layer is to be computed by a corresponding tile among the plurality of tiles.

11. The method according to claim 1, wherein The accelerator has a single tile to perform operations for two layers.

Citation Information

Patent Citations

  • System, apparatus and method for translating vector instructions

    CN103946797A

  • Prefetching weights for use in a neural network processor

    CN107454966A