Systems and methods for loading operands and outputting results using only a single-sided slave computing array

By loading and outputting operands and results from a single side in the calculation array, using collinear transmission channels and diagonal routing, the problem of array size limitation in the prior art is solved, and the scalability and routing efficiency of the array are improved.

CN114930351BActive Publication Date: 2025-07-18GROQ INC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202080092377.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-11-26
Filing Date
2020-11-25
Publication Date
2025-07-18
Estimated Expiration
2040-11-25

AI Technical Summary

Technical Problem

In existing computing array designs, loading and outputting operands and results from different sides leads to limited size of the computing array and increased wiring length and complexity.

Method used

By loading and outputting all operands and results from a single side of the calculation array, diagonal routing of result values is achieved through collinear weighting and activation of the transmission channel and result output channel, reducing routing complexity and allowing linear scaling of the array.

Benefits of technology

It realizes the scalability and routing efficiency of the computing array, reduces the wiring length and complexity, simplifies the data transmission path, and improves the scaling capability of the computing array.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114930351B_ABST
    Figure CN114930351B_ABST
Patent Text Reader

Abstract

A computing array is implemented in which all operands and results are loaded or output from a single side of the array. The computing array includes a plurality of cells arranged in n rows and m columns, each cell being configured to generate a processed value based on a weight value and an activation value. The cells receive the weights and activation values via collinear weight and activation transmission channels, each of which extends across a first side edge of the computing array to provide the weight values and activation values to the cells of the array. Additionally, result values generated at the top cells of each of the m columns of the array are routed through the array to be output from the same first side edge of the array at the same relative timing at which the result values are generated.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - reference to related applications

[0002] This application claims priority to U.S. Provisional Application No. 62 / 940,818, filed on November 26, 2019, the entire content of which is incorporated herein by reference. Technical Field

[0003] The present disclosure generally relates to computing arrays, and more particularly to routing of inputs and outputs of computing arrays. Background Art

[0004] In many computing arrays, operands and outputs are loaded and output from different sides of the computing array. For example, in many systolic array designs, different operands (e.g., weights and activations) are loaded via two different sides of the array, while the resulting values are output from a third side of the array. However, loading inputs and receiving results via multiple sides of a computing array may limit the size of the computing array relative to the memory and controller circuitry used to operate the computing array, and may increase the length and complexity of the wiring required to route the various inputs and outputs of the computing array. Summary of the Invention

[0005] A computing array is implemented in which all operands and results are loaded or output from a single side of the array. The computing array includes a plurality of cells arranged in n rows and m columns, each cell being configured to produce a processed value based on a weight value and an activation value. The cells receive the weight and activation values via collinear weight and activation transmission channels, each of which extends across a first side edge of the computing array to provide the weight value and the activation value to the cells of the array. Additionally, the result values generated at the top cells of each of the m columns of the array are routed through the array to be output from the same first side edge of the array at the same relative timing at which the result values are generated.

[0006] According to some embodiments, a system is provided that includes a computing array including a plurality of cells arranged in n rows and m columns, each cell being configured to produce a processed value based on a weight value and an activation value. The system further includes at least two collinear transmission channels, the at least two collinear transmission channels including at least a weight transmission channel and an activation transmission channel. The weight transmission channel and the activation transmission channel each extend across a first side edge of the computing array to provide the weight value and the activation value to the cells of the computing array.

[0007] In some embodiments, the computing array is configured to generate a plurality of result values based on the processed values produced by each cell. In some embodiments, the at least two collinear transmission channels further include a result output channel that extends across the first side edge of the computing array and outputs the plurality of result values generated by the computing array.

[0008] In some embodiments, the computing array is configured to: generate at the top cell of each of the m columns of the computing array a result value among a plurality of result values corresponding to the sum of the processing values generated by the cells of the corresponding column of the computing array; and output the m generated result values from a first side of the computing array via a result output channel.

[0009] In some embodiments, the computing array is configured to output the m generated results from a first side of the computing array using routing circuitry implemented in each of at least a portion of the cells of the array. In some embodiments, the routing circuitry is configured to propagate each of the m results through a plurality of cells along the corresponding column until reaching a cell along the diagonal of the computing array within the corresponding column, and to propagate each of the m results along the diagonal of the computing array from the corresponding cell across the m rows of the computing array such that each of the m results is output from the computing array from the first side of the computing array at the same relative timing at which the m results are generated. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] Figure 1 Shows a computing array including an array of multiply-accumulate units (MACC units) where operands are loaded from three different sides.

[0011] Figure 2A Shows a two-dimensional (2D) computing array where operands (e.g., weights and activations) and results are loaded / output from different sides.

[0012] Figure 2B Shows a 2D computing array according to some embodiments where operands and results are loaded / output from a single side.

[0013] Figure 3A Shows a top-level chip macro architecture including a computing array according to some embodiments.

[0014] Figure 3B Shows a high-level diagram of a processor including a memory and a computing array according to some embodiments.

[0015] Figure 4A and Figure 4B Shows loading weights into the cells of a computing array according to some embodiments.

[0016] Figure 5A Shows loading weights onto a computing array according to some embodiments.

[0017] Figure 5B Shows how control signals can be used when loading weights for the cells of a computing array according to some embodiments.

[0018] Figure 5C A diagram showing how control signals can be received by units of a computing array according to some embodiments.

[0019] Figure 6 Shows the propagation of vertical and horizontal portions of control signals through a computing array according to some embodiments.

[0020] Figure 7A Shows a computing array in which weights are loaded only in specific portions of the array according to some embodiments.

[0021] Figure 7B Shows a computing array in which weights are loaded only in specific portions of the array according to some embodiments.

[0022] Figures 8A to 8C Shows the order and timing of weight transmission according to some embodiments.

[0023] Figure 9A Shows a high-level diagram of a computing array in which weights and activations are loaded from different sides according to some embodiments.

[0024] Figure 9B Shows a high-level diagram of a computing array in which weight loading and activation loading are aligned according to some embodiments.

[0025] Figure 10A Shows a diagram of a computing array in which computed result values are calculated according to some embodiments.

[0026] Figure 10B Shows a result path for outputting a total result value at the top unit of each column of the array according to some embodiments.

[0027] Figure 11 Shows a high-level circuit diagram for routing result values to individual units within the array according to some embodiments.

[0028] Figure 12 Shows an example architecture for loading weights and activations into a computing array according to some embodiments.

[0029] For illustrative purposes only, the figures depict various embodiments. Those skilled in the art will readily recognize, based on the following discussion, that alternative embodiments of the structures and methods shown herein can be employed without departing from the principles described herein. Detailed Description

[0030] The accompanying drawings and the following description relate to preferred embodiments by way of illustration only. It should be noted that, based on the following discussion, alternative embodiments of the structures and methods disclosed herein will be readily recognized as viable alternatives that can be employed without departing from the principles claimed.

[0031] Reference will now be made in detail to several embodiments, examples of which are illustrated in the accompanying drawings. Note that, wherever feasible, like or similar reference numerals may be used in the drawings and may indicate like or similar functionality. The drawings depict embodiments of the disclosed system (or method) for illustrative purposes only. Those skilled in the art will readily recognize from the following description that alternative embodiments of the structures and methods shown herein may be employed without departing from the principles described herein.

[0032] Overview

[0033] The specially constructed 2D matrix functional unit (hereinafter referred to as the computational array) described herein loads its operands entirely from only one dimension (1D) (e.g., the left side) of the unit and generates results. The computational array may correspond to a systolic array for matrix multiplication, performing convolution operations, etc. In some embodiments, a computational array is used to implement a machine learning model.

[0034] Most computational array designs use three sides of the array to load operands (weights and activations) and output the generated results (new activations). For example, many systolic array designs load weights and activations into the array from the bottom or top and the left side of the array, respectively, and output the results from the top or bottom of the array. For example, Figure 1 shows a computational array including a multiply-accumulate unit (MACC unit) in which operands are loaded into the array from three different sides. As Figure 1 shown, activation values are loaded from the left side of the array, weight values are loaded from the bottom side of the array, and the results propagate through each column of the array to be output from the top side of the array.

[0035] When loading and outputting operands and results from different sides of the computational array, additional circuitry along different sides of the array is required to load and receive the operands and results, which potentially limits the ability to scale the size of the array. Figure 2A shows a 2D computational array in which operands (e.g., weights and activations) and results are loaded / output from different sides. Figure 2A The computational array of Figure 1 can load operands and output results in a manner similar to that shown in Figure 2A shown. As

[0036] On the other hand, by loading all operands from one side, direct linear scaling of the architecture can be achieved. Figure 2B A 2D computing array is shown in which operands and results are loaded / output from a single side (e.g., the left side of the array). By loading / output from only a single side, the quadratic scaling limitation is removed. For example, the width of adjacent circuitry (including memory, preprocessor, etc.) can be fixed, where only the height is scaled to match the height of the computing array since there is no need to accommodate changes in both dimensions of the computing array.

[0037] Using the techniques presented herein, a design of a 2D computing array is implemented in which all operands and results are loaded or output from a single side of the array. Advantages of single-side loading include scalability of the array. As the size of the computing array (e.g., the number of MACCs in the array) increases, the operand requirements for feeding the array also increase. Thus, when the array grows O(n^2), each side of the array grows O(n).

[0038] Figure 3A A top-level chip macro-architecture including a computing array is shown according to some embodiments. The macro-architecture 100 includes a memory 105 that stores activation inputs and weights to be used by the computing array 110 to generate result values. As Figure 3A shown, the computing array 110 receives activations and weight values from the memory 105 via one or more activation transmission lines and one or more weight transmission lines through a first side of the computing array 110 (shown as the left side in Figure 3A ). Additionally, the computing array 110 outputs result values (e.g., result values to be stored in the memory 105) to the memory 105 via one or more result transmission lines through the first side. In some embodiments, the received result values stored in the memory 105 can be used to determine activation values to be loaded onto the computing array 110 for future computations.

[0039] In some embodiments, the computing array 110 is a matrix multiplication unit (MXM) of an array including multiply-accumulate (MACC) units. In other embodiments, the computing array 110 is a convolution array for performing convolution operations, or other types of arrays.

[0040] In some embodiments, weights from the memory 105 can be temporarily stored in a weight buffer 115 before being loaded into the computing array. The weight buffer 115 is described in more detail below. In some embodiments, the weight buffer 115 can be bypassed such that weight values from the memory 105 are directly loaded onto the computing array 110 via the weight transmission lines, as will be discussed below.

[0041] Although Figure 3AThe compute array 110 is shown as being located near the memory 105, but it should be understood that in some embodiments, the compute array 110 can be separated from the memory 105 by additional components. For example, Figure 3B A high-level diagram of a processor including a memory and a compute array according to some embodiments is shown. The processor can be a tensor streaming processor (TSP) and is organized into multiple functional regions or units, each functional region or unit being configured to perform a specific function on the received data. These functional regions or units can include: a memory unit (MEM) configured to store data; a vector execution module (VXM) unit including arithmetic logic units (ALUs) configured to perform pointwise arithmetic or logical operations on the received data; an MXM unit including an array of MACC units for performing matrix multiplication; and a switching execution module (SXM) unit for allowing data to move between different channels across the processor. The memory 105 can correspond to the MEM unit, and the compute array 110 can correspond to Figure 3B the MXM unit shown in

[0042] In Figure 3B the processor shown in, data operands (including weights and activation values to be used by the compute array, as well as result values output from the compute array) are routed along a first dimension (e.g., in Figure 3BA data channel extending horizontally in any of the directions shown) transmits across the functional regions of the processor, thereby allowing data to be transmitted across the functional regions of the processor. For example, in some embodiments, each functional unit includes an array of cells or tiles. For example, the MEM unit may include an array of memory cells, while the VXM unit and the MXM unit each include an array of ALU and MACC units, respectively. The cells of each functional region of the processor are organized into multiple rows, where the cells of each row are connected by a corresponding data channel that includes multiple wires connecting adjacent cells of the row. The cells of each functional region of the processor can receive data operands via the data channel, perform one or more operations on the received data operands, and output the resulting data onto the data channel for transmission to the cells of subsequent functional regions. For example, in some embodiments, data operands can be read from the memory cells of the MEM unit, transmitted along the data channel (e.g., corresponding to the row of memory cells) for processing by the ALU unit of the VXM unit, and then transmitted to the MACC unit of the computing array (e.g., the MXM unit) for matrix multiplication, the result of which can be transmitted back to the cells of the MEM unit for storage, or transmitted to the cells of another functional unit for future processing. Additionally, the cells of a functional region can receive data operands via the data channel and pass the data operands through without additional processing (e.g., data operands are read from the memory cells of the MEM unit and transmitted across the data channel through the ALU unit of the VXM unit without being processed, for processing at the MXM for matrix multiplication).

[0043] Although Figure 3B the various regions of the processor are shown arranged in a particular order, it should be understood that in other embodiments, the various units of the processor may be arranged differently.

[0044] The ability to load weights and activation values from a single side of the computing array and output result values allows all data operands that are permitted to be transmitted between the memory and the computing array (and / or other functional units on the processor chip) to be transmitted along a data channel extending in a single dimension. In some embodiments, each cell of a functional unit is adjacent (or contiguous) to other cells in its row, and data is transmitted by connecting the adjacent cells of each row to form a collinear data channel for that row. This allows an increased amount of data to be transmitted across the functional units of the processor due to the collinear wiring scheme connecting adjacent cells and reduced congestion.

[0045] For example, the number of signals within and between functional units is limited by the "pitch" (distance between pairs of lines), which determines the line density that can be utilized (e.g., lines / mm). For example, on a chip with a 50 nm pitch, there can be a maximum of 20K lines per mm, or assuming a 50% utilization of the available line space, 10K per mm since it is usually not possible to use every single available line. In some embodiments, each unit of the processor can be approximately 1 mm high, allowing up to approximately 10K signals across each row of the unit. In some embodiments, a single data channel can have: (2 directions) × (138 bits per stream) × 32 streams = 8832 lines, which is less than the 10K / mm calculated above. In a processor chip with 20 rows, this allows (20 data channels) × (8832 lines per data channel) = 176640 lines to operate a 160 TB / s on-chip network capacity at 900 MHz.

[0046] However, routing congestion will consume line resources to connect non-adjacent components. Thus, an adjacent design style allows collinear data flows and minimizes line congestion, thereby allowing for a more efficient utilization of the available underlying ASIC line density (e.g., to achieve the above line density) and minimizing the total line length. For example, a compute array / MXM is configured to receive operand inputs and output result values from the same side of the array (e.g., a flow moving eastward carries operands from memory to the MXM, and a flow moving westward carries the result from the MXM back to memory), thereby allowing the compute array to be connected to other functional areas of the processor (e.g., memory) via parallel lines that do not require turning. However, if the result is produced in the opposite direction (i.e., if the operands and result are received / output on different sides of the MXM), then the signals will need to be routed orthogonally to return the result to the desired memory unit for storage. If the data path has to "turn" and is routed orthogonally, it consumes additional line resources, thereby eroding the available lines for the data path. To avoid this, the on-chip network uses bi-directional flow registers (eastward and westward) to shuttle operands and results across each channel (e.g., to allow the MXM to receive operands and send results across each row's data channels on the same side of the MXM). Other embodiments of the on-chip network can include rings or tori, for example, to interconnect the units of the functional areas and utilize the available line density on the ASIC while minimizing turns and the need for line congestion. In some embodiments (e.g., as Figure 3B shown), the functional units within each data channel are organized to interleave the functional units (e.g., MXM-MEM-VXM-MEM-MXM) to take advantage of this data flow locality between the functional units. This on-chip line density is significantly greater than the available off-chip pin bandwidth for communication between TSPs.

[0047] Computing Array Cell Structure

[0048] In some embodiments, as discussed above, the computing array includes an array of cells (e.g., n rows by m columns). Each cell may correspond to a basic computing primitive, such as a MACC function. Each cell may require up to three inputs and produce a single output. The inputs to the cell may include an input bias / offset or sub-result (if any) representing a partial result input from an adjacent cell, an input weight / parameter (e.g., a weight determined during a training process, or, if the computing array is used for training, representing some initial weights that will be updated during the forward and backward propagation phases of stochastic gradient descent (SGD) or other training techniques), and an input channel / activation corresponding to an activation from a previous layer of a neural network or an incoming channel of an input image. The cell processes the received inputs to generate an output feature corresponding to a partial sum of an output feature map. In some embodiments, these may subsequently be normalized and subjected to an "activation function" (e.g., a rectified linear unit or ReLU) to map them onto an input domain in which they can be used for subsequent layers.

[0049] For example, in a cell configured to perform a MACC function, the cell multiplies the received input weight by the received activation and adds the resulting product to the input bias (e.g., partial sum) received from an adjacent cell. The resultant sum is output as a bias or sub-result to another adjacent cell, or, if there is no adjacent cell for receiving the resultant sum, the resultant sum is output from the computing array as a result value. In some embodiments, each processing unit input is loaded in a single clock cycle. The loading of weights and activations is described in more detail below.

[0050] In some embodiments, each cell includes an array of sub-cells. For example, a cell may be configured to process weight and activation values each including multiple elements (e.g., 16-byte elements) and includes an array of sub-cells (e.g., 16 by 16 sub-cells) each configured to process a weight element (e.g., 1-byte weight element) and an activation element (e.g., 1-byte activation element) to generate a corresponding result. In some embodiments, the computing array includes an array of 20×20 cells, each cell having 16×16 sub-cells, resulting in a 320×320 element array.

[0051] Computing Array Weight Loading

[0052] In some embodiments, each cell of the computing array includes one or more registers for locally storing received weight values. This enables activations to be transmitted independently of the weights, thereby eliminating the need for both the activations and the weights to arrive at the cell at the same time, and simplifies routing of the signals to their desired locations. Additionally, the specific weights loaded onto a cell can be stored and used for multiple computations involving different activation values.

[0053] Figure 4A and Figure 4B illustrates loading weights into cells of a computing array according to some embodiments. Figure 4A illustrates how weights can be loaded into a computing array in many typical computing array systems. A weight transmission line 402 spans a row or column of the computing array that includes a plurality of cells 404, and transmits weight values received from a memory (e.g., directly or via a weight buffer). As Figure 4A shown, a plurality of capture registers 406 are positioned along the weight transmission line 402. Each capture register 406 is configured to capture the current weight value transmitted by the weight transmission line 402 and transmit the previously captured weight value on the weight transmission line 402 to a subsequent capture register 406. In some embodiments, the weight transmission line 402 includes one capture register 406 for each cell 404 of the row or column that the weight transmission line 402 spans. In each clock cycle, the capture register 406 of a particular cell can transmit its currently stored weight value to the next capture register 406 along the weight transmission line 402 corresponding to the subsequent cell, such that each weight value transmitted along the weight transmission line 402 propagates over multiple clock cycles.

[0054] Each cell of the computing array includes a weight register 408. During weight loading, each weight register 408 is configured to receive a weight value from the corresponding capture register 406 and store the received weight value for use by the cell in later computations. Each of the weight registers 408 receives a control signal that controls when each weight register 408 reads from its corresponding capture register 406. For example, the control signal for the weight register 408 is synchronized with the transmission of the weight value over the weight transmission line 402 such that when the capture register 406 receives the weight value to be loaded in the cell 404, each weight register 408 of the cell 404 reads the currently stored weight value of its corresponding capture register 406. The weight values stored in the weight registers 408 can be held by the cell 404 and used for computations over multiple cycles.

[0055] The coordination of data movement from memory to a computing array is referred to as "control flow". The control flow is executed by a controller that issues instructions describing the coordinated use of memory elements (e.g., use and / or bypass of weight buffers), weight reuse, and operations and data movement (e.g., loading of weights and activation values, output of result values).

[0056] Figure 4B Illustrates how weight values are loaded into a computing array according to some embodiments. Figure 4B The weight transmission line 402 in transmits weight values received from memory. In some embodiments, the weight transmission line 402 corresponds to a plurality of lines that form at least a portion of a data channel that corresponds to a row of the computing array that includes the cells 404. However, instead of including a capture register for each of the cells 404, weight data is transmitted over the weight transmission line 402 across multiple cells in a single clock cycle, where the capture register 406 is located between each of the plurality of cells. Each of the cells 404 is coupled to the capture register 406 such that the weight register 408 of the cell 404 can read the weight value that is currently being transmitted on the weight transmission line 402 from the same capture register 406. In some embodiments, each weight register 408 receives a control signal (e.g., a write enable signal) from the controller that indicates whether the weight register 408 is to read (e.g., from the capture register 406) the current value being transmitted on the weight transmission line 402. In this way, the transmission of the weight value on the weight transmission line 402 can be synchronized with the control signal provided to each of the weight registers 408 such that the weight value for a particular cell 404 is transmitted along the weight transmission line 402 in the same clock cycle (or with a predetermined offset) as the write enable control signal is provided to the weight register 408 of the particular cell 404, such that the weight register 408 of the cell can receive and transfer the weight value.

[0057] In some embodiments, during each of a plurality of clock cycles, the weight transmission line 402 stores the transmitted weight value (e.g., based on the received control signal) on the capture register 406 to be read by the cells 404 in the plurality of cells. In embodiments where the cell includes a plurality of sub-cells (e.g., 16 by 16 sub-cells), the transmitted weight value can include a weight value for each sub-cell. For example, in some embodiments, the weight transmission line 402 includes 16 lines, each line streaming 16 values (e.g., a vector of 16 values) onto the capture register 406, which are read by the cells 404 and used as the weight values for the 16 by 16 sub-cells of the cell. In some embodiments, the cell 404 receives the transmitted weight value that includes 16 vectors of 16 values and transposes each vector to provide the weight values to the corresponding columns of the sub-cells.

[0058] In other embodiments, multiple weight values for multiple cells may be transmitted on the weight transmission line 402 during a single clock cycle, where, in response to receiving a write enable signal, the weight register of each of the cells may receive the multiple weight values aggregated together from the capture register and extract and store the corresponding portions of the received weight values in the weight register (e.g., based on the address of the cell).

[0059] By configuring the weight transmission line 402 and the weight registers 408 of the cells such that the weight registers 408 read the weight values from the capture register located between multiple groups of cells, the amount of area required to implement the weight transmission line is reduced. Additionally, the reduced number of registers required to load the weight values reduces the total amount of clock power required to provide clock signals to the registers.

[0060] In some embodiments, the weight transmission line may be capable of transmitting weight values for a specific number of cells only over a single clock cycle (e.g., over a specific distance). Thus, in some embodiments, multiple capture registers may be positioned along the weight transmission line, corresponding to multiple groups of cells in each row of the computing array, dividing each row of the computing array into multiple sections. In some embodiments, the number of cells in each section corresponding to a capture register is based on the distance over which the weight transmission line can transmit weight values over a single clock cycle.

[0061] Figure 5A Illustrated is the loading of weights onto a computing array according to some embodiments. As Figure 5A shown, weight data is transmitted across the computing array on the weight transmission line 502. The weight transmission line 502 is configured to provide weight values to the cells of a specific row of the computing array. Each row of the computing array may include multiple cells (e.g., m cells), which are divided into groups of p cells each (e.g., 10 cells).

[0062] The weight transmission line 502 transmits weight values on a group of p cells 504 of the computing array during a single clock cycle. Capture registers 506 positioned along the weight transmission line 502 corresponding to each group of p cells (or between each group of p cells) capture the weight values transmitted by the weight transmission line 502 on that group of p cells and convey the captured weight values along the weight transmission line 502 to the next group of p cells (e.g., to the next capture register) in the subsequent clock cycle. In some embodiments, instead of the capture registers 506, different types of elements may be used to ensure the timing of the transmitted weight values across the transmission line. For example, latches may be used in some embodiments. In other embodiments, wave pipelining techniques may be used to maintain the timing of the propagating weight values at the correct rate.

[0063] As Figure 5AAs shown, each unit 504 in a set of p units is capable of reading and storing the weight values transmitted on weight transmission line 502 over the set of p units from capture register 506. In some embodiments, each of the p units receives a control signal that controls when the unit reads and stores the transmitted weight value from capture register 506. In some embodiments, the compute array loads p different weight values over multiple cycles (e.g., p cycles) for the set of p units, where different units load their respective weight values during each cycle. Since each capture register along weight transmission line 502 is used to provide weight values to multiple units (e.g., p units) within a row of the compute array, the number of capture registers required as the weight values travel between to reach each unit of the row is reduced, and the total number of clock cycles required for the row of units in the compute array to load the weight values is reduced. For example, loading weight values for two sets of p units can be performed in 2p + 1 clock cycles (e.g., as shown in Figure 5A ), where the additional 1 clock cycle accounts for the clock cycles required to transfer the weight value from the first capture register to the second capture register).

[0064] In some embodiments, each unit of the compute array receives a control signal from a controller. In other embodiments, the controller transmits the control signal to a subset of the units, and the subset of units then propagates the control signal to the remaining units of the compute array over one or more subsequent clock cycles. For example, in some embodiments, each unit can be configured to store the control signal it receives in a control signal register and propagate the control signal to adjacent units in the vertical direction, horizontal direction, or both in a subsequent clock cycle. The transmission of the weight values on weight transmission line 502 can be timed based on the propagation of the control signal such that each unit reads from the weight transmission line when the weight value for the unit is being transmitted.

[0065] Figure 5B Shows how control signals can be used when loading weights for units of a compute array according to some embodiments. In some embodiments, each unit receives a control signal that indicates whether the unit is to read the weight value currently being transmitted on the weight transmission line into the unit's weight register. For example, the control signal can include a "0" or a "1", where "0" indicates that the unit should not read the currently transmitted weight value into the weight register, and "1" indicates that the unit should read the currently transmitted weight value into the weight register.

[0066] In some embodiments, each unit stores the received control signal using a control signal register. The control signal register can store the value of the control signal and provide the control signal to the next unit of the compute array in a subsequent clock cycle. For example, as shown in Figure 5BAs shown, the control signal register of a cell can provide control signals to the subsequent cell in the next row of the computing array (e.g., the cell directly above the current cell), and to the subsequent cell in the next column of the computing array (e.g., the cell directly to the right of the current cell). Propagating control signals through adjacent cells simplifies the wiring required to provide control signals to the cells of the computing array, because the cells of the array can be connected to each other rather than requiring separate control lines for each cell. In some embodiments, the control signal register of a cell can provide control signals only to the cells in the next row (e.g., to cause the control signals to propagate up the columns of the cells in the computing array), and not to the cells in the next column.

[0067] In some embodiments, the cells of the computing array are configured to receive control signals from different directions for receiving data (e.g., weight data). Figure 5C A diagram showing how the cells of a computing array can receive control signals according to some embodiments. As Figure 5C shown, the computing array can be associated with an instruction control unit (ICU) 508. The ICU 508 is preferably located near the memory and the computing array. In other embodiments, the ICU 508 can be located on the side of the computing array opposite the memory. The ICU 508 is configured to receive at least one instruction and generate corresponding command signals for each column of the computing array. As Figure 5C shown, the ICU 508 transmits control signals to the cells of the computing array along an edge of the computing array (e.g., the bottom edge) via control signal lines 510, which is perpendicular to the edge (e.g., the left edge) through which data (e.g., weight values) is transmitted. In some embodiments, since the control signal data is typically much smaller in size (e.g., two orders of magnitude smaller) compared to the data of the weights and activation values used by the computing array, the amount of wiring required to transmit the control signals is much smaller than that for transmitting the weights and activation values. Thus, the control signals can be transmitted as weights and / or activations via different sides of the computing array without requiring significant additional circuitry along different sides of the computing array.

[0068] In some embodiments, since each cell is configured to propagate control signals to up to two adjacent cells (e.g., adjacent cells in the vertical direction and adjacent cells in the horizontal direction, e.g., as Figure 5BAs shown, the control signal can propagate diagonally through the computing array (e.g., starting from a corner of the array and propagating towards the opposite corner of the array). However, propagating the control signal from one corner of the computing array to the opposite corner may require loading weight values for all the cells of the computing array, without the option of loading weight values for only a portion of the cells of the computing array. For example, a control signal received by a corner cell of the computing array can cause that cell to load a weight value during a first clock cycle, and that weight value is then propagated to the remaining cells of the computing array, causing each of the remaining cells to load a weight value during subsequent clock cycles. In other embodiments, the control signal can be received by all the cells of a particular row or column and then propagated across the computing array. For example, each cell in the rightmost column of the computing array can receive a control signal that causes that cell to load a weight value during a first clock cycle, and that weight value is then propagated across the rows of the computing array during subsequent clock cycles (e.g., propagated to the second cell of each row during a second clock cycle, etc.). Propagating the control signal in this manner can allow weight values to be loaded for only a subset of the rows of the computing array.

[0069] In some embodiments, to allow weight values to be loaded for only a desired portion of the computing array (e.g., a particular rectangular sub-region of the computing array), the control signal can include two separate portions that propagate through the computing array in two different directions. For example, the control signal can include a vertical portion and a horizontal portion. Figure 6 Shows the propagation of the vertical and horizontal portions of the control signal through the computing array according to some embodiments. As Figure 6 shown, each cell 602 of the computing array 600 can receive a first control signal portion c1 corresponding to the horizontal portion and a second control signal portion c2 corresponding to the vertical portion.

[0070] Each of the horizontal portion c1 and the vertical portion c2 can include an indicator (e.g., a "0" or "1" value) of whether the cell should read the currently transmitted weight value into its weight register, where the cell reads the weight value only if both the vertical portion and the horizontal portion indicate that the weight value should be read (e.g., both c1 and c2 have a "1" value). On the other hand, if either c1 or c2 indicates that the cell should not read the currently transmitted weight value from the weight transmission line (e.g., c1 or c2 has a value of "0"), then the cell does not read the currently transmitted weight value.

[0071] The controller can provide a control signal portion to each cell on the edge of the computing array for propagation across the array. For example, the controller can provide a horizontal portion c1 of the control signal to each cell on the vertical edge (e.g., the left edge) of the computing array 600, and a vertical portion c2 of the control signal to each cell on the horizontal edge (e.g., the bottom edge) of the computing array 600, and then can propagate each portion across the array in the horizontal or vertical direction respectively in subsequent clock cycles. Each cell propagates only the vertical portion c2 of the control signal in the vertical direction and only the horizontal portion c1 in the horizontal direction.

[0072] By splitting the control signal into separate portions (e.g., a horizontal portion c1 and a vertical portion c2) and reading from the weight transmission line only when both portions of the control signal so indicate (e.g., both portions of the control signal are 1), the computing array can be configured to load weights for a specific portion of the computing array. Figure 6 The control signal loading scheme shown can be referred to as diagonal loading because the control signal that controls the loading of the weight values propagates along a diagonal that moves across the computing array over multiple clock cycles.

[0073] For example, Figure 7A illustrates a computing array in which weights are loaded only in specific portions of the array according to some embodiments. As Figure 7A shown, the computing array 700 receives control signal portions along the vertical edge (e.g., the right edge) and the horizontal edge (e.g., the bottom edge), which propagate across the array in the horizontal and vertical directions respectively. The control signal portions can include a "1" value or a "0" value, indicating whether the cell should load the weight value currently being transmitted along the associated weight transmission line. In this way, the horizontal control signal portion can specify the rows of the computing array for weight loading, while the vertical control signal portion can specify the columns of the computing array for weight loading.

[0074] Since each cell loads the weights transmitted along the weight transmission lines of the array based on both portions of the received control signal, only the cells of the computing array that are located in rows and columns that both have a control signal portion of "1" will load the weights, while the remaining cells of the computing array will not load the weights. In this way, by loading weights in the cells of the computing array based on the intersection of the horizontal and vertical control signals, weight loading can be performed for a specific region of the computing array without the need to load weight values for all cells of the computing array.

[0075] In some embodiments, the cells of the computing array are configured to receive the control signal from a single direction (e.g., vertically), rather than each cell receiving control signal portions along both the vertical and horizontal edges. Figure 7BShows a computing array in which weights are loaded only in specific portions of the array according to some embodiments. In some embodiments, the cells of the computing array are configured to receive control signals along each column. For example, as Figure 5C shown. As Figure 7B shown, each column of the computing array 710 receives a signal indicating whether that column should load the weight value transmitted along the associated weight transmission line. In some embodiments, since all cells of the column receiving the "1" command will load the weight value, cells in columns corresponding to rows outside the desired region may receive zero-valued weights. In other embodiments, the cells of the column receive a command indicating a row range, and each cell loads or does not load a weight value based on whether it is within the indicated row range. In some embodiments, configuring the computing array to receive control signals in only a single direction (e.g., vertically) can simplify wiring and routing. Thus, the computing array can be configured to separate the reception of control signals from the reception of data (e.g., weights and / or activation values), where the control signals are transmitted via a first direction (e.g., vertically) and the data is transmitted in a second, different direction (e.g., horizontally).

[0076] In some embodiments, the columns of the computing array are associated with an instruction control unit (ICU). The ICU is configured to receive instructions for one or more columns and determine the control signals to be sent to the cells of each column. For example, in some embodiments, the ICU receives an install weight instruction for the cells of the computing array. The install weight instruction may include at least a start column parameter and / or an end column parameter. The ICU parses the instruction and determines the command to be provided to the cells of each column, e.g., a "1" command if the column is between the start column and the end column indicated by the instruction, and a "0" command otherwise. In some embodiments, the instruction may also contain a parameter indicating the topmost row, which indicates the row of the column where the results of the cells will propagate down rather than up (to be described in more detail below).

[0077] Although Figure 7A and Figure 7B show only a single region of the computing array as loading weights, in other embodiments, the control signal portion may be configured to load weights for multiple non-contiguous regions of the computing array, based on which the rows and columns of the computing array are provided with horizontal and vertical control signal values of "1".

[0078] The ability to load weights only in specific sub-regions of the computing array allows for more efficient processing of models with low batch sizes. For example, for a computing array with 16×16 cells, if only a 2×16 region of the array is needed to implement the model, it only takes 2 clock cycles to reload the weight values onto the array, thus allowing the weight values of the model to be loaded every 2 clock cycles.

[0079] Weight loading order

[0080] Figures 8A to 8C Illustrates the order and timing of weight transfer to the cells of a computing array according to some embodiments. As described above, the transfer of weight values along the weight transfer lines is synchronized with the propagation of control signals across the cells of the computing array so that the correct weight values are loaded into each cell. Since the control signals for the cells can propagate diagonally across the computing array, the transfer of weight values across the weight transfer lines for each row or column of the computing array can be sorted and staggered to reflect the propagation of the control signals. For purposes of illustration, Figures 8A to 8C each of [figures] shows a computing array having nine cells arranged in three rows and three columns.

[0081] Figure 8A Illustrates providing control signals to a computing array according to some embodiments. Similar to Figure 7A the computing array 700 shown in [figure], the control signals are received in part by cells on the bottom and left edges of the array and propagate across the array up and to the right respectively over a plurality of subsequent clock cycles. Since the propagated control signals are received by different cells of the computing array at different times, the timing and order in which the weight values transferred on the weight transfer lines are loaded onto the cells of the computing array can depend on the direction and orientation of the weight transfer lines on the computing array. Although Figure 8A illustrates a horizontal control signal portion provided to the cells of a computing array via a horizontal direction, it should be understood that in some embodiments, the horizontal control signal portion can be provided to the leftmost cell of each row (e.g., the cells of the first column) via control signal lines extending in a vertical direction (e.g., such that all control signal lines providing control signals to the cells of the computing array extend in the vertical direction). In embodiments where horizontal and vertical control signals are propagated to subsequent cells in each row / column in each clock cycle, the horizontal control signal portion and the vertical control signal portion can be provided in a staggered manner (e.g., the cells of each row receive the horizontal control signal portion one clock cycle later than the cells of the previous row, and the cells of each column receive the vertical control signal portion one clock cycle later than the cells of the previous column) to maintain the timing at which each cell receives the corresponding horizontal control signal portion and vertical control signal portion indicating that the cell is to load weight data.

[0082] In other embodiments, the control signals are propagated to the cells of the computing array in only one direction (e.g., the vertical direction). For example, the bottom cells of each column of the computing array can be via corresponding control signal lines (e.g., as in Figure 5Creceives a control signal and propagates the received control signal to subsequent cells in its respective column. In some embodiments, the control signals for each column are staggered. For example, the control signals can be staggered such that the bottom cell of the first column receives a write enable control signal during a first clock cycle, and the bottom cell of the second column receives a write enable control signal during a second clock cycle, etc., such that each cell in a given row of the array receives a write enable control signal during a different clock cycle. This can be performed such that each cell in a given row of the compute array can load a different weight value (e.g., from a capture register along a weight transmission line for the row).

[0083] Figure 8B illustrates the order of weight loading in a compute array according to some embodiments, where weights are loaded column by column. Each of cells 1 to 9 in the compute array will be loaded with a corresponding weight value w1 to w9. The compute array can be associated with a plurality of weight transmission lines corresponding to each column in the columns of the compute array (e.g., three weight transmission lines corresponding to three columns of the compute array).

[0084] The weight values w1 to w9 are loaded on the weight transmission lines in an order that matches the propagation of the control signal through the cells of the compute array. For example, cell 9 can receive the vertical and horizontal portions of the control signal during a first clock cycle (e.g., the vertical and horizontal portions of the control signal as Figure 8A shown, or the staggered control signals received via the vertical direction as discussed above), which indicates that it should load the weight value being transmitted on the weight transmission line during a particular time period (e.g., during the first clock cycle). Thus, to match the timing of the control signal, during the first clock cycle, the weight value w1 is transmitted on the weight transmission line corresponding to the column of cell 1. During the second clock cycle, weights w2 and w4 are loaded into their respective cells. During the third clock cycle, weights w3, w5, and w7 are loaded, followed by weights w6 and w8 during the fourth clock cycle, and weight w9 during the fifth clock cycle. Thus, the weight values w1 to w9 are loaded in an order based on the columns of the compute array, staggering the weight values for each column based on the propagation of the control signal across the cells of the compute array.

[0085] Figure 8C illustrates the order of weight loading in a compute array according to some embodiments, where weights are loaded row by row. In the Figure 8C example shown, the weight transmission lines span each row of the compute array (e.g., three weight transmission lines corresponding to three rows of the compute array). Since the control signal is in the same manner as Figure 8Bpropagate in the same way across the cells of the computing array, so the same weight values are transmitted in each clock cycle (e.g., w1 is transmitted in the first clock cycle, w2 and w4 are transmitted in the second clock cycle, w3, w5, and w7 are transmitted in the third clock cycle, w6 and w8 are transmitted in the fourth clock cycle, and w9 is transmitted in the fifth clock cycle). However, since the weight transmission lines are oriented across rows rather than columns as shown in Figure 8B , the distribution of the weight values across different weight transmission lines that cross rows is transposed relative to how the weights are distributed across the weight transmission lines that cross columns. For example, while Figure 8B shows weights w1, w2, and w3 being transmitted on different weight transmission lines, in the configuration of Figure 8C , weights w1, w2, and w3 are transmitted at the same timing but on the same transmission lines. Similarly, while Figure 8B shows weights w1, w4, and w7 being transmitted on the same weight transmission line during different clock cycles, in the configuration shown in Figure 8C , weights w1, w4, and w7 are transmitted at the same timing but on different weight transmission lines.

[0086] Thus, the timing of how the weight values are transmitted on the weight transmission lines depends on the timing of the control signals propagating through the cells of the computing array. Additionally, the distribution of the weight values across different weight transmission lines depends on the orientation and direction of the weight transmission lines, where the distribution of the weight values is transposed when the weight transmission lines are horizontal relative to when the weight transmission lines are vertical.

[0087] By transposing the order of the loaded weight values, the weight loading can be aligned with the input activation stream. By aligning the weight loading and the activation loading, it is easier to scale the size of the computing array (e.g., as discussed with respect to Figure 2A and Figure 2B ). Additionally, the amount of wiring required to implement the weight and activation transmission lines for loading the weights and activation values into the computing array can be significantly reduced.

[0088] Figure 9A shows a high-level diagram of a computing array in which weights and activations are loaded from different sides. As shown in Figure 9AAs shown, the memory can load activation values onto the compute array via activation transmission lines that span the horizontal width of the compute array. The activation transmission lines are routed directly from the memory to the edge of the compute array (e.g., the left edge of the compute array) and extend across the compute array. On the other hand, weight values are loaded onto the compute array from a different direction (e.g., vertically from the bottom edge). Thus, in order to load weight values onto the compute array from the bottom edge, the weight transmission lines must extend to the bottom edge of the memory, across the distance between the bottom edge of the memory and the corresponding columns of the compute array, and across the length of the columns of the compute array. Therefore, the amount of wiring required to load weight values can be significantly greater than the amount of wiring required to load activation values. In some embodiments, the compute array is configured such that each cell of the compute array is longer in one dimension (e.g., rectangular rather than square). In cases where the weight transmission lines must extend across the longer dimension of the compute array (e.g., as Figure 9A shown), the additional amount of wiring required to load weight values onto the compute array is further increased. Additionally, in some embodiments, in order to maintain timing, the lengths of the transmission lines to each column of the compute array may need to match in length, thus requiring additional wiring even for the columns of the compute array closest to the memory.

[0089] On the other hand, Figure 9B FIG. shows a high-level diagram of a compute array in which weight loading and activation loading are aligned, according to some embodiments. As Figure 9B shown, both the weight transmission lines and the activation transmission lines can be routed collinearly directly from the memory to the compute array and span the horizontal width of the compute array, thus significantly reducing the amount of wiring required to route the weight and activation transmission lines. This can both reduce the area required to accommodate the circuitry for loading operands onto the compute array and potentially reduce the amount of latency when operating on the compute array. In embodiments where the compute array is longer in one dimension, the compute array and the memory can be placed such that the weight and activation lines span the shorter dimension of the compute array, which both reduces the amount of wiring required and allows the weight transmission lines to contain fewer capture registers (e.g., as Figure 5A shown), since each weight transmission line needs to span a shorter distance across the compute array. Additionally, since each of the weight transmission lines spans the horizontal width of the compute array, it is easier to maintain the uniformity of the lengths of the weight transmission lines compared to the case where the weight transmission lines are routed to the bottom edge of the compute array and up along the corresponding columns of the array. Further, since the weight transmission lines do not need to turn to reach the compute array 110, line congestion is reduced and the space on the chip available for routing the wiring can be utilized more efficiently.

[0090] Output of result values on the same side

[0091] In some embodiments, in addition to being able to load weights and activations from the same first side of the computing array, results generated by the computing array by processing the weight and activation values are also output from the first side. In some embodiments, the output result values can be stored in a memory and used as activation values for later computations. By outputting the result values from the same side of the computing array from which the activation values are loaded, the amount of wiring required to store the results and reload them as new activation values at a later time can be reduced.

[0092] In some embodiments, the result values of the computations are routed based on the diagonal of the computing array such that the result values of the computations can be output from the same side of the computing array as the weights and activation values are loaded while maintaining timing (e.g., all result values computed by the computing array can be output by the array a set amount of time after they are computed). The diagonal of the array divides the array into an upper triangle and a lower triangle.

[0093] Figure 10A A diagram of a computing array in which result values are computed is shown in accordance with some embodiments. As Figure 10A shown, weights and activation values (operands) are loaded from the same side of the computing array (e.g., via a plurality of weight transmission lines as described above, which may be collectively referred to as a weight transmission channel, and a plurality of activation transmission lines, which may be collectively referred to as an activation transmission channel). For example, in some embodiments, the weight transmission channel and the activation transmission channel extend across a side edge of the computing array and are coupled to the first cell in each row of the computing array. Additionally, result values are computed by summing the outputs of the cells in each column of the array, thereby producing result values at the top of each column of the array. In some embodiments, the cells at the top of each column in the computing array may include one or more post-processing circuits that perform one or more post-processing functions on the generated result values.

[0094] Since the result values are determined at the top cells of each column of the array, routing the result values to be output from the top edge of the array can be relatively straightforward (since the top cells of each column of the array are adjacent to the top edge). However, in order to route the result values to be output from the same side of the computing array from which the operands are loaded (e.g., the left edge, via a plurality of result output lines, or collectively, a result output channel), the routing should be configured such that the time at which each result is to be output by the computing array after each result is computed is constant, regardless of in which column the result value is computed.

[0095] Figure 10BA result path is shown for outputting a sum result value at the top cell of each column of the array. The results propagate in the opposite direction (e.g., downward), are reflected about the diagonal, and emerge from the same side as the loaded operands. By shifting the result value downward and reflecting it from the diagonal, the number of clock cycles required to output the result value from each column of the array is made constant.

[0096] Thus, in the embodiments shown in Figure 10A and Figure 10B , the weights are first loaded from the left side of the array and stored in their respective cells (e.g., using the techniques described above with respect to FIGS. 4 and 5). The activations are transmitted into the array from the same side and are processed based on the weight values. Once processed, each cell produces a sub-result, which is summed across the cells of each column. The final result value is produced by the last cell (e.g., the top cell) in each column, corresponding to the sum of the sub-results of all the cells in the column.

[0097] As Figure 10B shown, when a result value is produced at the top cell of each column of the array, the result value is first reflected downward along the same column. Upon reaching the diagonal, the result value is then reflected off the diagonal and routed out of the computing array toward the first side of the computing array (e.g., to the left). By reflecting the result value downward and off the diagonal, the number of cells that each result value passes through from the first side to where it is output is the same for all columns of the computing array. Thus, even if some result values are generated in columns of the computing array that are closer to the first side, the amount of time after the result value has been computed until it is output from the computing array is constant for all columns.

[0098] At least a portion of the cells of the computing array includes routing circuitry configured to route the final result value of the computing array (e.g., determined at the top cell of each column) to a first set of computing arrays for output. As Figure 10A and Figure 10B shown, the cells of the computing array are divided into diagonal cells 1002, routing non-diagonal cells 1004, and non-routing cells 1006. When the final result value is produced (e.g., at the top cell of each column of the array), the routing circuitry of the diagonal cells 1002 and the routing non-diagonal cells 1004 of the computing array route the final result value to the first side of the computing array, and the non-routing cells 1006 are not involved in the routing of the final result value.

[0099] Figure 10A and Figure 10BThe techniques shown for routing output data (e.g., final result values) out of the first side of a computing array can be applied to a computing array having n rows by m columns. The computing array generates m result values at the top cell of each of the m columns, and these result values propagate down each column (through the active non-diagonal cells of the column) until they reach the diagonal cell. In embodiments where n and m are not equal, the diagonal cell can be designated as the i-th topmost cell of each column, starting from the column opposite the first side (e.g., the first topmost cell of the column furthest from the first side, the second topmost cell of the column second furthest from the first side, etc.). The result values are "reflected" from the diagonal cell and propagate along the row of their respective diagonal cell, ensuring that all results are output from the first side of the computing array at the same timing. In some embodiments, the computing array can be configured to have a number of rows greater than or equal to the number of columns to ensure there are sufficient rows from which to output result values. In some embodiments, the matrices to be multiplied using the computing array can be transposed, e.g., m < n.

[0100] Figure 11 A high-level circuit diagram of a routing circuit within a separate cell of an array for routing result values is shown, according to some embodiments. Each cell of the array includes a routing circuit. The routing circuit receives a sub-result value of the cell (e.g., a MACC result, which corresponds to a processed value generated by processing a weight value and an activation value, summed with the sub-result value or partial sum received from the previous cell (if any) in the column), and transmits the result to the next cell in the column (e.g., the cell above) if there is a next cell. If there is no next cell (e.g., the current cell is the top cell of the column), then the sub-result value of the cell will be the result value to be output by the computing array (e.g., the final result value).

[0101] The routing circuit stores indications regarding whether the cell is on the top row of the array and whether the cell is on the diagonal of the array. In some embodiments, a controller transmits indications to each cell within the array regarding whether it is a top cell or a diagonal cell. This allows different cells within the array to be top cells and / or diagonal cells. For example, if only a portion of the array is used for computing (e.g., as Figure 7A and Figure 7BIf so (as shown), then the cells designated as top cells and diagonal cells will be different compared to the case where the entire array or the entire plane of the array is used for computing. In some embodiments, the computing array receives instructions (e.g., setup weight instructions) indicating a start column, an end column, and / or a top row. The ICU can determine which cells are top cells or diagonal cells based on the instruction parameters and transmit appropriate indications (e.g., as part of a command signal) to each cell. In other embodiments, each cell can receive one or more instruction parameters (e.g., indications of the top row, start column, and / or end column) and determine whether it is a top cell or a diagonal cell. Additionally, each cell can determine (or receive an indication) as to whether it is a routing cell or a non-routing cell (e.g., whether it is above or below the diagonal).

[0102] If a cell is on the top row of the array, the sub-result of the cell corresponds to the result value of the column of the array to be output (e.g., the final result value). Thus, the routing circuit reflects the result down along the same column. On the other hand, if the cell is not a top cell, the result of the cell is not the final result value to be output, and the routing circuit instead propagates the result of the cell to the next cell in its column (e.g., upward) for further computation and receives the result received from the cell above (corresponding to the result value of a previous computation not yet output) and propagates it downward to the cell below.

[0103] If a cell is on the diagonal of the array, the routing circuit is configured to receive the result value (e.g., the MACC result of the cell if the cell is also a top cell, or the result from the cell above) and reflect it to the left. On the other hand, if the cell is not on the diagonal, the routing circuit receives the result from the cell to the right and propagates it to the left (to the subsequent cell, or outputs it from the left side of the computing array).

[0104] Although Figure 11 illustrates routing the result value in a particular direction, it should be understood that the same techniques can be used to route the result value in other directions or towards other sides of the computing array. Additionally, although Figure 11 illustrates using a multiplexer in each cell, in other embodiments, the routing of the result value by the cell can be performed in other ways. For example, in some embodiments (e.g., where the size of the computing array is fixed), the cells can be hardwired to route the result value in a particular direction based on their position in the computing array. In some embodiments, each cell can use one or more switches or other types of circuitry to route the result value. In some embodiments, the routing circuit can be configured to check whether the cell is not a diagonal cell before routing the result to the cell below.

[0105] In some embodiments, only the diagonal cells 1002 and the active non-diagonal cells 1004 of the computing array include routing circuits, while the inactive cells 1006 of the computing array do not include routing circuits. Figure 7A and / or Figure 7B In other embodiments of the present invention, which cells are diagonal cells, routing non-diagonal cells, and non-routing cells can be changed based on the configuration of the computing array. In this way, all cells can include corresponding routing circuits. In some embodiments, the routing circuits of non-routing cells can be powered off or run in a lower power state.

[0106] Thus, using the above technique, the computation array is able to load activation and weight values and output result values all from the same first side of the array. By routing all inputs and outputs through the first side of the array, the size of the array can be scaled more easily and the amount of wiring required can be greatly reduced.

[0107] although Figure 11 Routing circuitry is shown implemented as part of each cell of the computational array, but it should be understood that in other embodiments, each sub-cell of each cell may include corresponding routing circuitry for routing result values generated by the sub-cell to be output by the computational array. For example, in some embodiments where each cell includes multiple sub-cells (e.g., an array of sub-cells), each sub-cell may include a circuit similar to Figure 11 The routing circuit of the routing circuit shown in , and wherein routing of result values is performed at a sub-cell level rather than a cell level (eg, based on top and diagonal sub-cells within an array).

[0108] Figure 12 An example architecture for loading weights and activations into a computation array is shown in accordance with some embodiments. Figure 12 A row of cells 1202 within a computation array 1200 is shown. An activation transmission line 1204 enters the boundary of the computation array via a first side (e.g., left side) of the array and runs across the row of cells. In some embodiments, a capture register 1206 positioned across the activation transmission line captures the transmitted activation values so that they can be loaded into an activation register 1208 within each cell of the row. In some embodiments, the activation values are propagated to each cell in the row over consecutive clock cycles (e.g., a particular activation value is loaded into the first cell of the row during a first clock cycle, loaded into the second cell of the row during a second clock cycle, etc.). In other embodiments, other elements or techniques such as latches or wave pipelining may be used in place of capture registers. In some embodiments, the activation transmission line 1204 may transmit multiple activation values that are aggregated together, where each cell may extract a particular activation value of the aggregate activation value to be used for calculation.

[0109] The weight transmission line 1210 enters the boundary of the computing array via the first side (e.g., the left side) of the array and runs across the rows of cells. The weight distribution register 1212 positioned along the weight transmission line receives the transmitted weight values, which can be read by the weight register 1214 of the cell. In some embodiments, each weight register 1214 of the cell is configured to receive a control signal indicating when the weight register reads the current weight value within the weight distribution register. In other embodiments, the weight distribution register determines which cell will receive the weight value based on the address of the processing unit and the received control signal. Since the weight distribution register 1212 can distribute the received weights to any cell in the row, the weights can be quickly loaded into a specific cell without propagating through the computing array. In some embodiments, the weight distribution register receives different weight values in each cycle while providing a write enable control signal to consecutive cells of the row, causing one cell of the row to load the corresponding weight value in each clock cycle (e.g., the first cell of the row loads the first weight value during the first clock cycle, the second cell of the row loads the second weight value during the second clock cycle, etc.).

[0110] Each cell can process the received weight and activation value (e.g., multiply) to produce a processed value, which is summed with the partial sum 1216 received from the row below (if there is a row below). If the cell does not belong to the top row of the array, the summed partial sum value is propagated to the next cell in the row above. On the other hand, if the cell belongs to the top row of the array, the sum of the processed value and the partial sum 1216 forms the result value to be output. Additionally, each cell is configured to receive result data (e.g., result value) from the cell above and / or the cell to the right at the routing circuit 1218 (which may correspond to the routing circuit shown in Figure 11 and propagate the result down or to the left based on whether the cell is in the top row of the array or on the diagonal of the row.

[0111] In some embodiments, since the cells of each row depend on the result values generated by the previous row (e.g., the row below) to determine their own result values, the activation values of the rows of the array can be loaded in a staggered manner (e.g., each row is one activation value “ahead” of its row above).

[0112] In some embodiments, the cells of the computing array 1200 are configured to start loading the activation values and calculating the results before all the weight values have been loaded into the computing array 1200. For example, as Figure 12As shown, the activation values are propagated across each row 1202 over multiple cycles (e.g., via capture register 1206). Once the weight value of a cell is received, each cell in the row can load the activation value and start computing the result. For example, since the computing array is capable of loading different weight values for the next cell in the row at each clock cycle, once the first cell in the row loads the weight value (e.g., during the same clock cycle, one clock cycle later, or some other predetermined time offset), the computing array can start propagating the first activation value across the cells in the row. The loaded weight value can also be used to process the subsequently received activation values for each cell in the row. Since the computing array can start loading activation values and processing results on a particular row once the first cell in the row loads the weight value, the computing array is able to perform computations more efficiently with different batches of weights because it does not have to wait until all the weights in a new batch are loaded before starting to process the activations. This reduction in latency enables the computing array to handle applications where a given set of weights is only used to process a small number of activations and then updated with new weights.

[0113] The use of the routing circuitry enables the final output value of the computing array (e.g., generated at the top cell of each column) to be output from the first side while maintaining the relative timing of the output result values. For example, in some embodiments, for a computing array with 20 rows, it would take 20 cycles to propagate the results computed by the cells in the column for the first activation value to the top cell of the column to produce the final result value for the column, and 20 additional cycles to output the result from the first side of the computing array (e.g., 20 - i cycles to propagate the final result value of the i-th column of the array down to the diagonal cell of the column, and i cycles to propagate the value from the diagonal cell to the first side of the array). Additionally, the last column of the array determines its result for a given activation value column m cycles after the first column (where m is the total number of columns in the array), resulting in an additional m cycles between outputting the result values of the first and last columns of a given activation value column from the array.

[0114] In some embodiments where the instruction / control signals are propagated across each column of the computing array over multiple cycles (e.g., 1 cycle per cell), the weight values and activation values are propagated in a staggered manner (e.g., as Figure 8C shown) to match the timing of the propagated control signals. Although the previous illustrations show using a particular technique to load each of the weight values and activation values onto the computing array, in some embodiments, the manner of loading the weights and activation values can be interchanged (e.g., the weight values are loaded along a transmission line with a capture register for each cell, and the activation values are loaded along a transmission line with an assignment register configured to distribute the value to multiple cells, or where each cell is able to directly read the activation value from the transmission line), or some combination thereof.

[0115] Weight buffer

[0116] Efficient use of a computing array requires weights to be loaded onto the computing array at a rate that matches the rate at which the computing array can receive weight values. For example, using the Figure 6 control signal propagation scheme shown in, weights can be loaded into an n×n computing array or a subset of the computing array, and all n×n cells can be loaded with weights in n clock cycles. By matching the loading of weights to the rate at which the computing array can receive weights, weights can be loaded onto the computing array without interruption (e.g., over n cycles).

[0117] However, in some embodiments, driving weights to the computing array at full bandwidth consumes a large amount of valuable data bus bandwidth. To enable fast weight loading without interrupting the loading of activation values for computation, a weight buffer can be used. In some embodiments, the bandwidth for loading weight values onto the weight buffer is less than the bandwidth at which the weight values can leave the weight buffer to be loaded onto the computing array. For example, the weight values loaded into the weight buffer can be directed to one of a plurality of buffers, each buffer corresponding to one or more rows of the computing array. The weight values can then be loaded in parallel from the plurality of buffers onto different rows of the computing array, enabling a large number of weight values to be loaded at once.

[0118] For example, when the computing array processes the loading of weights and activation values, future weight values to be loaded onto the array can be cached in the weight buffer in preparation for loading onto the computing array over a short time frame. This “burst” high-bandwidth weight loading allows for good performance when processing a model because it allows the weight values of the model to be loaded without interruption.

[0119] Thus, in some embodiments, the weight buffer provides a capacitor-like ability to store weight values until they are ready to be loaded onto the computing array. The weight values can be stored in the weight buffer over time and then quickly released to be loaded onto the computing array. In some embodiments, the weight buffer can also provide a pin extender function by providing additional local wire bandwidth (e.g., to allow the transfer of multiple weight values across multiple cells in a single clock cycle).

[0120] In some embodiments, the weights stored in the weight buffer pass through a preprocessor that allows switching of the weights to arrange them appropriately within the computing array, reusing weights to create useful constructs for convolution, and / or preprocessing the numerical values.

[0121] The use of a weight buffer can thus facilitate the efficient use of the computational resources of a computational array by allowing weights to be loaded into the weight buffer asynchronously and / or over many cycles, while serving as a capacitor-like hardware structure that enables the rapid loading of stored weight values onto the computational array. This potentially simplifies scheduling as it allows controller time flexibility in loading weight values over an extended period of time.

[0122] While the use of a weight buffer can allow for more efficient weight loading, in some embodiments it may be desirable to bypass the weight buffer or eliminate the weight buffer entirely (e.g., to save power and / or simplify circuit design). For example, in some embodiments, weights are loaded into an array of n×n cells over more than n cycles.

[0123] Other Considerations

[0124] The foregoing description of the embodiments has been presented for purposes of illustration; it is not intended to be exhaustive or to limit the patent rights to the exact forms disclosed. Many modifications and variations are possible in light of the above disclosure, as will be recognized by those of ordinary skill in the relevant art.

[0125] Some portions of this specification describe embodiments in terms of algorithms and symbolic representations of operations on information. Those skilled in the data processing arts typically use these algorithmic descriptions and representations to effectively convey the substance of their work to others skilled in the art. These operations, while described functionally, computationally, or logically, are to be understood as being implemented by a computer program, an equivalent circuit, microcode, or the like. Additionally, without loss of generality, these arrangements of operations are sometimes conveniently referred to as modules. The described operations and their associated modules may be embodied in software, firmware, hardware, or any combination thereof.

[0126] Any step, operation, or process described herein may be performed or implemented singly or in combination with other devices using one or more hardware or software modules. In one embodiment, a software module is implemented using a computer program product that includes a computer-readable medium containing computer program code that can be executed by a computer processor to perform any or all of the described steps, operations, or processes.

[0127] Finally, the language used in the specification has been principally selected for readability and instructional purposes, and it may not have been selected to delineate or circumscribe the patent rights. Accordingly, it is intended that the scope of the patent rights not be limited by this detailed description, but rather by any claims that may be presented based on this application. Thus, the disclosure of the embodiments is intended to be illustrative and not to limit the scope of the patent rights set forth in the appended claims.

Claims

1. A system that uses only a single side slave computing array to load operands and output results, comprising: A computing array, the computing array including a plurality of cells arranged in n rows and m columns, each cell being configured to generate a processed value based on a weight value and an activation value; At least two collinear transmission channels, the at least two collinear transmission channels including at least a weight transmission channel and an activation transmission channel; Wherein, each of the weight transmission channel and the activation transmission channel extends across a first side edge of the computing array to provide a weight value and an activation value to the cells of the computing array; And A result output channel, the result output channel extending across the first side edge of the computing array and outputting a plurality of result values generated by the computing array at the first side edge.

2. The system according to claim 1, wherein: The computing array is configured to generate the plurality of result values based on the processed values generated by each cell.

3. The system according to claim 2, wherein The computing array is configured to: Generate, at the top cell of each column of the m columns of the computing array, a result value among the plurality of result values corresponding to the sum of the processed values generated by the cells of the corresponding column of the computing array; Output the m generated result values from a first side of the computing array via the result output channel.

4. The system according to claim 3, wherein, The computing array is configured to output the m generated results from a first side of the computing array by: Propagating each of the m results through a plurality of cells along the corresponding column until reaching a cell along the diagonal of the corresponding column within the computing array, and Propagating each of the m results from the cell along the diagonal of the computing array along the m rows of the computing array such that each of the m results is output from the computing array from the first side of the computing array.

5. The system according to claim 4, wherein, Each of the m results is output from a different row of the n rows of the computing array, where m ≤ n.

6. The system according to claim 2, wherein Each cell of the computing array stores an indication of whether it is a top cell or a diagonal cell of the computing array, and includes routing circuitry configured to route the received result values based on whether the cell is a top cell or a diagonal cell of the computing array.

7. The system according to claim 1, further comprising a memory storing a plurality of weight values and a plurality of activation values, wherein, The computing array is coupled to the memory via the at least two collinear transmission channels, and wherein the weight transmission channel and the activation transmission channel provide a weight value and an activation value to the cells of the computing array from the plurality of weight values and the plurality of activation values stored in the memory.

8. The system according to claim 1, further comprising a controller circuit configured to transmit a plurality of control signals propagating along the columns of the computing array to the cells of the computing array.

9. The system according to claim 1, wherein For a row of the computing array, the weight transmission channel includes a capture register and a transmission line coupled to a set of cells of the row of the computing array, wherein at least one cell of the set of cells is capable of reading the weight value stored at the capture register in response to the cell receiving a write enable control signal.

10. The system according to claim 9, wherein, During each of a plurality of clock cycles, different weight values are stored into the capture register for reading and storing by different ones of the set of units.

11. The system according to claim 1, wherein, Each unit in a row of the computing array receives the same activation value, and wherein units in a particular column of the row of the computing array are configured to: receive the activation value via the activation transmission channel during a first clock cycle, and propagate the received activation value to units in a subsequent column of the row of the computing array during a second clock cycle.

12. The system according to claim 11, wherein, The units of the row are configured to receive corresponding weight values via the weight transmission channel during each of a plurality of clock cycles, and wherein, after a predetermined period of time after each unit has received its corresponding weight value, the activation value is propagated to each unit of the row.

13. The system according to claim 1, wherein, n is equal to m.

14. The system according to claim 1, wherein Each unit of the computing array includes an array of sub-units.

15. The system according to claim 1, wherein, The stored weight values include a plurality of weight vectors.

16. A method of loading operands and outputting results using only a single-sided slave computing array, comprising: Transmitting a plurality of weight values and a plurality of activation values over at least two collinear transmission channels including at least a weight transmission channel and an activation transmission channel; Receiving, at a computing array including a plurality of units arranged in n rows and m columns, the plurality of weight values and the plurality of activation values from the weight transmission channel and the activation transmission channel, wherein each of the weight transmission channel and the activation transmission channel extends across a first side edge of the computing array to provide a weight value from the plurality of weight values and an activation value from the plurality of activation values to each unit of the computing array; Generating, at each unit of the computing array, a processed value based on the received weight value and the received activation value; and Outputting, at the first side edge, a plurality of result values generated by the computing array via a result output channel that extends across the first side edge of the computing array.

17. The method according to claim 16, further comprising: Generating, at the computing array, the plurality of result values based on the processed values generated by each unit.

18. The method according to claim 17, wherein, Generating the plurality of result values at the computing array based on the processed values generated by each unit includes: Generating, at a top unit of each of the m columns of the computing array, a result value of the plurality of result values corresponding to a sum of the processed values generated by units of the corresponding column of the computing array; Outputting the m generated result values from a first side of the computing array via the result output channel.

19. The method according to claim 18, wherein Outputting the m generated result values from a first side of the computing array via the result output channel includes: Propagating each of the m results through a plurality of units along the corresponding column until reaching a unit along a diagonal of the computing array within the corresponding column, and Propagating each of the m results from the unit along the diagonal of the computing array along the m rows of the computing array such that each of the m results is output from the computing array from the first side of the computing array.

20. The method according to claim 18, wherein, Each of the m results is output from a different one of the n rows of the computing array, where m ≤ n.

Citation Information

Patent Citations

  • Sparse convolutional neural network accelerator

    US20180046906A1