Initializing on-chip operations
By configuring integrated circuit hardware blocks to follow predetermined actions based on scheduled operations and counters, the need for decoding logic and complex wiring is eliminated, facilitating efficient and simultaneous execution of operations across the chip.
Patent Information
- Application Number
- JP2023081795
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-08-22
- Filing Date
- 2023-05-17
- Publication Date
- 2025-05-20
- Estimated Expiration
- 2040-08-20
AI Technical Summary
Existing integrated circuit technologies require decoding logic to determine the destination of data packets, leading to complex wiring and increased packaging area, especially for large numbers of hardware blocks.
Configuring hardware blocks to act on data according to predetermined actions without decoding logic, using a scheduled operation approach where data is transferred between blocks based on explicit schedules and counters to manage state changes.
This method reduces the need for decoding logic and complex wiring, allowing efficient configuration of multiple hardware blocks within an integrated circuit without direct connections to data sources, enabling simultaneous execution of operations across the chip.
Smart Images

Figure 0007680495000001 
Figure 0007680495000002 
Figure 0007680495000003
Abstract
Description
[Technical field]
[0001] The present disclosure relates to initializing integrated circuit operation and the individual operations of various integrated circuit hardware blocks. [Background technology]
[0002] Data transmitted by processor and microcontroller chips often includes encoded information, such as a header that specifies where the data should be sent. Each processor or microcontroller that receives such data must therefore include decoding logic to decode the header and make a decision whether the received data should be saved, installed to initialize the processor or microcontroller, or forwarded to another circuit. Summary of the Invention [Means for solving the problem]
[0003] In general, the disclosure relates to initializing the configuration of a semiconductor chip, where operations to be performed on the chip are explicitly scheduled (the operations are sometimes said to be "deterministic"). More specifically, the disclosure relates to a semiconductor chip divided into separate hardware blocks, where data is transferred between the hardware blocks according to an explicit schedule. Rather than configuring the hardware blocks to include decode logic or similar functionality that determines, based on the content of the received data, whether the received data should be installed, saved in memory, or forwarded to another hardware block, the hardware blocks are instead configured to act on the data in advance according to predetermined actions. In this manner, the hardware blocks may be characterized as "agnostic" with respect to the ultimate destination of received data.
[0004] In general, in some aspects the subject matter of this disclosure is embodied in a method of configuring an integrated circuit including a plurality of hardware tiles, the method including: establishing a data transfer path through the plurality of hardware tiles by configuring each hardware tile except a last hardware tile of the plurality of hardware tiles to be in a data transfer state, where configuring each hardware tile except a last hardware tile to be in a transfer state includes installing a respective transfer state counter that specifies a corresponding predetermined length of time the hardware tile is in the data transfer state; supplying a respective program data packet along the data transfer path to each hardware tile of the plurality of hardware tiles, the program data packet comprising program data for the hardware tile; and installing the respective program data for each hardware tile of the plurality of hardware tiles.
[0005] Implementations of the method may include one or more of the following features. For example, in some implementations, a forwarding state counter for each hardware tile except for a last hardware tile of the plurality of hardware tiles is installed upon receiving a first data packet that has traversed the data forwarding path. The first data packet may include a program data packet that includes program data for the last hardware tile of the plurality of hardware tiles.
[0006] In some implementations, installing a respective forwarding state counter for each hardware tile includes defining the forwarding state counter in a trigger table for the hardware tile. When the forwarding state counter for each hardware tile reaches a corresponding predetermined length of time, the trigger table may trigger the installation of program data for the hardware tile and may cause the hardware tile to exit the data forwarding state. For each hardware tile that includes a respective forwarding state counter, the corresponding predetermined length of time for the forwarding state counter is a function of the number of subsequent hardware tiles in the data forwarding path.
[0007] In some implementations, each hardware tile of the plurality of hardware tiles stores respective program data for the hardware tile in a local memory.
[0008] In some implementations, each hardware tile that includes a respective transfer state counter forwards at least one program data packet to at least one other hardware tile in a data transfer path.
[0009] In some implementations, each hardware tile includes a systolic array of circuit elements.
[0010] In some implementations, multiple tiles are arranged in a one- or two-dimensional array.
[0011] In some implementations, the method further includes installing a respective kick-off counter in at least some of the plurality of hardware tiles that specifies a corresponding predefined length of time before the hardware tile initiates an operation defined by the program data installed on the hardware tile. The respective kick-off counter of each hardware tile except for a last hardware tile of the plurality of hardware tiles may be installed upon receiving a first data packet. The predefined length of time for each kick-off counter may be different. The predefined length of time for each kick-off counter may be a function of the number of hardware tiles in the data transfer path. The predefined length of time for each kick-off state counter may be defined such that multiple hardware tiles execute their respective program data simultaneously. Installing the respective kick-off counter of each hardware tile may include defining the kick-off counter in a trigger table of the hardware tile.
[0012] In general, in some other aspects, the subject matter of this disclosure may be embodied in a method of configuring an integrated circuit including a plurality of hardware tiles, the method including establishing a data transfer path through each hardware tile of the plurality of tiles except for a last hardware tile of the plurality of tiles, where establishing the data transfer path includes sequentially configuring each hardware tile of the data transfer path by (a) installing program data for the tile, (b) configuring the tile to be in a transfer state, and (c) installing a program kick-off counter that specifies a corresponding predetermined length of time that the hardware tile is in the data transfer state.
[0013] Implementations of these methods may include one or more of the following features: For example, in some implementations, for a particular tile in the data transfer path, the predetermined amount of time is a function of the number of tiles in the plurality of tiles that do not yet have the program data installed.
[0014] In some implementations, after each program kick-off counter reaches a corresponding predetermined length of time, the tile on which the program kick-off counter is installed begins to perform operations according to the program data installed on the tile.
[0015] Various implementations include one or more of the following advantages. For example, in some implementations, the processes described herein allow for configuration of multiple hardware blocks located internally within an array of hardware blocks to be configured without requiring the internal hardware blocks to be directly wired to their data sources. In some implementations, the processes described herein allow for configuration of hardware blocks without the need to encode destination data within data packets. In some implementations, the processes described herein allow for hardware blocks to be configured without the need to install decoding logic within the hardware blocks.
[0016] The details of one or more embodiments are set forth in the accompanying drawings and the description below. Other features and advantages will become apparent from the description, drawings, and claims. [Brief description of the drawings]
[0017] [Figure 1] FIG. 1 is a schematic diagram illustrating an example integrated circuit device configured to operate according to a scheduled operation. [Diagram 2] 4 is a flowchart illustrating an example process for initializing a hardware tile with configuration data. [Diagram 3] FIG. 11 is a flow diagram illustrating an example of a process for initializing a group of hardware tiles. [Figure 4] FIG. 13 is a schematic diagram illustrating an example of a trigger table. [Diagram 5] FIG. 1 is a schematic diagram illustrating an example of a dedicated logic circuit that can be configured to operate according to a scheduled operation. [Figure 6] FIG. 6 is a schematic diagram illustrating an example of a tile used in the ASIC chip of FIG. [Figure 7A] FIG. 2 is a schematic diagram outlining data flow through the ASIC at different times for an exemplary process performed by the ASIC. [Figure 8A] FIG. 2 is a schematic diagram outlining data flow through the ASIC at different times for an exemplary process performed by the ASIC. [Figure 9A] FIG. 2 is a schematic diagram outlining data flow through the ASIC at different times for an exemplary process performed by the ASIC. [Figure 10A] FIG. 2 is a schematic diagram outlining data flow through the ASIC at different times for an exemplary process performed by the ASIC. [Figure 11] FIG. 2 is a schematic diagram outlining data flow through the ASIC at different times for an exemplary process performed by the ASIC. [Figure 12A] FIG. 2 is a schematic diagram outlining data flow through the ASIC at different times for an exemplary process performed by the ASIC. [Figure 13A] FIG. 2 is a schematic diagram outlining data flow through the ASIC at different times for an exemplary process performed by the ASIC. [Figure 7B] FIG. 7B is a schematic diagram detailing the data flow within a single tile of the ASIC at times relevant to FIG. 7A. [Figure 8B] FIG. 8B is a schematic diagram detailing the data flow within a single tile of the ASIC at times relevant to FIG. 8A. [Figure 9B] FIG. 9B is a schematic diagram detailing the data flow within a single tile of the ASIC at times relevant to FIG. 9A. [Figure 10B] FIG. 10B is a schematic diagram detailing the data flow within a single tile of the ASIC at times relevant to FIG. 10A. [Figure 12B] FIG. 12B is a schematic diagram detailing the data flow within a single tile of the ASIC at the time associated with FIG. 12A. [Figure 13B] FIG. 13B is a schematic diagram detailing the data flow within a single tile of the ASIC at the time associated with FIG. 13A. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0018] In general, the present disclosure relates to initializing the configuration of a semiconductor chip, where operations performed on the chip are explicitly scheduled (the operations are sometimes said to be "deterministic"). In one example, the semiconductor chip may be divided into individual hardware blocks, and data is transferred between the hardware blocks according to an explicit schedule. More specifically, the individual hardware blocks may operate according to individualized operation schedules to execute a coordinated program executed by the semiconductor chip as a whole. In other words, the individual hardware blocks perform their respective operations at scheduled times according to a clock (e.g., a counter), rather than, for example, performing operations in response to control signals or according to an unscheduled sequential list of process instructions. Each hardware block represents an associated set of replicated logic, such as a subset of the electrical circuits (e.g., logic circuits) on the chip that are configured to perform a particular set of tasks independent of the tasks performed by other hardware blocks. These operations include, but are not limited to, data transfer operations, initialization operations, matrix operations, vector operations, scalar operations, logical operations, memory access operations, external communication operations, or combinations thereof.
[0019] Rather than configuring a hardware block to include decode logic or similar functionality that determines whether received data should be installed (e.g., to initialize program operations within the hardware block), saved to memory, or forwarded to another hardware block, the hardware blocks of the present disclosure are instead configured to pre-address data in a particular manner. For example, a hardware block may be pre-configured to install received data (e.g., to initialize program operations within the hardware block), save received data to memory, or forward received data to another hardware block. In this manner, a hardware block may be characterized as pre-configured to address data in a particular manner independent of the data received. A hardware block may be pre-configured to address data in a particular manner at / during a pre-defined time during execution of a schedule of operations, i.e., a hardware block may be configured to change the way it addresses data at a pre-defined time.
[0020] Each hardware block executes an individual schedule of hardware block-specific operations. The individual schedules for each hardware block together represent a complete program (e.g., neural network operations) executed by the entire chip. However, prior to executing such operations, program data specifying the particular operations to be executed is distributed to and installed on each hardware block. To provide program data to each hardware block without including decoding logic, one option is to wire the source of the program data directly to each block. For a large number of hardware blocks, that amount of wiring may require a significant amount of packaging area and may become an untenable solution.
[0021] Alternatively, as described in this disclosure, only a portion of the hardware blocks (e.g., the outermost hardware blocks in a two-dimensional array of blocks) may be directly connected to the source of program data. To reach hardware blocks located internally to the array, a group of hardware blocks may be placed in a data transfer state such that each group establishes a data transfer path to an internal block. Each hardware block in the transfer state may automatically reconfigure to a new non-data transfer state after a predetermined amount of time specific to that hardware block. For example, one or more hardware blocks may automatically reconfigure to a data storage state that stores data received in the hardware block. Alternatively, or in addition, one or more hardware blocks may automatically reconfigure to a data initialization state that initializes a program to be executed by the hardware block, the program being defined by the received data. In some cases, the hardware block may execute a scheduled operation specified by the program data at a predetermined counter time.
[0022] In some implementations, the semiconductor chip including the hardware block is an application specific integrated circuit (ASIC) designed to perform machine learning operations. The ASIC includes, for example, an integrated circuit (IC) customized for a specific application. For example, the ASIC may be designed to perform machine learning model operations including, for example, recognizing objects in an image as part of a deep neural network, machine translation, speech recognition, or other machine learning algorithm. When used, for example, as an accelerator for a neural network, the ASIC may receive inputs to the neural network and compute neural network inferences for the inputs. Data inputs to a neural network layer, for example, either an input to the neural network or an output of another layer of the neural network, may be referred to as actuation inputs. The inferences may be computed according to respective sets of weight inputs associated with the layers of the neural network. For example, some or all layers may receive a set of actuation inputs and process the actuation inputs according to a set of weight inputs for the layer to generate outputs. Furthermore, the repetitive nature of the computational operations performed to compute neural network inferences facilitates explicitly scheduled chip operations.
[0023] FIG. 1 is a schematic diagram illustrating a simplified example of an integrated circuit chip 10 according to the present disclosure. The chip 10 may be a general purpose integrated circuit or a dedicated integrated circuit. For example, the chip 10 may be an ASIC, a field programmable gate array (FPGA), a graphics processing unit (GPU), or any other suitable integrated circuit. The chip 10 includes a number of hardware blocks 14, also referred to herein as "tiles" (labeled "A", "B", and "C") arranged in an array. The hardware tiles 14 may be coupled to each other using a data bus 16. Although only three hardware tiles 14 are shown in FIG. 1, the chip 10 may include other numbers of hardware tiles, such as, for example, 10, 20, 30, 40, 50, 100, or 200, among others. Additionally, although the hardware tiles 14 are shown as arranged in a linear array, the hardware tiles 14 may also be arranged in other configurations, such as a two-dimensional array having multiple rows and multiple columns. The chip 10 also includes a communication interface 12. The communication interface 12 is coupled to at least a first hardware tile 14 (e.g., Tile A) using a data bus 17. The communication interface 12 may include, for example, one or more sets of serializer / deserializer (SerDes) interfaces and a general purpose input / output (GPIO) interface. The SerDes interfaces are configured to receive data for the ASIC 10 (e.g., configuration data or program data for a hardware tile as described herein) and to output data from the ASIC 10 to external circuitry. The chip 10 is shown in a simplified manner for purposes of illustration and description. However, in some implementations, the chip 10 includes additional components, such as memory and other circuitry appropriate for the purpose of the chip 10.
[0024] A hardware tile 14 represents a set of replicated logic, such as a subset of the electrical circuitry (e.g., logic circuitry) on chip 10 designed to perform a particular set of tasks independent of the tasks performed by other hardware tiles 14. Each hardware tile 14 may represent the same or different types of circuitry, e.g., hardware tiles A, B, and C may represent tiles of a dedicated chip designed to perform machine learning functions (described in more detail below). For example, hardware tiles A, B, and C may represent computational nodes of a neural network configured to perform matrix operations.
[0025] The operations of the hardware tiles 14 may be performed at predetermined times according to a common clock signal on the chip 10. Each hardware tile 14 operates according to its own individualized operation schedule. Each operation schedule represents a portion of a program (e.g., a "subprogram") to be executed by the entire chip 10, and each operation schedule represents a portion of the program to be executed by a corresponding individual hardware tile 14. An operation schedule includes a set of program operations to be executed by a hardware tile 14 at a predetermined counter value. In other words, an operation schedule can be considered as a list of timers that trigger a particular operation (e.g., see FIG. 4) to be executed by a particular hardware tile 14 at a pre-scheduled "chip time". For example, each schedule can include a list of execution counter values (e.g., execution times) that include an associated operation to be executed at the counter value. In some examples, each operation is represented by a scheduled counter value and data such as an instruction code that specifies the operation to be executed by the particular hardware tile 14 at the scheduled counter value. The operations executed by a hardware tile 14 may be coordinated with respect to the operations executed by other hardware tiles. The action schedule may be implemented on each tile using a "trigger table," which is described in more detail below.
[0026] As shown in FIG. 1, each hardware tile 14 may include a control circuit 18, a local memory 20, and one or more computational units 22. In some implementations, a program specified by an operation schedule may be stored in the local memory 20 of the hardware tile 14. The computational units 22 represent electrical circuits configured to perform specific calculations, e.g., addition, subtraction, multiplication, logical operations, and the like. The control circuit 18 may be configured to read and execute operations of the program stored in the memory 20. For example, the control circuit 18 may include control elements such as multiplexers and flip-flops that route data between the memory 20, input buffers, or bus 16 and the appropriate computational units 22 to perform the scheduled operations. As discussed herein, the operation schedule of the program may act as a series of timers that trigger the control circuit 18 to begin executing a particular function at a particular counter value. Thus, the operation schedule may trigger the control elements of the control circuit 18 to route data in the hardware tile 14 to the appropriate computational units 22 to perform the scheduled operations at the scheduled times.
[0027] As described herein, program data defining an operation schedule for the hardware tiles must first be provided to each hardware tile 14 and initialized on each hardware tile 14. To reach tiles that are not directly coupled to the communication interface 12, a group of hardware tiles 14 (e.g., tiles A and B) may be individually placed in a data transfer to establish a data transfer path. Program data for the individual tiles 14 may then be sent from the communication interface 12 along the transfer path. After each hardware tile 14 receives its corresponding program data, the tile 14 may automatically reconfigure after a predetermined amount of time into a new, non-data transfer state. In this new state, the tile 14 may be initialized to perform the operations defined by the program data or, alternatively, to perform some other function.
[0028] 2 is a flow chart illustrating an example process (30) for loading and initializing a hardware tile to execute a predefined program. The process (30) is described with respect to the example chip 10 shown in FIG. 1, but is applicable to other chips that include an array of hardware tiles. In this disclosure, a configuration data packet includes configuration data that causes a tile to change state, e.g., from a state in which the tile is operable to read data in a received data packet to a state in which the tile is operable to forward the received data packet, or from a state in which the tile is operable to forward the received data packet to a state in which the tile is operable to read data in the received data packet. In this disclosure, a program data packet includes program data that defines an operation schedule for the hardware tile.
[0029] In a first step (32), a data transfer path is established along a group of hardware tiles 14 (e.g., tiles A, B, and C) in the array of hardware tiles 14. Establishing the data transfer path allows data from the communications interface 12 to reach internal tiles 14 that are not directly coupled to the interface 12 by a communications bus 16. Establishing the data transfer path includes configuring (33) each hardware tile 14 except for the last hardware tile in the array (e.g., hardware tile C) to be in a forwarding state for a corresponding predetermined amount of time. While in the forwarding state, the tile 14 is operable to forward received data packets over one or more data links to one or more of the other tiles 14 in the data transfer path.
[0030] The process of configuring each hardware tile 14 except the last hardware tile 14 in the array to be in a forwarding state (33) includes sequentially sending corresponding configuration data to each hardware tile 14 except the last tile 14 that configures the tile 14 to be in a forwarding state. Sending the configuration data can be accomplished in a number of different ways. For example, in some implementations, sending the configuration data includes sending a single data packet having multiple headers to a first tile in the array. Each header of the data packet may include tile-agnostic data that configures the tile to change state. Upon receiving the data packet, the first tile in the array reads the first header of the data packet and changes its state based on the configuration data of the first header. For example, the configuration data may change the state of the tile to a forwarding state. The data packet may then be passed by the first tile to a next (second) tile in the array, where the second tile reads the second header of the data packet and changes its state based on the configuration data of the second header. For example, the configuration data may change the state of the second tile to a forwarding state. This may continue for each tile in the array, or for only the first N-1 tiles of an N-tile array.
[0031] In some implementations, separate data packets are sent to the tiles, each data packet having its own configuration data. For example, a first data packet with configuration data may be sent to a first tile, which may then install the configuration data (e.g., to change the first tile state to a forwarding state). A second data packet with configuration data may then be sent to the first tile, which forwards the second data packet to a second tile in the array. Upon receiving the second data packet, the second tile may install the configuration data (e.g., to change the second tile state to a forwarding state). This may continue for each tile in the array, or for only the first N-1 tiles of the N-tile array.
[0032] An exemplary process of establishing a forwarding data path along hardware tiles A, B, and C is also illustrated in the time flow diagram of FIG. 3. For example, in one example, a data source, such as communication interface 12, begins at time t1 by sending a first data packet including configuration data to the first hardware tile A shown in FIG. 1. Upon receiving the first data packet including the configuration data, the first hardware tile A does not store the first data packet. Instead, in response to receiving the first data packet, the first hardware tile A is configured to operate in a data forwarding state. In the data forwarding state, hardware tile A forwards the data it receives rather than storing the received data in local memory. For example, the configuration data may configure tile A to forward data to the next adjacent tile in the array, such as tile B. Thus, the configuration data includes data that, when installed in the first tile A, places the first tile in a forwarding state. The first hardware tile A is able to read the first data packet because it is set to a "listening for configuration packets" state. That is, the first hardware tile A is already configured at this point to read data packets arriving at the first hardware tile. Such configuration state for each tile 14 may, in some implementations, be set after a reset operation in which the chip 10 is reset. Thus, the first packet sent, i.e., the packet that constitutes the "data transfer state", contains tile-agnostic data that simply configures the tile to transfer data and stop "listening".
[0033] More generally, each hardware tile in the array of hardware tiles may be initialized to a state called a "listening for configuration packets" state, also referred to herein as a "listening" state. As mentioned above, this listening for configuration state may be performed after a reset operation in which the chip 10 is reset. Upon receiving a data packet in the listening for configuration state, each tile in the listening state reads and "consumes" a header containing tile-agnostic data that configures the tile to change state, e.g., switch to a forwarding state, and ceases "listening" for configuration data. Thus, tiles configured to be in a listening state may not require decode logic. Additionally, note that in some implementations, since each block consumes one header, for a particular state change operation, there are as many headers containing tile configuration data provided to the array as there are hardware blocks along the path. For example, modifying N tiles in a similar manner (e.g., to create a forwarding path consisting of N tiles) may require sending a series of N consecutive data packets, each of which includes configuration data that changes the tile to a different state (e.g., configuration data that changes the tile from a listening state to a forwarding state).
[0034] Thus, after sending the first data packet, the data source (e.g., communications interface 12) sends a second data packet including configuration data to the first hardware tile A at time t2. Because hardware tile A is in the forwarding state, tile A forwards the second data packet to another tile in the array. For example, tile A may forward the second data packet to hardware tile B. Upon receiving the second data packet including the configuration data, the second hardware tile B, which is in the listening state, does not store the second data packet. Instead, in response to receiving and reading the second data packet, the second hardware tile B is configured to operate in a data forwarding state. In the data forwarding state, hardware tile B forwards the data it receives rather than storing the received data in local memory. For example, the configuration data in the second data packet may configure tile B to forward data to the next adjacent tile in the array, such as tile C.
[0035] In view of the above, both tiles A and B are in a forwarding state establishing a data transfer path along hardware tiles A, B, and C. Because hardware tile C is the last tile in the array, there is no need to configure hardware tile C to be in a forwarding state. Rather, tile C can remain configured in a "listening" state where it is ready to receive and install program data for performing scheduled operations, as described herein.
[0036] Because it may also be desirable to receive and install program data among the remaining hardware tiles, hardware tiles configured in the data forwarding state are scheduled to sequentially exit the data forwarding state at predetermined times. For example, after the last hardware tile (e.g., hardware tile C) has received and installed its program data, tiles A and B may be configured to sequentially exit their respective data forwarding states. In this manner, each hardware tile in the forwarding state shifts back to a state in which it can receive and install program data intended for that hardware tile.
[0037] Configuring the hardware tiles to exit the forwarding state may include, for example, installing a respective forwarding state counter on each tile that was placed in the forwarding state. The installation of the forwarding state counter in the tile may occur as part of the change of the tile's state from the listening state to the forwarding state, for example, when the tile receives a configuration data packet. The forwarding state counter includes a counter that counts down (or up) until a predefined time period has elapsed. While the forwarding state counter counts down (or up), the hardware tile on which the counter is installed remains in the forwarding state. When the forwarding state counter reaches the predefined time period, the counter may induce a trigger to fire that causes the hardware tile to exit the forwarding state. For example, in some implementations, a trigger may fire that causes the hardware tile 14 to reconfigure such that for any new data packets received at the tile 14, the data packets are saved to a local memory on the tile 14 (e.g., memory 20). Alternatively, or in addition, the trigger may cause the hardware tile 14 to reconfigure to install data from the memory or to install data from any new data packets that the tile 14 receives. Alternatively, or in addition, the trigger may cause the hardware tile 14 to reconfigure so that the tile returns to a listening state. The counter installed on the tile 14 may be synchronized with a global clock on the chip 10 or a local clock running on the tile 14. The predefined time period may include multiple clock cycles. For example, the predefined time period may include 2, 5, 10, 20, 50, or 100 clock cycles, among others. The predefined time period may be in the range [2, 100] clock cycles, e.g., in the range [10, 50] clock cycles, such as 20 clock cycles. The predefined time period for the counter may be defined in the configuration data that is read by the tile upon receiving the configuration data packet.
[0038] Hardware tiles 14 configured in a forwarding state (e.g., tiles A and B) should not all exit their respective forwarding states at once. Instead, because the hardware tiles 14 in the array share a forwarding data path, they each receive their program data packets at different times. Thus, each tile exits its forwarding state at a different time, and the default length of time for each forwarding state counter is different. The default length of time specified for each forwarding state counter is a function of the number of hardware tiles in the data forwarding path to which the current hardware tile forwards data packets. For example, in the data forwarding path established by tiles A, B, and C in FIG. 1, there are two hardware tiles (B and C) to which the first tile A forwards data. Thus, the default period of time associated with the forwarding state counter on tile A is a function of the time to send program data to tiles B and C and to install the program data in tiles B and C. With reference to Figure 3, this may occur, for example, at a time equal to the difference between when a first hardware tile A is configured in a forwarding state (at t1) and when the first hardware tile A receives a data program packet intended to be installed on tile A (at t5), i.e., at a time t5-t1. Similarly, in the data forwarding path established by tiles A, B, and C of Figure 1, there is one hardware tile (C) to which a second tile B forwards data. Thus, the predefined time period associated with the forwarding state counter on tile B is a function of the time to send program data to tile C and install the program data on tile C. With reference to Figure 3, this may occur, for example, at a time equal to the difference between when a second hardware tile B is configured in a forwarding state (at t2) and when the second hardware tile B receives a data program packet intended to be installed on tile B (at t4), i.e., at a time t4-t2.The particular time associated with each forwarding state counter may be pre-calculated and included in the configuration data provided with the first and second data packets. Hardware tile C is not set to a forwarding state and is therefore not configured to include a forwarding state counter.
[0039] After the transfer data path is established, the data source (e.g., communications interface 12) provides their respective program data packets to the hardware tiles 14 along the data transfer path (34). The program data packets include program data that defines an operational schedule for the hardware tiles. The program data packets may be provided sequentially such that the last hardware tile in the data transfer path (e.g., the hardware tile furthest from the data source) receives its program data first, while the first hardware tile in the data transfer path (e.g., the hardware tile closest to the data source) receives its program data last. As the program data packets are received, they may be installed on the hardware tiles (36).
[0040] 1, in the data transfer path established by hardware tiles A, B, and C, communications interface 12 sends a first program data packet at time t3 to hardware tile A, which then forwards the first program data packet to hardware tile B, which then forwards the first program data packet to hardware tile C. Upon receiving the first program data packet, hardware tile C may save the first program data packet in local memory (e.g., memory 20) and / or install the program data contained within the packet.
[0041] Communications interface 12 then sends the second program data packet to hardware tile B at time t4. Before hardware tile B can save and / or install the program data contained in the second program data packet, a forwarding state counter installed on hardware tile B fires a trigger that causes hardware tile B to change from a forwarding state to a new state configured to save and / or install data packets that hardware tile B receives. As described herein, this trigger may occur at time t4-t2. In this manner, the second program data packet is retained in hardware tile B rather than being forwarded to hardware tile C.
[0042] Communications interface 12 then sends a third program data packet to hardware tile A at time t5. Before hardware tile A can save and / or install the program data contained in the second program data packet, a forwarding state counter installed on hardware tile A fires a trigger that causes hardware tile A to change from a forwarding state to a new state configured to save and / or install data packets that hardware tile A receives. As described herein, this trigger may occur at time t5-t1. In this manner, the third program data packet is retained in hardware tile A rather than being forwarded to hardware tile B.
[0043] In some implementations, it is desirable for all hardware tiles in the array to execute their installed / initialized programs simultaneously. In such cases, the hardware tiles that first receive and initialize / install their program data wait until the other hardware tiles in the array have also received and initialized / installed their respective program data. The time at which all hardware tiles begin executing their installed program data is referred to as the "kick-off" time. To ensure that each hardware tile kicks off at the same time, installing the data program packet (36) may include, for example, configuring each hardware tile to include a corresponding program kick-off counter, also referred to herein as a kick-off state counter (37).
[0044] The kickoff state counter includes a counter that counts down (or up) until a predefined time period has elapsed. The predefined time period specified for each hardware tile specifies the time until the actions defined in the program data are initiated within that tile. While the kickoff state counter counts down (or up), the hardware tile on which the counter is installed remains in a holding state during which no actions are executed. When the kickoff state counter reaches the predefined time period, the counter may induce a trigger to be initiated that causes the hardware tile to begin executing the actions defined by the program data within the tile. The predefined time period for each kickoff state counter is calculated to be a value that results in all hardware tiles in the array simultaneously executing their installed / initialized programs. As described herein, the predefined time period for each kickoff state counter may be different.
[0045] For example, in some implementations, a trigger may fire that causes the hardware tile 14 to begin executing operations defined in program data previously received at the tile 14 and stored in the tile's local memory (e.g., memory 20). The predefined time period may include multiple clock cycles. For example, the predefined time period may include 2, 5, 10, 20, 50, or 100 clock cycles, among others. As described herein, the kickoff state counter installed on the tile 14 is synchronized with other kickoff state counters with respect to a global clock.
[0046] Because the hardware tiles 14 in the array share a transfer data path, they each receive and install their program data packets at different times. That is, each tile 14 waits a different amount of time before it can begin to perform an operation, and the default length of time for each kickoff state counter is different. The default length of time specified for each kickoff state counter may be, for example, a function of the number of hardware tiles that must still receive and install their program data. This ensures that all tiles in the array execute their installed / initialized programs simultaneously. For example, in the data transfer path established by tiles A, B, and C of FIG. 1, after hardware tile C receives its program data, tile C waits until it installs its own program data and until the other two hardware tiles (A and B) receive and install their respective program data. Thus, the default period of time associated with the kickoff state counter on tile C may be, for example, a function of the time required to transmit program data to tiles A and B and install the program data in tiles A and B, as well as the time to install the program data in tile C. Referring to FIG. 3, this may be done, for example, at a time equal to the difference between when all hardware tiles kick off (at t6) and when the last hardware tile C receives a data program packet intended to be installed on tile C (at t3), i.e., at a time t6-t3.
[0047] Similarly, after hardware tile B receives its program data, it waits until it installs its own program data and until one other hardware tile (tile A) has received and installed its respective program data. Thus, the predefined time period associated with the kickoff state counter on tile B may be, for example, a function of the time to send program data to and install the program data on tile A, as well as the time to install the program data on tile B. With reference to Figure 3, this may occur, for example, at a time equal to the difference between when all hardware tiles kick off (at t6) and when hardware tile B receives a data program packet that is intended to be installed on tile B (at t4), i.e., at a time t6-t4.
[0048] Similarly, after hardware tile A receives its program data, it waits to install its own program data before it can kick off execution of the program data. Thus, the predefined time period associated with the kickoff state counter on tile A may be, for example, a function of the time required to install the program data on tile A. With reference to Figure 3, this may occur, for example, at a time equal to the difference between when all hardware tiles kick off (at t6) and when hardware tile A receives a data program packet that is intended to be installed on tile A (at t5), i.e., at a time t6-t5.
[0049] A specific time associated with each kickoff state counter may be pre-calculated and included in the program data provided with each program data packet. As described herein, a different specific time associated with each kickoff state counter may be calculated and defined such that all tiles execute their installed / initialized programs simultaneously. In some implementations, each tile in the array may be configured to include a corresponding kickoff state counter. For example, tiles A, B, and C may each be configured to include a corresponding kickoff state counter with a different predefined time period to wait to execute its stored program data.
[0050] In some implementations, at least some of the tiles in the array may be configured to include a corresponding kick-off state counter. For example, if there are N tiles in the array, then N-1 tiles may be configured to include a corresponding kick-off state counter. This may include configuring all of the tiles in the array except the first tile of the array to include a corresponding kick-off state counter. Using the example of the present application, tiles B and C may each be configured to include a corresponding kick-off state counter with a different predefined period of time to wait to execute its stored program data. However, tile A may not be configured to have a corresponding kick-off state counter. In this case, tiles B and C may wait until their kick-off state counters are triggered to execute their received program data, but tile A may immediately execute the program data upon receiving the program data such that tiles A, B, and C each execute their respective program data simultaneously.
[0051] As described herein, the chips described herein differ from traditional processors in that instructions are issued every cycle and contain source and destination registers for various configurations of the chip's functional units. Instead, each hardware tile on the chip is controlled by a logical set of states known as configuration states. A configuration state represents whether the hardware tile is configured to transfer data, store data in memory, install data, or execute a program, among other functions. Depending on the particular state the hardware tile is configured to, the state specifies control signals for the various multiplexers in the hardware tile, as well as read and write operations to the memory and registers in the hardware tile. The configuration states of the hardware tiles are updated (e.g., switched to a transfer state or other state) through a trigger table.
[0052] 4 is a schematic diagram of an example of a trigger table 40. The trigger table 40 is an addressable set of configuration state updates, each of which becomes applied when a particular condition is met. For example, when a forwarding state counter is installed in a hardware tile as described herein, a condition and associated configuration may be added to the trigger table 40. When the forwarding state counter counts down to zero (or counts up to a predefined period of time), the condition added to the table 40 is met, so a trigger is fired and the configuration associated with that condition is applied to the hardware tile. A kick-off state counter may be implemented similarly.
[0053] The trigger table 40 includes a number of entries 50. Each entry 50 may include, for example, a trigger ID 52, an address 54, a configuration space update 56, and an enable flag 58 and one or more additional flags 60. The trigger ID 52 represents a trigger type and acts as a pointer to any associated state. The combination of the address 54 and the configuration space update 56 represents how the configuration state of the hardware tile should be updated. The enable flag 58 specifies whether the trigger is currently active and may fire at any time. The additional flags 60 may specify other aspects related to the trigger, such as whether the trigger fired within the last clock cycle. The trigger table 40 may include multiple entries, including, for example, 8, 16, 32, or 64 entries. The trigger table 40 may be implemented by locally storing different configuration states of the trigger table 40 in memory and using one or more multiplexers in the hardware tile to select a state.
[0054] A particular implementation of modifying the configuration state of a tile has been described. However, other implementations are possible. For example, in some implementations, rather than creating a forwarding path and then sequentially loading the program data into the tile as described herein, program data may be installed on the tile simultaneously with configuring the tile to be in a forwarding state and installing a kick-off state counter and starting a program in the tile. Using the tile structure shown in FIG. 1, for example, an alternative or additional exemplary process for loading program data is represented. For example, in some implementations, the communication interface 12 sends a first data packet to a first tile (e.g., Tile A) that includes configuration data (e.g., as a header of the data packet) and program data (e.g., as a payload of the data packet). The configuration data, when read by Tile A, causes Tile A to store the program data from the data packet in a local memory and / or configure logic to perform a predefined set of operations. Additionally, the configuration packet may configure Tile A to be placed in a forwarding state while simultaneously initializing a kick-off state counter. As described herein, the kickoff state counter may be a counter that counts a predetermined amount of time until the tile is forced to change its configuration state so that a program installed on the tile begins to execute (i.e., "kicks off"). The amount of time that the kickoff state counter in tile A counts depends on the amount of data required for the tile that follows tile A, rather than the amount of data required for the tile that appears before tile A. For example, if there are two tiles following tile A, the value of the kickoff state counter in tile A is determined based on the time it takes to transfer configuration and program data to the next two tiles.
[0055] Additionally, note that configuring a tile to install program data while simultaneously placing the tile in a transfer state may include setting up a separate timer for the transfer state that is set to 0 clock cycles.
[0056] After the first tile is in the forwarding state, the communication interface 12 sends out a second data packet. This second data packet may pass through the first tile in the forwarding state and may arrive at a second tile (e.g., tile B) in the array. The second data packet is similar to the first data packet. For example, the second data packet includes configuration data (e.g., as a header of the second data packet) and program data (e.g., as a payload of the second data packet). The configuration data, when read by the second tile, causes the second tile to store the program data from the second data packet in a local memory and / or configure a logic circuit to perform a predefined set of operations. Additionally, the configuration packet may configure the second tile to be placed in the forwarding state while simultaneously initializing a kickoff state counter. The value of the kickoff state counter in the second tile is determined by the amount of data to be forwarded and configured by the remaining tiles after the second tile in the array.
[0057] In this implementation, each tile in the path is configured in a similar manner to the first and second tiles except for the last tile in the path (e.g., tile C). For the last tile, the communication interface 12 sends a last data packet. The last data packet includes configuration data and program data. In contrast to the previous data packets, the last tile is not configured to be in a forwarding state. Instead, the last tile is configured to install program data and begin to perform operations using the program data after installation is complete. At the same time, the kick-off state counters for the previous tiles in the path (e.g., tiles A and B) reach their limits, causing their respective tiles to begin performing operations using the program data saved locally to those tiles. In this manner, the tiles may be said to be directly programmed by the communication interface, where the forwarding path is set up once the program data is installed on the tiles and the kick-off state counters are initialized, rather than being indirectly programmed by first setting up the forwarding path and then sending down the program data for each tile after the forwarding path is established, as also described herein.
[0058] 5 is a schematic diagram illustrating an example of a dedicated logic circuit that can be configured and initialized to operate according to scheduled operations as described herein. For example, the dedicated logic circuit may include an ASIC 100. The ASIC 100 includes various different types of hardware blocks that can be configured to perform the overall operation of the ASIC 100 according to individual operation schedules. Exemplary hardware blocks that can operate according to individual operation schedules include tiles 102 and vector processing units 104 (similar to hardware tiles 14 in FIG. 1) and communication interfaces 108 (similar to communication interfaces 12 in FIG. 1).
[0059] More specifically, the ASIC 100 includes dedicated circuitry configured to cause one or more of the tiles 102 to perform operations, such as, for example, multiplication and addition operations. Specifically, each tile 102 can include a computational array of cells (e.g., similar to the computational unit 22 of FIG. 1 ), in which each cell is configured to perform a mathematical operation (see, for example, the example tile 200 shown in FIG. 6 and described herein). In some implementations, the tiles 102 are arranged in a grid pattern and are arranged along a first dimension 101 (e.g., rows) and a second dimension 103 (e.g., columns). For example, in the example shown in FIG. 5 , the tile 102 is divided into four different portions (110a, 110b, 110c, 110d), each portion including 288 tiles arranged in a grid of 18x16 tiles. In some implementations, the ASIC 100 shown in FIG. 5 may be understood as including a single systolic array of cells divided / arranged as separate tiles, where each tile includes a subset / sub-array of cells, local memory, and bus lines (see, for example, FIG. 6).
[0060] The ASIC 100 also includes a vector processing unit 104. The vector processing unit 104 includes circuitry configured to receive outputs from the tiles 102 and calculate vector computation output values based on the outputs received from the tiles 102. For example, in some implementations, the vector processing unit 104 includes circuitry (e.g., multiplier circuitry, adder circuitry, shifters, and / or memory) configured to perform a multiply-accumulate operation on the outputs received from the tiles 102. Alternatively or additionally, the vector processing unit 104 includes circuitry configured to apply a non-linear function to the outputs of the tiles 102. Alternatively or additionally, the vector processing unit 104 generates a normalization value, a merged value, or both. The vector computation output of the vector processing unit can be stored in one or more tiles. For example, the vector computation output can be stored in a memory uniquely associated with the tile 102. Alternatively or additionally, the vector computation output of the vector processing unit 104 can be transferred to a circuitry external to the ASIC 100, for example, as an output of the computation. Additionally, operation of separate operation schedules for tiles 102 and vector processing units 104 coordinate the transfer of tile outputs to the vector processing units 104 .
[0061] In some implementations, the vector processing unit 104 is segmented, whereby each segment includes circuitry configured to receive outputs from a corresponding set of tiles 102 and calculates a vector computation output based on the received outputs. For example, in the example shown in FIG. 5, the vector processing unit 104 includes two rows spanning along the first dimension 101, with each row including 32 segments 106 arranged as 32 columns. Each segment 106 includes circuitry (e.g., multiplier circuitry, adder circuitry, shifters, and / or memory) configured to perform a vector computation, as described herein, based on the outputs (e.g., sum of products) from a corresponding column of tiles 102. The vector processing unit 104 may be located in the center of the grid of tiles 102 as shown in FIG. 5. Other arrangements of the vector processing unit 104 are possible.
[0062] The ASIC 100 also includes a communication interface 108 (e.g., interfaces 1010A, 1010B). The communication interface 108 includes one or more sets of serializer / deserializer (SerDes) interfaces and a general purpose input / output (GPIO) interface. The SerDes interfaces are configured to receive instructions (e.g., operation schedules for individual hardware blocks of the ASIC 100) and / or input data for the ASIC 100 and output data from the ASIC 100 to external circuitry. For example, the SerDes interfaces can be configured to send and receive data (e.g., operation schedules and / or input / output data) at 32 Gbps, 56 Gbps, or any suitable data rate via the set of SerDes interfaces included within the communication interface 108. For example, the ASIC 100 may execute a boot program when turned on. The GPIO interface may be used to load an operation schedule onto the ASIC 100 to execute a particular type of machine learning model.
[0063] The ASIC 100 further includes a communication interface 108, a vector processing unit 104, and a plurality of controllable bus lines (see, e.g., FIG. 6) configured to carry data between the plurality of tiles 102. The controllable bus lines include, for example, wires extending along both a first dimension 101 of the grid (e.g., rows) and a second dimension 103 of the grid (e.g., columns). A first subset of the controllable bus lines extending along the first dimension 101 can be configured to transfer data in a first direction (e.g., the right side of FIG. 5). A second subset of the controllable bus lines extending along the first dimension 101 can be configured to transfer data in a second direction (e.g., the left side of FIG. 5). A first subset of the controllable bus lines extending along the second dimension 103 can be configured to transfer data in a third direction (e.g., the top side of FIG. 5). A second subset of the controllable bus lines extending along the second dimension 103 can be configured to transfer data in a fourth direction (e.g., the bottom of FIG. 5). As described above, the separate operating schedules of different hardware blocks can coordinate access to shared resources, such as the controllable bus lines, to prevent communication errors within the ASIC 100.
[0064] Each controllable bus line includes a plurality of conveyor elements, such as flip-flops, used to convey data along each line according to a clock signal. Transferring data over the controllable bus lines can include shifting data from a first conveyor element of the controllable bus line to a second adjacent conveyor element of the controllable bus line at each clock cycle. In some implementations, data is conveyed over the controllable bus lines on a rising or falling edge of a clock cycle. For example, data present on a first conveyor element (e.g., a flip-flop) of the controllable bus line at a first clock cycle can be transferred to a second conveyor element (e.g., a flip-flop) of the controllable bus line at a second clock cycle. In some implementations, the conveyor elements can be periodically spaced apart from one another at a fixed distance. For example, in some cases, each controllable bus line includes a plurality of conveyor elements, each of which is disposed within or proximate to a corresponding tile 102.
[0065] Each controllable bus line also includes a number of multiplexers and / or demultiplexers. The multiplexers / demultiplexers of the controllable bus lines are configured to transfer data between the bus lines and components of the ASIC chip 100. For example, the multiplexers / demultiplexers of the controllable bus lines can be configured to transfer data to and from the tiles 102, the vector processing units 104, or the communication interfaces 108. Transferring data between the tiles 102, the vector processing units 104, and the communication interfaces can be coordinated by an operation schedule. The operation schedule can coordinate which ASIC 100 hardware blocks are transmitting to or receiving data from the controllable bus lines in each counter. The operation scheduled at any given counter time may determine, for example, what data is transferred from a source (e.g., a memory 102 or a vector processing unit 104 in a tile 102) to a controllable bus line, or alternatively, what data is transferred from the controllable bus line to a sink (e.g., a memory 102 or a vector processing unit 104 in a tile 102).
[0066] The controllable bus lines are configured to be controlled at a local level such that each tile, vector processing unit, and / or communication interface includes its own set of control elements for operating the controllable bus lines passing through that tile, vector processing unit, and / or communication interface. For example, each tile, 1D vector processing unit, and communication interface may include a corresponding set of conveyor elements, multiplexers, and / or demultiplexers for controlling data transfer to and from that tile, 1D vector processing unit, and communication interface. Thus, an operation schedule for each tile, 1D vector processing unit, and communication interface can trigger a respective hardware block to provide appropriate control signals to its conveyor elements to route data according to the scheduled operation.
[0067] To minimize latency associated with the operation of the ASIC chip 100, the tiles 102 and vector processing units 104 can be arranged to reduce the distance that data travels between the various components. In certain implementations, both the tiles 102 and the communication interface 108 can be separated into multiple portions, with both the tile portions and the communication interface portions arranged to reduce the maximum distance that data travels between the tiles and the communication interface. For example, in some implementations, a first group of tiles 102 can be arranged in a first portion on a first side of the communication interface 108, and a second group of tiles 102 can be arranged in a second portion on a second side of the communication interface. As a result, the farthest tile may be half the distance from the communication interface compared to a configuration in which all tiles 102 are arranged in a single portion on one side of the communication interface.
[0068] Alternatively, the tiles may be arranged as a different number of portions, such as four portions. For example, in the example shown in FIG. 5, the tiles 102 of the ASIC 100 are arranged as a number of portions 110 (110a, 110b, 110c, 110d). Each portion 110 includes a similar number of tiles 102 arranged as a grid pattern (e.g., each portion 110 may include 256 tiles arranged in 16 rows and 16 columns). The communication interface 108 is also divided into a number of portions, namely a first communication interface 1010A and a second communication interface 1010B arranged on either side of each portion 110 of the tiles 102. The first communication interface 1010A may be coupled to the two tile portions 110a, 110c on the left side of the ASIC chip 100 through a controllable bus line. The second communication interface 1010B may be coupled to the two tile portions 110b, 110d on the right side of the ASIC chip 100 through a controllable bus line. As a result, the maximum distance that data travels to and from the communications interface 108 (and therefore the latency associated with data propagation) can be halved compared to a configuration in which only a single communications interface is available. Other coupling arrangements of the tiles 102 and communications interfaces 108 are possible to reduce data latency. The coupling arrangement of the tiles 102 and communications interfaces 108 can be programmed by providing control signals to the conveyor elements and multiplexers of the controllable bus lines.
[0069] In some implementations, one or more tiles 102 are configured to initiate read and write operations to controllable bus lines and / or other tiles in the ASIC 100 (referred to herein as "control tiles"). The remaining tiles in the ASIC 100 can be configured to perform calculations based on input data (e.g., to compute layer inference). In some implementations, the control tiles include the same components and configurations as other tiles in the ASIC 100. The control tiles can be added as additional tiles, additional rows, or additional columns of the ASIC 100. For example, in a symmetric lattice of tiles 102, each tile 102 is configured to perform calculations on input data, and one or more additional rows of control tiles can be included to accommodate read and write operations for tiles 102 that perform calculations on input data. For example, each portion 110 may include 18 rows of tiles, with the last two rows of tiles including control tiles. Providing separate control tiles increases the amount of memory available to other tiles used to perform calculations in some implementations. Providing separate control tiles may also aid in the coordination of data transmission operations between operation schedules. For example, using control tiles to control read and write operations to controllable bus lines and / or other tiles in the ASIC 100 may reduce the number of separate schedules that need to be checked for scheduling conflicts. In other words, if the operation schedules for the control tiles are coordinated to avoid "double booking", i.e., using certain controllable bus lines at the same counter time, there is a reasonable guarantee that no communication errors will occur on the controllable bus lines. However, a separate tile dedicated to performing the controls described herein is not required, and in some cases, a separate control tile is not provided. Instead, each tile may store instructions in the tile's local memory to initiate read and write operations for that tile.
[0070] 5 includes tiles arranged in 18 rows and 16 columns, the number of tiles 102 and the arrangement of the tiles within a given portion may vary. For example, in some cases, each portion 110 may include an equal number of rows and columns.
[0071] Further, although the tiles 102 are shown in FIG. 5 as being divided into four parts, they may be divided into other different groups. For example, in some implementations, the tiles 102 are grouped into two different parts, such as a first part above the vector processing unit 104 (e.g., closer to the top of the page shown in FIG. 5) and a second part below the vector processing unit 104 (e.g., closer to the bottom of the page shown in FIG. 5). In such an arrangement, each part may include, for example, 596 tiles arranged as a grid of 18 tiles (along the direction 103) by 32 tiles (along the direction 101). Each part may include other total numbers of tiles and may be arranged as arrays of different sizes. In some cases, the partition between each part is defined by the hardware elements of the ASIC 100. For example, as shown in FIG. 5, the parts 110a, 110b may be separated from the parts 110c, 110d by the vector processing unit 104.
[0072] A schematic diagram illustrating an example of a tile 200 used in the ASIC chip 100 is shown in FIG. 6. Each tile 200 includes a local memory 202 and a computational array 204 coupled to the memory 202. The local memory 202 includes a physical memory located proximate to the computational array 204. The computational array 204 includes a plurality of cells 206. Each cell 206 of the computational array 204 includes a circuit configured to perform a computation (e.g., a multiply-accumulate operation) based on data inputs, such as activation inputs and weight inputs, to the cell 206. Each cell can perform a computation (e.g., a multiply-accumulate operation) for a cycle of a clock signal. The computational array 204 can have more rows than columns, or more columns than rows, or an equal number of columns and rows. For example, in the example shown in FIG. 6, the computational array 204 includes 64 cells arranged in 8 rows and 8 columns. Other computational array sizes are possible, such as computational arrays having 16, 32, 128, or 256 cells, among others. Each tile can contain the same number of cells and / or the same size computation array. The total number of operations that can be performed in parallel for an ASIC chip depends on the total number of tiles with the same size computation array in the chip. For example, in the ASIC chip 100 shown in FIG. 5, which contains about 1150 tiles, this means that about 92000 calculations can be performed in parallel per cycle. Examples of clock speeds that may be used include, but are not limited to, 225 MHz, 500 MHz, 950 MHz, 1 GHz, 1.25 GHz, 1.5 GHz, 1.95 GHz, or 2 GHz. The computation array 204 of each individual tile is a subset of the larger systolic array of the tile, as shown in FIG. 5.
[0073] The memory 202 included in the tile 200 may include, for example, a random access memory (RAM), such as an SRAM. Other memories may be used instead. Each memory 202 may be configured to store 1 / n of the total memory associated with the n tiles 102 of the ASIC chip. The memory 202 may be provided as a single chip or multiple chips. For example, the memory 202 shown in FIG. 6 is provided as four single-port SRAMs, each SRAM coupled to the computational array 204. Alternatively, the memory 202 may be provided as two single-port SRAMs or eight single-port SRAMs, among other configurations. The combined capacity of the memory may be, for example, but not limited to, 16 kB, 32 kB, 64 kB, or 128 kB after error correction coding. By providing the physical memory 202 local to the computational array, the wiring density for the ASIC 100 may be significantly reduced in some implementations. Alternative configurations in which memory is centrally located within the ASIC 100 may require wiring for each bit of memory bandwidth as opposed to a locally provided configuration as described herein. The total number of wires required to service each tile of the ASIC 100 may significantly exceed the available space within the ASIC 100. In contrast, providing dedicated memory for each tile may substantially reduce the total number of wires required to service the area of the ASIC 100.
[0074] The tile 200 also includes controllable bus lines, which may be categorized into a number of different groups. For example, the controllable bus lines may include a first group of general purpose controllable bus lines 210 configured to transfer data between tiles in each cardinal direction. That is, a first group of controllable bus lines 210 may include bus lines 210a configured to transfer data in a first direction along the first dimension 101 of the grid of tiles (referred to as "East" in FIG. 6 ), bus lines 210b configured to transfer data in a second direction along the first dimension 101 of the grid of tiles (referred to as "West" in FIG. 6 ), where the second direction is opposite to the first direction, bus lines 210c configured to transfer data in a third direction along the second dimension 103 of the grid of tiles (referred to as "North" in FIG. 6 ), and bus lines 210d configured to transfer data in a fourth direction along the second dimension 103 of the grid of tiles (referred to as "South" in FIG. 6 ), where the fourth direction is opposite to the third direction. The general-purpose bus lines 210 can be configured to carry control data, activation input data, carry data to and from the communication interface, carry data to and from the vector processing unit, and carry data (e.g., weight inputs) that are stored and / or used by the tile 200. The tile 200 may include one or more control elements 221 (e.g., flip-flops and multiplexers) for controlling the controllable bus lines and thus routing data to and from the tile 200 and to the memory 202.
[0075] The controllable bus lines may also include a second group of controllable bus lines, also referred to herein as computational array partial sum bus lines 220. The computational array partial sum bus lines 220 may be configured to carry data output from computations performed by the computational array 204. For example, the bus lines 220 may be configured to carry partial sum data obtained from rows in the computational array 204, as shown in FIG. 6. In such a case, the number of bus lines 220 matches the number of rows in the array 204. For example, in an 8×8 computational array, there are eight partial sum bus lines 220, with each partial sum bus line 220 coupled to an output of a corresponding row in the computational array 204. The computational array output bus lines 220 may be further configured to couple to another tile in the ASIC chip, e.g., as an input to the computational array of another tile in the ASIC chip. For example, the array partial sum bus lines 220 of tile 200 can be configured to receive an input (e.g., partial sum 220a) of a computational array of a second tile located at least one tile away from tile 200. The output of computational array 204 may then be added to partial sum lines 220 to create a new partial sum 220b, which may be output from tile 200. Partial sum 220b may then be passed to another tile or, alternatively, to a vector processing unit. For example, each bus line 220 may be coupled to a corresponding segment (such as segment 106 in FIG. 5) of the vector processing unit.
[0076] As described with respect to FIG. 5, the controllable bus lines can include circuits such as conveyor elements (e.g., flip-flops) configured to allow data to be transmitted along the bus lines. In some implementations, each controllable bus line includes a corresponding conveyor element for each tile. As described further with respect to FIG. 5, the controllable bus lines can include circuits such as multiplexers configured to allow data to be transferred between various tiles of the ASIC chip, vector processing units, and communication interfaces. The multiplexers can be located wherever there is a source or sink of data. For example, in some implementations, as shown in FIG. 6, control circuitry 221 such as a multiplexer can be located at the intersections of the controllable bus lines (e.g., at the intersections of generic bus lines 210a and 210d, generic bus lines 210a and 210c, generic bus lines 210b and 210d, and / or generic bus lines 210b and 210c). Multiplexers located at bus line intersections can be configured to transfer data between the bus lines located at the intersections. The control circuitry 221 can route data to the appropriate components in the tile 102 (e.g., route activation data or layer weights to or from the SRAM 202 to the appropriate cells 206 in the computational array 204) or perform the operations of an operational schedule by routing output and input data to and from the controllable bus lines.
[0077] 7A-13B are schematic diagrams illustrating an exemplary process in which the ASIC 100 is used as a hardware accelerator for computing neural network inference. 7A, 8A, 9A, 10A, 11, 12A, and 13A are schematic diagrams outlining data flow through the ASIC 100 at different times in each of the processes. 7B, 8B, 9B, 10B, 12B, and 13B are schematic diagrams illustrating data flow within a single tile (e.g., a control tile or other tile 102) of the ASIC 100 at times associated with 7A, 8A, 9A, 10A, 12A, and 13A, respectively. Ovals in 7A-13B indicate the presence of repeating elements that are not shown. A compass 300 is provided in each of 7A-13B to provide an orientation of the data flow. The labels "N", "W", "S", and "E" do not correspond to actual geographic directions, but instead are used to indicate different relative directions that data may flow within the grid. The controllable bus lines that convey data in the directions indicated by the labels "N", "W", "S", and "E" are referred to herein as North Flow Bus Lines, West Flow Bus Lines, South Flow Bus Lines, and East Flow Bus Lines.
[0078] The arrangement of tiles 102 and vector processing units 104 in Figures 7A-13A is similar to the arrangement shown in Figure 5. For example, half of the tiles 102 can be disposed on a first side of the vector processing units 104 and the other half of the tiles 102 can be disposed on a second, opposite side of the vector processing units 104. Although communications interface 108 is shown in Figures 7A-13A as being generally disposed on the right side of the tile grid, it can be disposed on either side of the tile grid as shown generally in Figure 5.
[0079] In a first step, as shown in FIG. 7A, input values (e.g., activation inputs and / or weight inputs) for a first layer of a model (e.g., a neural network model) are loaded onto one or more tiles 102 (e.g., all tiles 102) in the ASIC 100 from the communication interface 108. That is, data, such as configuration data packets, or program operation packets, as described herein with respect to FIGS. 1-3, are received from the communication interface 108. From the communication interface 108, the input values follow a data path along a controllable bus line (e.g., a generic controllable bus line, as described herein) to one or more control tiles. Data can be routed between the different bus lines by using multiplexers (e.g., see routing element 221 in FIG. 6) where the different bus lines cross. For example, as shown in FIG. 7A, input data flows along a data path with a travel on a west-flow generic controllable bus line and then a travel on a south-flow generic controllable bus line. A multiplexer can be used at the intersection of the West and South flow bus lines to transfer input data from the flow bus lines to the South flow bus lines. In some implementations, weight inputs for a second inference can be loaded into one or more control tiles while a previous first inference is being executed by the ASIC 100. In other words, the operation schedule of the control tile is coordinated with the operation schedule of the other tiles 102 that are computing an inference, so that at the same counter time that the other tiles 102 are computing a first inference, the control tile 102 prepares new activation data and / or weights for the next inference to be sent to the other tiles 102 to compute the next inference.
[0080] FIG. 7B is a schematic diagram showing a detailed view of an example of a tile 102 from the ASIC 100. As shown in FIG. 7B, the tile 102 can include a memory 302 in which input values are stored. The memory 302 can include any suitable memory as described herein with respect to FIG. 6. As described above, the memory 302 can be used to store configuration state data (e.g., from a configuration data packet) or program data such as an individual operation schedule of the tile. The input values are obtained from one or more south flow general purpose controllable bus lines 310d that pass next to the tile 102 or pass through the tile 102 itself. Data from the south flow controllable bus line 310d can be transferred to the memory 302 by using a multiplexer. The other general purpose controllable bus lines (310a, 310b, 310c) are not used during this step.
[0081] The tile 102 also includes a computational array of cells 306 directly coupled to the memory 302. As described herein, the computational array of cells 306 may be a subset of a larger systolic array of cells that make up the tile of the ASIC. The cells 306 are arranged in an array, with a single cell 306 shown in FIG. 7B at (i, j) = (0, 0), where the parameter i represents the cell row position in the array and j represents the cell column position in the array. In the example shown in FIG. 7B, the computational array has 8 rows and 8 columns; however, other sizes are possible. Each cell 306 of the computational array may include circuitry configured to perform a calculation based on data received at the tile. For example, each cell 306 may include a multiplier circuit, an adder circuit, and one or more registers. The output of each cell 306 may be passed as a partial sum to adjacent cells in the computational array or to cells in a computational array of another tile in the ASIC 100. The computational array of cells 306 is used in subsequent steps.
[0082] Tile 102 also includes a controllable bus line 320 for providing data from a previous tile. For example, controllable bus line 320 may carry partial sum output data obtained from a computational array of a previous tile in ASIC 100 and provide the partial sum output data as input to cells of a computational array in tile 102. In this step, controllable bus line 320 is not used.
[0083] The tile 102 also includes a controllable bus line 330 for providing an activation input value as an input to the cell 306 of the computational array. For example, the activation input value may be provided to a multiplier circuit in the cell 306. The activation input value may be obtained from the communication interface 108 or a cell in another tile in the ASIC 100. Data from the controllable bus line 330 may be transferred to the cell 306 by using a multiplexer. The controllable bus line 330 is not used in the exemplary steps shown in Figures 7A and 7B.
[0084] As described herein, in some implementations, one or more tiles 102 are dedicated to storing program data, such as operation schedules, and / or output information from the vector processing unit 104. In some implementations, the computation arrays in one or more control tiles may not be used to perform computations. Alternatively, one or more control tiles may be configured to store program data, such as operation schedules, in addition to performing computations on input data, such as received weight inputs and activation values. In some implementations, weight inputs are loaded into the memory of each tile 102 where the weight inputs are used, without first storing the weight inputs in a subset of one or more control tiles.
[0085] In a second step, as shown in FIG. 8A, at scheduled counter values, the weight inputs 301 are loaded into individual cells 306 of the computation array in the tile 102. Loading the weight inputs 301 into the individual cells 306 may include transferring the weight inputs 301 from the memory of one or more control tiles to the corresponding tile 102 to which the weight inputs 301 belong. The weight inputs 301 may communicate along generic controllable bus lines to the tile 102 and transfer to the memory via multiplexers coupled to the bus lines and the memory. FIG. 8B is a detailed diagram of an example of a tile 102. The weight inputs 301 may be stored in the memory 302 for the duration of a model execution, which may include the calculation of multiple inferences. As an alternative to loading the weight inputs 301 from one or more control tiles, the weight inputs 301 may be preloaded directly into the memory of the tile 102 from the communication interface 108. To prepare the model for execution, the weight inputs 301 for each tile 102 may be loaded from the memory 302 of the tile 102 into each cell 306 of the computational array in that tile 102. For example, the weight inputs 301 may be loaded into registers 400 (also called "back registers") in the cells 306. The back registers allow the cells 306 to perform computations on their current weight inputs while the next weight input is loaded into the back register. Although loading the weight register is shown for only one cell 306 in FIG. 8B, the weight registers of other cells in the computational array may also be loaded during this step.
[0086] In a third step, as shown in FIG. 9A, at the scheduled counter value, the activation value 500 may be introduced to the tile 102 and stored in the memory 302. The activation value 500 may be transferred over multiple clock cycles. A calculation is then performed by the calculation array of each tile 102 using the received activation value 500 and the weight input 301 from the memory 302 in the tile 102. For example, the calculation may include multiplying the activation value by the weight input and then adding the result with a product of a different weight input and the activation value. In some implementations, the activation value 500 is communicated to the tiles 102 on the controllable bus lines 330 and communicated between the tiles 102. Each of the controllable bus lines 330 may extend along the same direction. For example, as shown in FIG. 9B, the controllable bus lines 330 extend laterally along a lattice dimension that is orthogonal to the lattice dimension along which the controllable bus lines 320 extend. Further, as indicated by the arrows 501 in Figure 9A and the arrows 501 on the controllable bus lines 330 in Figure 9B, the activation input data 500 travel in the same (e.g., east flow) direction on the bus lines 330. Alternatively, in some implementations, some activation input values 500 travel in a first direction (e.g., east flow direction) on some controllable bus lines 330 and some other activation input values 500 travel in a second opposite direction (e.g., west flow direction) on some other controllable bus lines 330.
[0087] In some implementations, the number of controllable bus lines 330 that extend through each tile 102 is determined by the size of the computational array. For example, the number of controllable bus lines 330 that extend through each tile 102 may be at least equal to the number of rows of cells in the computational array. In the example shown in FIG. 9B, given that there are eight rows of cells 306 in the computational array of the tile 102, there are eight controllable bus lines 330 that pass through the tile 102. In some implementations, each separate controllable bus line 330 transfers an activation input value 500 to a cell 306 in a corresponding row of the computational array. For example, in an 8×8 computational array of cells 306 in tile 102, a first controllable bus line 330 transfers activation input values 500 to cells 306 in the first row of the array, a second controllable bus line 330 transfers activation input values 500 to cells 306 in the second row of the array, and so on for other controllable bus lines 330, with the last controllable bus line 330 transferring activation input values 500 to cells 306 in the last row of the array. Additional controllable bus lines (e.g., partial sum bus lines) may pass through each tile to provide partial sums from another tile, receive results of computations within the tile to combine with the provided partial sums, and output new partial sums to a new tile or vector processing unit.
[0088] In some implementations, the controllable bus line 330 routes the activation input value 500 to a circuit configured to perform a calculation within the cell 306. For example, as shown in FIG. 9B, the controllable bus line 330 is configured to route the activation input value 500 to a multiplier circuit 502 within the cell 306. The activation input value 500 can be routed to the multiplier circuit 502 by using a multiplexer on the controllable bus line 330.
[0089] In some implementations, after the activation input value 500 and the weight input value 301 are determined to be in place (e.g., after a predetermined number of counter cycles required to perform a loading operation), the cell 306 of the computation array in the tile 102 performs a computation using the received activation input value 500 and the weight input value 301 from the memory 302 in the tile 102. For example, as shown in FIG. 9B, the weight input value 301 already stored in the register 400 is transferred to a register 504 (also called a "front register"). The weight input value 301 is then multiplied by the received activation input value 500 using a multiplier circuit 502.
[0090] As described herein, the activation input value 500 is communicated on the controllable bus line 330. In some implementations, the controllable bus line 330 is a general-purpose controllable bus line. In some implementations, the controllable bus line 330 can be used exclusively for providing activation inputs. For example, as shown in FIG. 9B, an activation input value can be provided to the tile 102 (e.g., to a cell 306 of a computational array within the tile 102) by the line 330, while other general-purpose controllable bus lines 310b can be used to provide other data and / or instructions to the tile 102.
[0091] In a fourth step, as shown in FIG. 10B, at the scheduled counter value, the result of the calculation of the weight input value 301 and the activation input value 500 in each cell 306 is passed to a circuit 602 in the cell 306 to generate an output value 600. In the example of FIG. 10B, the circuit 602 includes an adder circuit. The adder circuit 602 in each cell 306 is configured to add the product of the multiplier circuit 502 to another value obtained from another tile 102 in the ASIC 100 or another cell 306 in the computation array. The value obtained from the other tile 102 or another cell 306 can include, for example, an accumulation value. Thus, the output value 600 of the adder circuit 602 is a new accumulation value. The adder circuit 602 can then send the new accumulation value 600 to another cell located in a lower adjacent cell (e.g., in the south flow direction) of the computation array in the tile 102. The new accumulation value 600 can be used as an operand for the addition in the lower adjacent cell. At the inboard last row of cells in the computation array, the new accumulation value 600 may be forwarded to another tile 102 in the ASIC 100, as shown in Figure 10A. In another example, the new accumulation value 600 may be forwarded to another tile 102 that is at least one tile away from the tile 102 where the new accumulation 600 was generated. Alternatively, as also shown in Figure 10A, the new accumulation value 600 from the last row of cells in the computation array is forwarded to the vector processing unit 104.
[0092] The accumulated values 600 transferred into or out of the tile 102 may travel along the controllable bus lines 320. Each of the controllable bus lines 320 extends along the same direction. For example, as shown in FIG. 10B, the controllable bus lines 320 extend vertically along a grid dimension that is orthogonal to the grid dimension along which the controllable bus lines 330 extend. Furthermore, as indicated by the arrows 604 in FIG. 10A and 604 in FIG. 10B, the accumulated values 600 travel in either a north-flow direction or a south-flow direction on the controllable bus lines 320 depending on the location of the vector processing unit 104 relative to the tile 102 in which the accumulated values 600 were generated. For example, in a tile 102 located above the vector processing unit 104 in FIG. 10A , the accumulated value 600 moves on the controllable bus line 320 in a south flow direction toward the vector processing unit 104, while in a tile 102 located below the vector processing unit 104, the accumulated value 600 moves in a north flow direction toward the vector processing unit 104.
[0093] In a fifth step as shown in FIG. 11, the data received by the vector processing unit 104 at the scheduled counter value (e.g., accumulated value) is processed by the vector processing unit 104 to provide a processed value 900. The processing of the data in the vector processing unit 104 may include adding a bias to the data received at the vector processing unit 104, performing additional accumulation operations, and / or applying a non-linear function (e.g., a normalization function or a sigmoid function known in neural network systems) to the received data. Other operations may also be applied by the vector processing unit 104. The vector processing unit 104 may include circuitry arranged in multiple segments 106, each segment 106 configured to process data received from a corresponding column of the tile 102 and generate a corresponding processed value 900.
[0094] In a sixth step, as shown in FIG. 12A, at the scheduled counter value, the processed value 900 from the vector processing unit 104 is transferred and stored in one or more tiles of the ASIC 100, e.g., a subset of tiles of the ASIC 100. For example, the processed value 900 can be sent to the control tile 103, which is located immediately next to the vector processing unit 104. Alternatively, or in addition, the processed value 900 can be sent to one or more of the other tiles 102 in the ASIC 100. The processed value 900 can be transferred to one or more tiles via a general-purpose controllable bus line, such as the controllable bus line 310c. After the processed value 900 reaches the tile (e.g., the control tile or other tiles 102), it can be stored in the memory 202 of the tile. For example, the processed value 900 can be transferred to the memory 902 using a multiplexer associated with the controllable bus line 310c. The step of storing the processed values 900 can be performed after the inferences of each model layer are obtained. In some implementations, the processed values 900 can be provided as input values to the next layer of the model.
[0095] In a seventh step, the processed values 900 may be exported to the ASIC 100 at the scheduled counter values, as shown in Figures 13A and 13B. For example, the processed values 900 may be transferred from the memory 202 of one or more control tiles to the communication interface 108. The processed values 900 may be communicated to the communication interface 108 over a controllable bus line (e.g., controllable bus line 310c and / or 310d). The processed values 900 may be transferred to the controllable bus line via a multiplexer associated with the controllable bus line.
[0096] The processed values 900 may be exported to the ASIC 100, for example, when inferences have been obtained for a final layer of the model, or when the model is partitioned among multiple ASICs and inferences have been obtained for a final layer associated with the ASIC 100. The processed values 900 may be received by a SerDes interface of the communications interface 108 and exported to another destination, including, for example, but not limited to, another ASIC 100 or a field programmable gate array chip.
[0097] In the exemplary process described with respect to FIGS. 7A-13B, activation values and weight inputs may need to be fully propagated throughout the computational array of each tile before the cell computation is performed, or the cell may perform the computation before all values are fully propagated. In either case, the operation schedule of the individual tiles can be adjusted so that the computation occurs at the correct time. For example, if a particular machine learning program requires that activation values and weight inputs be fully propagated through the computational array of each tile before the cell computation is performed, the operation instructions can schedule the computation to occur at a time that ensures that the activation values and weights are fully propagated. Furthermore, while the ASIC 100 has been described as having weight inputs sent to columns of the computational array and activation inputs sent to rows of the computational array, in some implementations, weight inputs are sent to rows of the array and activation inputs are sent to columns of the array.
[0098] Additionally, although computational arrays have been described herein as using individual summing circuits in each cell, groups of cells in a computational array (e.g., all of the cells in a column) may be directly coupled to a single summing circuit, which sums the outputs received from the cells in the group, thus reducing the number of summing circuits required to store the outputs.
[0099] The subject matter and functional operations described herein may be implemented in digital electronic circuitry, computer hardware including structures disclosed herein and their structural equivalents, or a combination of one or more of them. The subject matter described herein may be implemented as one or more computer programs, i.e., as one or more modules of computer program instructions encoded on a tangible, non-transitory program carrier for execution by or for controlling the operation of a data processing apparatus. Alternatively or additionally, the program instructions may be encoded on a human-generated propagated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal generated to encode information to be transmitted to a suitable receiver device for execution by the data processing apparatus. The computer storage medium may be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or a combination of one or more of them.
[0100] The term "data processing apparatus" encompasses all kinds of apparatus, devices, and machines for processing data, including, by way of example, a programmable processor, a computer, or multiple processors or computers. The apparatus may include special purpose logic circuitry, e.g., a Field Programmable Gate Array (FPGA) or an ASIC. In addition to hardware, the apparatus may also include code that creates an execution environment for the computer program in question, e.g., code that constitutes a processor firmware, a protocol stack, a database management system, an operating system, or any combination of one or more of these.
[0101] The processes and logic flows described herein may be performed by one or more programmable computers executing one or more computer programs to perform functions by processing input data and generating output. The processes and logic flows may also be performed by, and an apparatus may be implemented as, special purpose logic circuitry, e.g., an FPGA, an ASIC, or a GPGPU (general purpose graphics processing unit).
[0102] Although many specific implementation details are described herein, these should not be interpreted as limitations on the scope of any invention or what may be claimed, but as descriptions of features that may be specific to certain embodiments of a particular invention. Some features described in the context of separate implementations herein can also be combined and implemented in a single embodiment. Conversely, various features described in the context of a single embodiment can also be implemented separately in multiple embodiments, or in any suitable subcombination. Furthermore, although each feature may be described above as working in some combinations and may be claimed as such from the outset, one or more features from a claimed combination may, in some cases, be deleted from the combination, and the claimed combination may be directed to a subcombination or a variation of the subcombination.
[0103] Similarly, although operations are shown in the figures in a particular order, this should not be understood as requiring such operations to be performed in the particular order or sequentially shown, or that all of the illustrated operations be performed, to achieve desirable results. In some situations, multitasking and parallel processing may be advantageous. Furthermore, the separation of various system modules and components in the above-described embodiments should not be understood as requiring such separation in all implementations, and it should be understood that the described program components and systems may generally be integrated into a single software product or packaged as multiple software products.
[0104] Specific embodiments of the subject matter have been described. Other implementations are within the scope of the following claims. For example, although bus lines have been described as "controllable," not all bus lines need have the same level of control. For example, there may be various degrees of controllability, and some bus lines may only be controlled if some bus lines are limited in terms of the number of tiles from which data can be obtained or to which data can be sent. In another example, some bus lines may be dedicated to providing data along a single direction, such as north, east, west, or south, as described herein. In some cases, the actions recited in the claims may be performed in a different order and still achieve desirable results. As an example, the processes depicted in the accompanying figures do not necessarily require the particular order or sequence depicted to achieve desired results. In some situations, multitasking and parallel processing may be advantageous. [Explanation of symbols]
[0105] 10 Chips 12 Communication Interface 14 Hardware Tile 16 Data bus, communication bus 17 Data Bus 18 Control circuit 20 Local Memory 22 Calculation unit, arithmetic unit 40 Trigger Table 50 entries 52 Trigger ID 54 Address 56 Configuration space update 58 Activation Flag 60 additional flags 100 ASIC 101 The First Dimension 102 tiles 103 The Second Dimension 104 Vector Processing Unit 106 Segments 108 Communication Interface 110, 110a, 110b, 110c, 110d part 200 tiles 202 Local Memory 204 Computational Array 206 Cells 210 Controllable Bus Line 210a, 210b, 210c, 210d bus lines 220 Bus Line 220a, 220b partial sum 221 Control circuit 300 Compass 301 Weight Input 302 Memory 306 Cell 310a, 310b, 310c, 310d Controllable bus lines 320, 330 Controllable bus lines 500 Activation Value 501 Arrow 502 Multiplier Circuit 504 Register 602 Circuit 604 Arrow 900 Value after processing 902 Memory 1010A, 1010B Interface
Claims
1. 1. A method of configuring a hardware tile, comprising: receiving, by the hardware tile, a first data packet, the first data packet including configuration data, the hardware tile being in a first state, the first state being a state in which the hardware tile listens for a configuration packet; Based on the configuration data of the first data packet, switching the hardware tile to a second state for a predetermined amount of time; exiting the second state upon expiration of the predetermined amount of time. A method for providing the above.
2. the second state includes a data transfer state; The method of claim 1 , wherein the hardware tile is configured to forward data received at the hardware tile while in the data forwarding state.
3. receiving, by the hardware tile, a second data packet while the hardware tile is in the data transfer state; transferring the second data packet from the hardware tile while the hardware tile is in the data transfer state; 3. The method of claim 2, comprising:
4. The method of claim 1 , wherein when the hardware tile is in the first state, the hardware tile is configured to read data packets received at the hardware tile.
5. The method of claim 4 , wherein switching the hardware tile to the second state comprises reading configuration data of the first data packet.
6. 6. The method of claim 5, wherein switching the hardware tile to the second state comprises installing and executing a state counter that specifies a predetermined length of time the hardware tile is in the second state.
7. 7. The method of claim 6, wherein the step of exiting the second state occurs when the state counter reaches the predetermined amount of time.
8. The method of claim 5 , wherein reading the configuration data of the first data packet comprises reading a header of the first data packet.
9. The method of claim 1 , wherein exiting the second state comprises initiating a trigger that causes the hardware tile to be reconfigured into a new state.
10. The method of claim 9 , wherein in the new state, the hardware tile is configured to store in a memory data packets received at the hardware tile.
11. The method of claim 9 , wherein in the new state, the hardware tile is configured to install data from memory.
12. The method of claim 9 , wherein in the new state, the hardware tile is configured to install data from a data packet received at the hardware tile.
13. The method of claim 1 , wherein exiting the second state comprises initiating a trigger that causes the hardware tile to be reconfigured back to the first state.
14. switching the hardware tile to the second state includes installing and executing a state counter that specifies the predetermined length of time that the hardware tile is in the second state; the state counter is synchronized with a global clock of the chip on which the hardware tile is mounted, or The method of claim 1 , wherein the state counter is synchronized with a local clock of the hardware tile.
15. and exiting the second state includes initiating a trigger that causes the hardware tile to be reconfigured to a new state; The method of claim 1 , wherein the reconfiguring to a new state includes installing and executing a state counter that specifies the predetermined length of time the hardware tile is in the new state.
16. A system comprising an integrated circuit including the hardware tile of the method according to any one of claims 1 to 15.
Citation Information
Patent Citations
Multiprocessor system
JP1987208158A
Information processing system
JP2004021867A
Multiprocessor system with improved secondary interconnection network
JP2018125044A
Configuring Coarse-Grained Reconfigurable Arrays (cgra) for Dataflow Instruction Block Execution in Block-Based Dataflow Instruction Set Architectures (isa)
JP2018527679A
Synchronization in multiple tile processing array
JP2019079529A