Byte stream processing pipeline for hardware integrated circuits
By introducing an improved instruction set with byte addressing capabilities, the complexity of programming and debugging in memory access of hardware integrated circuits is solved, enabling flexible memory access and efficient machine learning computation, especially in image processing operations, supporting fine-grained stride operations and data reuse.
Patent Information
- Application Number
- CN202480021120.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-03-24
- Filing Date
- 2024-03-25
- Publication Date
- 2025-11-07
AI Technical Summary
Existing hardware integrated circuits have data padding and access alignment requirements when accessing memory, making them difficult to program and debug. Furthermore, existing instruction sets struggle to support flexible memory access modes, especially in machine learning computations, particularly image processing operations, where they cannot effectively utilize data locality and perform fine-grained stride operations.
It adopts an improved instruction set architecture (ISA), introduces byte addressing capabilities, and allows local access from any given byte address of the input memory of the compute block through the ByteAddressingMode instruction. It supports automatic zero padding, simplifies memory access operations, and provides unified, semantically consistent fields to improve access efficiency.
It enables flexible access to memory, supports fine-grained stride operations and space-time data reuse, improves the efficiency of machine learning computations, and reduces programming and debugging complexity, especially in image processing and other data science applications.
Smart Images

Figure CN120917423A_ABST
Abstract
Description
BACKGROUND
[0001] This specification generally relates to memory operations for hardware integrated circuits.
[0002] Neural networks are machine learning models that employ one or more layers of nodes to generate, for received inputs, an output such as a classification. Some neural networks include one or more hidden layers in addition to an output layer. Some neural networks can be convolutional neural networks (CNNs) configured for image processing or recurrent neural networks (RNNs) configured for speech and language processing. Different types of neural network architectures can be used to perform various tasks related to classification or pattern recognition, prediction involving data modeling, and information clustering.
[0003] Neural network layers can have a corresponding set of parameters or weights. The weights are used to compute neural network inference by processing inputs (e.g., a batch of inputs) through a neural network layer to generate a corresponding output for the layer. A batch of inputs and a kernel set can be represented as tensors, i.e., multi-dimensional arrays, of inputs and weights, respectively. Hardware accelerators are application-specific integrated circuits used to implement neural networks. The circuits include memory whose locations correspond to elements of tensors that can be traversed or accessed using control logic of the circuits. SUMMARY
[0004] This document describes aspects of an enhanced instruction set that can be executed by a processor or hardware accelerator to provide an efficient byte-level data streaming process. The disclosed technology and instruction set also allow for corresponding synchronization of those streaming processes for various compute-intensive operations, including various machine learning (ML) workloads for inference or prediction operations.
[0005] One aspect of the subject matter described in this specification can be embodied in a computer-implemented method for accessing memory banks of a hardware accelerator. The method includes receiving an instruction for a compute tile of the hardware accelerator. The instruction can be executed at the compute tile to cause performance of operations including: i) identifying, in the instruction, an opcode for an operation involving access to a first memory of one or more banks of the compute tile; ii) generating, based on the opcode, a byte addressing sequence for accessing consecutive bytes of different lengths stored at the one or more banks; iii) performing, based on the byte addressing sequence, a plurality of byte-addressable memory accesses at the first memory to obtain a plurality of input vectors from the one or more banks of the first memory; and iv) processing, based on the instruction, each of the plurality of input vectors at the hardware accelerator.
[0006] These and other implementations can each optionally include one or more of the following features. For example, in some implementations, the byte addressing sequence is operable to access consecutive bytes of different lengths starting from any given byte address across the one or more banks of the first memory. The method includes generating a byte access request based on the opcode, the byte access request providing byte-granularity access for obtaining data of at least one byte stored at a given row of a bank among the one or more banks of the first memory.
[0007] In some implementations, generating the byte addressing sequence includes generating the byte addressing sequence corresponding to a plurality of byte access requests, each byte access request for accessing a distinct / different portion of byte-level data stored at a respective address location across the one or more banks of the first memory. In some cases, each byte access request can be for accessing a respective distinct portion of byte-level data stored at the respective address location. The opcode can be for execution at the compute tile to iterate over respective elements of the input tensor.
[0008] In some implementations, the method further includes: i) executing the plurality of byte access requests using the byte addressing sequence; ii) performing a plurality of byte-granularity accesses at the first memory based on the executed plurality of byte access requests; and iii) iterating over respective elements of the input tensor by performing the plurality of byte-granularity accesses at the first memory. In some implementations, each distinct portion of byte-level data stored at a respective address location across the one or more banks of the first memory corresponds to a respective element of the input tensor; and each respective address location storing a distinct portion of byte-level data is mapped to a corresponding element of the input tensor.
[0009] In some implementations, a portion of the byte-level data corresponds to an input vector of the plurality of input vectors; and the plurality of byte-addressable memory accesses for obtaining the plurality of input vectors corresponds to at least one of: i) the plurality of byte-granularity accesses and ii) the plurality of byte access requests. In some aspects, processing the plurality of input vectors includes: processing the plurality of input vectors by a neural network layer of a neural network implemented at the hardware accelerator; and generating an output of the neural network layer in response to processing the plurality of input vectors by the neural network layer.
[0010] Other implementations of this and other aspects include corresponding systems, apparatus, and computer programs to perform the actions of the methods, encoded on computer storage devices. One or more computer systems can be so configured by virtue of software, firmware, hardware, or a combination thereof installed on the system that in operation cause the system to perform the actions. One or more computer programs can be so configured by virtue of having instructions that, when executed by data processing apparatus, cause the apparatus to perform the actions.
[0011] The subject matter described in this specification can be implemented in particular embodiments so as to realize one or more of the following advantages.
[0012] Unlike previous approaches to data access that are limited to row addressing schemes for accessing physical memory, the disclosed technology provides an instruction set architecture (ISA) that allows for enhanced byte-level addressing functionality for memory, such as on-chip static random access memory (SRAM) of an example application-specific integrated circuit, e.g., a hardware accelerator, a graphics processing unit (GPU), or a neural network (NN) processor, including a tensor processing unit (TPU) or a vector processing unit (VPU). The application-specific integrated circuit can also be a hardware accelerator or a neural network processor that uses a particular type of memory structure to store inputs for ML workloads.
[0013] Existing ISAs have data padding and access alignment requirements and can be difficult to program and debug. In contrast, the byte addressing provided by the improved instruction set of the present document enables local access of contiguous bytes of various lengths from any given byte address of an input memory of a compute tile. The technology and systems disclosed in this specification can provide new combinations of byte access patterns to support ML computations. Such data access patterns allow for various types of fine-grained stride operations for ML image processing operations as well as spatial and temporal data reuse to support ML inference and training workloads for a range of data science applications.
[0014] The improved ISA includes one or more “ByteAddressingMode” instructions with integrated fields that allow for automatic zero padding when accessing data from an input memory of a compute tile. For example, based on the disclosed instruction set, a processor of a compute tile can issue a byte access request that provides byte-granularity access for obtaining data of at least one byte stored at a given row of a physical memory bank among one or more banks of a first memory. Thus, a single byte can be retrieved without requiring unrelated data padding (such as with another byte padding request) and without requiring access alignment across a particular set of rows or banks.
[0015] The details of one or more implementations of the subject matter described in this specification are set forth in the accompanying drawings and the description below. Other potential features, aspects, and advantages of the subject matter will become apparent from the description, the drawings, and the claims. BRIEF DESCRIPTION OF DRAWINGS
[0016] Figure 1 is a block diagram of an example computing system for implementing a neural network machine learning model.
[0017] Figure 2 An example processing pipeline for routing inputs and outputs between a memory and compute cells of a hardware integrated circuit is shown.
[0018] Figure 3 Example fields and message structures of an example instruction set are shown.
[0019] Figure 4 Code for an example byte addressing operation for an instruction set is shown.
[0020] Figure 5 Aspects of an example byte addressing operation are shown.
[0021] Figure 6 An example process for byte stream processing using an example byte addressing operation of an instruction set is shown.
[0022] Figure 7 Examples of input tensors, parameter tensors, and output tensors are shown.
[0023] Like reference numbers and designations in various figures indicate like elements. DETAILED DESCRIPTION
[0024] Figure 1 is a block diagram of an example computing system 100 for implementing a neural network model at a hardware integrated circuit such as a machine learning hardware accelerator. The computing system 100 includes one or more compute tiles 101, a host 120, and a higher-level controller 125 (“controller 125”). As described in more detail below, the host 120 and the controller 125 cooperate to provide data sets and instructions to the one or more compute tiles 101 of the system 100.
[0025] In some implementations, the host 120 and the controller 125 are the same device. The host 120 and the controller 125 can also perform distinct functions but are integrated in a single device package. For example, the host 120 and the controller 125 can form a central processing unit (CPU) that interacts with or cooperates with a hardware accelerator having multiple compute tiles 101. In some implementations, the host 120, the controller 125, and the multiple compute tiles 101 are included or formed as different sections on a single integrated circuit die. For example, the host 120, the controller 125, and the multiple compute tiles 101 can form a specialized system on a chip (SoC) that is optimized for executing neural network models for processing machine learning workloads.
[0026] Each compute tile 101 typically includes a controller 103 that provides one or more control signals 105 to cause inputs (or activations) of an input vector 102 to be stored at or accessed from memory locations of a first memory 108 (“memory 108”). Likewise, the controller 103 can also provide one or more control signals 105 to cause weights (or parameters) of a matrix structure of weights 104 to be stored at or accessed from memory locations of a second memory 110 (“memory 110”). In some examples, the memory 108 is identified as “narrow” and the memory 110 is identified as “wide.” In some implementations, the input vector 102 is obtained from an input tensor and the matrix structure of weights is obtained from a parameter tensor. Each of the input tensor and the parameter tensor can be a multi-dimensional data structure, such as a multi-dimensional matrix or tensor. This is described in more detail below with reference to FIG. 2. Figure 7 More detail is described.
[0027] Each memory location of the memories 108, 110 can be identified by a corresponding memory address, such as a logical address having a corresponding mapping to a physical row of a physical memory bank of the memory. The compute tile 101 can derive a set of contiguous addresses (e.g., virtual / logical addresses) from a request group. For example, the set of contiguous addresses can be derived with reference to a logical memory corresponding to the physical memory 108 of the compute tile 101.
[0028] The logical memory can have multiple logical ports, and each port can be connected to or associated with a different requestor requesting access to the physical resources of the memory 108. To handle a given access request, for each port, the tile 101 (or its controller 103) determines, based on an address in the request, a bank to which the request is to be routed.
[0029] For each bank, the compute tile 101 can include an arbiter that arbitrates access to the bank from multiple ports, e.g., according to a bank generation function that is uniquely configured to mitigate some (or all) of the requests being routed to the same physical memory bank. The logical memory, the ports of the logical memory, and the arbiter can be implemented in software, hardware, or both. In some implementations, the logical memory and its ports, and the arbiter are controlled based on control signals generated by the controller 103.
[0030] Each of the memories 108, 110 can be implemented as a series of physical banks, units, or any other related storage medium or device. Each of the memories 108, 110 can include one or more registers, buffers, or both. In some implementations, the memory 108 is an input / activation memory, while the memory 110 is a parameter memory. In some other implementations, inputs or activations are stored at the memory 108, the memory 110, or both; and weights are stored at the memory 110, the memory 108, or both. For example, inputs and weights can be transferred between the memory 108 and the memory 110 to facilitate certain neural network computations. In some implementations, each of the memories 108 and 110 is referred to as a bank memory.
[0031] Each compute bank 101 also includes an input activation bus 106, an output activation bus 107, and a compute unit 112 that includes one or more hardware multiply-accumulate circuits (MACs) in each of the compute cells 114 a / b / c. The controller 103 can generate control signals 105 to obtain operands stored at the memories of the compute bank 101. For example, the controller 103 can generate control signals 105 to obtain: i) an example input vector 102 stored at the memory 108 and ii) weights 104 stored at the memory 110. Each input obtained from the memory 108 is provided to the input activation bus 106 for routing (e.g., direct routing) to the compute cells 114 a / b / c in the compute unit 112. Similarly, each weight obtained from the memory 110 is routed to the compute cells 114 a / b / c of the compute unit 112.
[0032] As described below, each compute cell 114 a / b / c performs computations that produce partial sums or accumulated values to generate an output for a given neural network layer. An activation function can be applied to the output set to generate an output activation set for the neural network layer. In some implementations, the output or output activations are routed via the output activation bus 107 for storage and / or transfer. For example, an output activation set can be transferred from a first compute bank 101 to a different second compute bank 101 for processing at the second compute bank 101 as input activations for a different layer of the neural network.
[0033] Generally, each compute tile 101 and system 100 can include additional hardware structures to perform computations associated with multi-dimensional data structures such as tensors, matrices, and / or data arrays. In some implementations, the inputs of an input vector (or tensor) 102 and the weights 104 of a parameter tensor can be pre-loaded into the memory 108, 110 of a compute tile 101. The inputs and weights are received as sets of data values that arrive at a particular compute tile 101 from a host 120 (e.g., an external host) via a host interface or from a higher level of control such as a controller 125.
[0034] Each of the compute tiles 101 and controllers 103 can include one or more processors, processing devices, and various types of memory. In some implementations, the processors of the compute tiles 101 and controllers 103 include one or more devices such as a microprocessor or central processing unit (CPU), a graphics processing unit (GPU), an application-specific integrated circuit (ASIC), or a combination of different processors. Each of the compute tiles 101 and controllers 103 can also include other computing and storage resources such as buffers, registers, control circuitry, etc. These resources cooperate to provide additional processing options for performing one or more determinations and computations described in this specification.
[0035] In some implementations, the processing units of the controllers 103 execute programmed instructions stored in memory to cause the controllers 103 and compute tiles 101 to perform one or more functions described in this specification. The memory of the controllers 103 can include one or more non-transitory machine-readable storage media. The non-transitory machine-readable storage media can include a solid-state memory, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (e.g., EPROM, EEPROM, or flash memory), or any other tangible medium that is capable of storing information or instructions.
[0036] The system 100 receives instructions that define a particular computational operation to be performed by the compute tiles 101. The system 100 also receives data such as inputs, activations, weights (or parameters) as operands for the computational operation. The system 100 can include one or more buses 109 for routing the instructions and data. For example, the system 100 can include a first bus 109-1 that provides instructions, including relevant commands, operation codes, and operation parameters (not weights), to each of the one or more compute tiles 101 of the system 100. The system 100 can also include a second bus 109-2 that provides data or operands to each of the one or more compute tiles 101 of the system 100.
[0037] Bus 109 can be configured to provide one or more interconnected data communication paths between two or more compute tiles of system 100. For example, first bus 109-1 can be a ring bus that traverses each compute tile to transmit data sets and individual instructions (or multiple instructions) to one or more compute tiles to perform an ML workload, while second bus 109-2 can be a mesh bus that interconnects two or more tiles to provide data or operand sets between two or more compute tiles 101.
[0038] In some implementations, bus 109-1 and bus 109-2 are distinct data buses. In some other implementations, bus 109-1 and bus 109-2 are the same data bus. In some cases, to process an example workload, a first compute tile 101 arbitrates and performs access requests (e.g., read / write requests) to memory 108, where the requests are based on external data communications originating from outside the first compute tile 101, such as from a second, different compute tile 101. Such external communications can be received at the first tile 101 via bus 109-2 (e.g., a mesh bus).
[0039] Each compute tile 101 is an individual computing unit that cooperates with other tiles 101 in system 100 to accelerate computation across one or more layers of a multi-layer neural network or across one or more sections of another machine learning construct. Each compute tile 101 can function as an individual computing unit. For example, each compute tile 101 is an independent computing component configured to independently perform a subset of tensor or ML computations relative to other compute tiles. In some implementations, compute tiles 101 can also share execution of tensor computations associated with a given instruction.
[0040] In some implementations, a host can generate a set of parameters (i.e., weights) and corresponding inputs for processing at a neural network layer. For example, a host can generate a set of compressed parameters (CSP) and a corresponding mapping vector, such as a non-zero map (NZM), that maps neural network inputs to non-zero parameters in the set of CSPs for a given operation. For example, host 120 can send parameters to a compute tile 101 via a host interface for further processing at the tile. Controller 103 can execute programmed instructions to analyze a data stream associated with received weights and inputs that includes compressed parameters and corresponding mapping vectors.
[0041] The controller 103 causes the input and weights of the data stream to be stored at the compute tile 101. For example, the controller 103 can store the inputs (including any associated mapping vectors) and weights / parameters (including any associated compressed sparse parameters) in the local tile memory of the compute tile 101. This is described in more detail below. The controller 103 can also analyze the input data stream to detect operation codes (“opcodes”). The system 100 can support various types of opcodes, such as opcodes that indicate a vector matrix multiplication operation, an element-wise vector operation, and an opcode type that indicates whether compressed sparse parameters / weights or uncompressed parameters / weights are to be used for a given operation.
[0042] Based on one or more opcodes, the controller 103 can activate or execute a bank generation function to arbitrate requests for access to the tile memory of the compute tile 101. For example, the controller 103 utilizes the bank generation function to arbitrate two or more requests such that each of the two or more requests is processed against a different physical bank of the tile memory. The controller 103 can employ a predetermined bank selection technique that is programmed or encoded at the controller 103 prior to an inference determination being performed at the compute tile 101.
[0043] In some implementations, a given compute operation involves multiple requesters, each of which needs to access resources of the memory 108. For example, a compute workload performed at the compute tile 101 can trigger memory access requests due to a tensor traversal operation that requires read-write access to respective address locations of the memory 108. The system 100 can use the enhanced byte-level addressing functionality of the disclosed ISA to process these requests to access one or more bytes of data stored at address locations of the memory 108, 110. As described below, these address locations can correspond to elements of an input (or parameter / weight) tensor that is processed as part of the workload.
[0044] In addition to tensor read / write operations, processing workloads can involve processing read or write access requests to: i) move data (e.g., parameters) from memory 108 (e.g., first on-chip SRAM) to memory 110 (e.g., second on-chip SRAM), and ii) move data from memory 110 to memory 108. Tensor operations can be indicated by an opcode in an instruction (e.g., a single instruction) received at compute tile 101. For example, a “ByteAddressingMode” instruction can include one or more opcodes for tensor operations to be executed at compute tile 101 to iterate over respective elements of an input tensor. Generally, tensor operations can be executed to perform a particular ML computation operation involving: i) consuming (or reading) tensors from memory 108, 110, and ii) producing (or writing) tensors to memory 108, 110.
[0045] As used in this document, the first on-chip SRAM can refer to one or more memory units that each operate on and store data having a size (or width) equal to or less than, for example, 16 bits, 32 bits, or 64 bits, while the second on-chip SRAM can refer to one or more memory units that each operate on and store data having a size (or width) greater than 64 bits or the size of data operated on by the first SRAM units. For example, the size or width for on-chip SRAM 108 can be between 8 bits and 64 bits, while the size or width for on-chip SRAM 110 can be between 64 bits and 256 bits. Of course, other sizes and data widths beyond these example ranges are also contemplated for each of memory 108, memory 110.
[0046] Figure 2 An example processing pipeline for routing inputs and outputs between memory and compute cells of a hardware integrated circuit is shown. Generally, pipeline 200 uses input bus 106 to route inputs obtained from memory locations of memory 108 to one or more compute cells 114. In some implementations, the inputs routed to compute cells 114 via input bus 106 are inputs obtained from memory 108 using various combinations of byte access requests based on the enhanced byte-level addressing functionality of the ISA disclosed in this specification.
[0047] The pipeline 200 routes the outputs to memory locations of the memory 108 using the output bus 107. These outputs can be numerical outputs or quantitative outputs generated from computational operations (e.g., multiplication or other arithmetic) performed at one or more computational cells 114. Each tile 101 can include a nonlinear unit (NLU) 210 that applies a nonlinear activation function to the outputs that are the result of the arithmetic performed at the computational unit 112. For example, the computational tile 101 causes the NLU 210 to apply an activation function (e.g., sigmoid, tanh) to generate an output activation for the computation at a given neural network layer. The pipeline 200 routes the outputs (or output activations) to memory locations of the memory 108 using the output bus 107.
[0048] The pipeline 200 utilizes a hardware architecture in which the input bus 106 is coupled (e.g., directly coupled) to each of the plurality of hardware computational cell groupings of the application-specific integrated circuit. Although not shown in the example of FIG. 1, the system 100 can provide a first operand from a location in the memory 108 and a second operand from a location in the memory 110, as described above with reference to the example of FIG. 2. Figure 2 Figure 1 The operands 204 are used for the computation performed at the cells 114, and the cells 114 generate outputs corresponding to the results of the computation. In some implementations, the computation is for a machine learning operation, such as for processing an input by a neural network layer of an artificial neural network.
[0049] In this implementation, the computational tile 101 can provide a first operand to the subset of cells 114 corresponding to an input or activation (e.g., a0, a1, a2, etc.) of an input feature map. For example, after a byte-level data access operation performed with respect to the memory bank of the memory 108, a respective input of the input vector 102 is provided to each MAC in the subset via the input bus 106 of the computational tile 101. The system 100 can provide different respective sets of inputs based on the byte-level access operations performed locally at each computational tile.
[0050] The system 100 can process, execute, or otherwise perform different combinations of byte access requests to obtain or write sets of operands for a given computational operation. For example, across multiple computational tiles 101, the system 100 can perform a broadcast operation to route a set of inputs (or operands) to the cells 114 of a given tile. The system 100 can use respective input groupings and corresponding weights at each computational tile 101 to compute products for a given neural network layer. At a given computational tile 101, the products are computed at each MAC in the subset by multiplying a respective input (e.g., a1) with a corresponding weight (e.g., w1) using the multiplication circuitry of the MAC.
[0051] The system 100 can generate an output of a layer based on an accumulation of a plurality of respective products computed at each MAC of the cells 114a / b / c in a subset of cells 114a / b / c of the compute unit 112. As explained below at least with reference to Figure 7 The multiplication operations performed within the compute tile 101 can involve, as explained below at least with reference to
[0052] In the example of Figure 2 The shift register 202 can provide shift functionality such that the input of the operand 204 is broadcast onto the input bus 106 and routed to the MAC(s) of the cell 114. In some implementations, the shift register 202 enables one or more input broadcast modes at the compute tile 101. For example, the shift register 202 can be used to broadcast the input sequentially, e.g., one-by-one, from the memory 108 (e.g., a first broadcast mode), broadcast the input concurrently, e.g., in parallel, from the memory 108 (e.g., a second broadcast mode), or some combination of these broadcast modes. The shift register 202 can be a consolidated function of the memory 108 and can be implemented in hardware, software, or both.
[0053] In some implementations, the weight (w3) of the operand 204 can have a zero weight value. When the controller 103 determines that the weight (w3) has a zero value, to conserve processing resources, the multiplication between the input (a2) and the weight (w3) can be skipped such that these operands are not routed to or consumed by the cell 114a / b / c. As noted above, the determination to skip this particular multiplication operation can be based on the mapping vector that maps the discrete inputs (an) of the input vector to the individual weights (wn) of the parameter tensor.
[0054] Figure 3 Example fields and message structures of example instruction sets are shown. A first set of fields and message structures are associated with a first instruction set 302, while a second, different set of fields and message structures are associated with a second, different instruction set 304. The instruction set 302 represents an existing method for data access that is limited to a row addressing scheme for accessing physical memory. For example, the instruction set 302 can represent an existing instruction set that is flawed at least because it includes operations that require data padding and access alignment requirements that are difficult to program and debug. This will be described below with reference to the example of Figure 4
[0055] In the example of Figure 3 In the example, instruction set 304 represents an ISA that allows enhanced byte-level addressing capabilities to be performed for the memory (e.g., on-chip SRAM) of the example-specific processor. For example, the processor could be a TPU, hardware accelerator, or neural network processor that uses a specific type of memory structure to store inputs for computational operations such as ML workloads. As indicated above, the processor may include multiple computation blocks 101, each of which may be configured to perform some (or all) portions of a computational operation (such as processing a batch of inputs through layers of an artificial neural network).
[0056] The inputs in this batch may correspond to an image, and the pixels of this image may be represented as an input tensor (e.g., a multidimensional input tensor). Each pixel in the input tensor may correspond to a neural network input provided as input to a layer of a neural network. The neural network input may also be an activation value corresponding to the neural network output generated as the output of a neural network layer. Each input (or activation) value of the tensor corresponds to an element along a specific dimension (e.g., X, Y dimension) of the multidimensional tensor. Accordingly, each element of the tensor may correspond to a memory location of memory 108 (or memory 110), and each memory location may have a corresponding address.
[0057] In some implementations, each computation block 101 includes a corresponding tensor traversal unit (TTU) for generating addresses for traversing elements along a given dimension of the tensor. See below for reference. Figure 5 Describe an example TTU. Each input can be represented by a set of bits (e.g., bytes) stored at a given memory address location in memory 108. Using the disclosed instruction set 304, system 100 can cause one or more TTUs in the TTU to generate addresses at the byte-level granularity. This is described in more detail below.
[0058] Generally, ISA 302 and 304 may include data fields and corresponding messages associated with those data fields. For example, data fields may specify the attributes or details of an operation, while messages may indicate the type of operation. Figure 3 In the example, for a given operation, ISA 302 requires multiple fields 306 and message 308. For instance, memory access using ISA 302 opcodes typically requires multiple data fields, such as row information fields or hardware component fields, which generally translates to increased processing overhead at the compute block level.
[0059] In contrast, ISA 304 represents an improved instruction set that provides a streamlined approach for specifying attribute and type information for various memory access operations. For example, the integrated aspects of ISA 304 provide a simplified and modular ISA that provides uniform, semantically consistent fields to allow for different memory access patterns. Rather than relying on multiple information fields, ISA 304 incorporates a uniform information field 310 and a uniform message field 312, which improves overall performance by avoiding the need to store and process values for unrelated fields and reducing the data communications needed to transmit those values.
[0060] In some implementations, the opcodes generated by ISA 304 have an information field 310 that can include a byte count parameter that indicates the number of bytes to access from a memory location of the compute tile 101. In some cases, the information field 310 is included in the opcode to indicate that the byte address mode for the compute tile has been activated. For example, the opcode in an instruction sent to the compute tile can convey a “ByteAddressMode” message that provides integrated fields for byte-level access to data in the memory 108, 110. This is described in more detail below with reference to Figure 4
[0061] For ISA 304, the uniform information field 310 and message field 312 cooperate to provide byte-granular addressing that enables access to contiguous bytes of various lengths from any given byte address at the memory bank of the compute tile 101. In some implementations, the compute tile 101 can include one or more TTUs, and ISA 304 is configured such that each TTU can have a matching ByteAddressMode field or instruction. The message field 312 represents a streamlined approach for indicating the type of pipeline (or operation type) in which byte-level access can occur. For example, the message field 312 can indicate tensor traversal for a tensor read / write operation.
[0062] In some implementations, the message field 312 can indicate byte-level access to an operation such as a read or write access request to: i) move data (e.g., parameters) from a first on-chip SRAM 108 to a second on-chip SRAM 110, and ii) move data from the memory 110 to the memory 108. One or more of these move operations can be used to support parameter caching. In some cases, to handle an example workload, a first compute tile 101 arbitrates and performs an access request (e.g., read / write request) to the memory 108, where the request is based on an external data communication originating from outside the first compute tile 101.
[0063] Figure 4 Example code is shown for a tensor operation 400 involving performing an input memory read operation (402) using the byte addressing features of the improved instruction set disclosed herein. In some implementations, the code for operation 400 corresponds to instructions for a memory addressing technique that is executed to access data from a portion of on-chip SRAM that stores inputs and activations for later processing by a neural network. The instructions (and / or code) for operation 400 can be derived from the ISA 304 for an example TPU that includes one or more compute tiles 101.
[0064] As noted above, existing ISAs (e.g., ISA 302) have data padding and access alignment requirements and are difficult to program and debug. The alignment aspect involves a requirement that, for a given access request, the number of bytes of data accessed must be a multiple of some particular integer value, where the integer value is greater than 1 (e.g., 4). The data padding aspect corresponds to the alignment aspect and involves a requirement to pad a certain number of bytes onto one or more bytes returned for a given request. The number of bytes padded onto the returned bytes is based on the integer multiple.
[0065] For example, existing / previous instruction sets can be tightly coupled to the microarchitecture of a circuit such that a memory unit of this circuit is required to return a minimum of four bytes even if a particular request only needs one, two, or three bytes of data. Thus, if a given request wants to access one byte and the previous instruction set imposes a requirement of four bytes (i.e., integer value = 4), the padding and alignment constraints of this previous ISA will require three bytes of padding. For this operation, the example compute tile 101 will be required to return four bytes of data, but three padding bytes will then be discarded. In some implementations, this inefficient padding and alignment requirement is associated with a row addressing scheme that imposes unnecessary constraints on read accesses to the memory 108, 110.
[0066] In contrast to existing / previous instruction sets, the techniques and systems disclosed in this specification provide new combinations of byte access patterns to support ML computations. More specifically, the disclosed techniques support more flexible memory access patterns and do not limit the opportunity for different access patterns in new domains such as image processing. In some implementations, the improved instruction set and byte access techniques (e.g., ISA 304) are used to optimize stencil operations involving data reuse patterns. For certain image processing applications, such operations are strong candidates for exploiting data locality.
[0067] For example, referring again to the tensor operation 400, data access patterns allow for various types of fine-grained stride (404) operations for workloads such as ML image processing operations, as well as spatial and temporal data reuse (404) to support ML inference and training workloads for a range of data science applications. In Figure 4 In examples, aspects of stride operations can be controlled via a “step” parameter 406. Generally, strides are handled in conjunction with memory accesses to obtain data such as image pixel values for image processing ML workloads.
[0068] Referring to the hardware configuration of the physical rows and banks of memory, images or pixel values associated with images can be stored across the physical rows and banks of memory 108, 110. In this example, the stride is a component or parameter of a filter (or kernel) of a neural network. The stride is used to modify the amount that the filter is moved over an image or video. For example, when the stride is set to 1, the compute tile 101 moves the filter over the region one pixel (or input) at a time. Likewise, when the stride is 2, the compute tile 101 moves the filter over the region two pixels at a time. Thus, the filter can be shifted based on the stride value of the layer, and in some implementations, the system 100 can repeat this process across multiple compute tiles 101 of different layers until the inputs for different regions of the image have corresponding dot products.
[0069] Moving the filter over the inputs of a region of an image based on a stride (or “step”) value can include retrieving, obtaining, or otherwise accessing the inputs from various locations in memory 108, 110 according to the stride value. In Figure 4 In examples, the stride can be represented as a “step size” used when sliding a convolutional filter over an input image (e.g., step = 2). By enabling local block access to consecutive bytes of various lengths, the improved byte addressing ISA of the present specification (e.g., ISA 304) supports and allows for a range of stride (or step values) to be applied from any given byte address of input memory at a compute tile.
[0070] As described above, the improved ISA includes one or more ByteAddressingMode instructions that have an integrated field related to memory access granularity and allow for automatic zero padding when accessing data from the input memory of a compute tile. For example, the instruction byte_address_mode (408) can include one or more "access_bytes" fields to indicate the number of bytes to write in the address space of memory 108, 110 at a given iteration. In some implementations, the bytes can be written to a contiguous address space (e.g., linear access) or a non-contiguous address space. Automatic zero padding can be implemented via hardware logic, software logic, or both on the fly.
[0071] With respect to the access_bytes field, the instruction can include one or more fields such as access_bytes_loop_map, last_access_bytes, and default_access_bytes. The effective byte size for a given data packet can be indicated by a valid_byte_size field or parameter in the particular byte_address_mode instruction or message. In some implementations, the value of the access_bytes field in the instruction for byte-level access can be greater than the maximum effective byte size per packet if the successive packet delivery is to store data in a contiguous address space. Thus, access_bytes can be greater than the amount of data carried by each ring / grid packet.
[0072] In these examples, the hardware compute tile can consume these ring packets and write the packets contiguously to the specified memory address. The packets can be consumed concurrently with determining how much or which portion(s) of a particular packet are valid for consumption by a particular hardware unit. For example, there can be 20 incoming ring packets. Each packet carries 32B of data, of which only 25B of the 32B is valid. In this example, the access_bytes field / instruction can be 500B / byte for the ring bus consumer pipeline. Thus, the ring consumer will consume all 20 ring packets, extract 25B from each ring packet, and a total of 500 of these bytes will be written in a contiguous manner at the start address specified by the TTU.
[0073] Using the access_bytes_loop_map, the system 100 decides whether to access the default_access_bytes or the last_access_bytes. For each TTU iteration, the system reads these access bytes (e.g., default or last) according to the particular access_bytes_loop_map. For each TTU clock cycle, a new start address can be generated to retrieve one or more access_bytes. In some implementations, the valid_byte_size is specific to a ring bus or mesh bus data packet. Each packet can include, for example, 32B of data, and a count value can be provided with the valid_byte_size field / instruction to specify the portion of the 32B data packet that is valid for consumption by a particular tile.
[0074] In the byte addressing mode, one or more tiles of the system 100 can dynamically create write transactions for byte-level access based on the valid_byte_size of a given packet and the outstanding access_bytes. In some implementations, control logic at a given tile 101 is configured to determine whether a single packet needs to be divided into multiple write transactions, and in response to determining that a packet needs to be divided, generate control signals to support or perform the division operation and generate write transactions.
[0075] For example, the system 100 utilizes this division feature to implement support for immediate packet dispersion, which allows for more efficient use of the available bandwidth of the bus 109. In some implementations, a write transaction can include multiple byte-level requests. In some cases, the byte-level requests correspond to a byte addressing sequence. Each address of the sequence identifies a memory location for storing (and later accessing) a byte of data, such as an input to be made for a machine learning workload.
[0076] As described above, the bus 109-1 can be a ring bus that traverses two or more compute tiles to transfer data sets, such as inputs and weights, to the compute tiles for execution of an ML workload. These data sets can alternatively be referred to as ring packets. For example, a data set can be an inbound ring packet that arrives at a compute tile 101 via the ring bus 109-1 of the system 100. The input of the ring packet received at the compute tile 101 is stored in the memory 108, which is later accessed and routed to the compute unit 112 for the computations needed to perform the ML workload.
[0077] The improved ISA and its corresponding byte_address_mode instruction allow for managing input data in a ring packet at byte-level granularity. The byte-level scheme implemented by the improved ISA provides a mechanism for storing cached data in coarse-grained structures such as ring packets or partial sums using fewer clock cycles and overall bandwidth compared to a row addressing scheme in which data is written over several cycles due to padding and alignment requirements.
[0078] For example, during a cross feed operation, a tile can receive a packet of 32 bytes (32B) of data from bus 109-1 (e.g., a ring bus). The structure of memory 108 can be limited to 4B per row, thus requiring an alignment across eight rows to receive 32 bytes of data. But in some cases, the example address of memory 108 to which the data will be written (e.g., 0x0...4) will not be aligned to eight rows. With a previous instruction set, such as ISA 302 or related specification method, writing data to this unaligned address would require the system to first write the 32 bytes to an eight row-aligned address using one or more cycles. The system then moves the data back to the target address (i.e., 0x0...4) by using an example SRAM1 to SRAM1 operation in which the data is moved within (or locally) memory 108. With the previous instruction set, the bandwidth of the SRAM1 to SRAM1 pipeline at tile 101 can be 4 bytes per cycle (4B). Thus, for the previous instruction set, the end-to-end cycle time for a cross feed operation (ring to SRAM1 write) would require at least 9 cycles, which is unnecessarily high.
[0079] The improved instruction set of ISA 304 allows for SRAM1 (or SRAM2) memory writes to any byte address (e.g., at memory 108, 110) without the extraneous overhead of additional cycles and bandwidth consumption. System 100 can perform certain write operations in fewer clock cycles and without the row alignment constraints or additional operations required by the previous instruction set. System 100 can use the configuration allowed by ISA 304 to perform a series of write operations in a single cycle regardless of how the memory banks of memory 108, 110 are organized at the tile. For example, when the requested bundle size is 8 (e.g., 32 bytes), the set of addresses generated by the TTU at the tile need not be eight row-aligned (e.g., 0, 8, 16... row addresses) to complete the write of the bundle in a single cycle.
[0080] Additionally, rather than discarding packets, the improved ISA's byte addressing mode supports efficient byte-granular filtering of data packets received at tile 101, such as in a ring consumer mode of bus 109. System 100 is operable to decouple packet boundaries from filtering and focus on the data itself without padding. One rationale behind byte-granular filtering is to configure the new ISA so that the TTU can decouple from conventional packet-boundary-based data processing schemes.
[0081] The improved ISA's byte addressing mode and filtering features can use specific fields that allow bus 109 to deliver valid bytes of data in a linearized order. This allows tile 101 to selectively filter bytes in or out of the byte stream as the tile consumes packets of inbound data. To support filtering and other byte-granular operations, example byte address mode instructions of ISA 304 can include fields such as first_discard_byte_count, discard_byte_count (filter_out_bytes), access_bytes, and last_discard_byte_count. In some implementations, each tile can include a register and associated control logic that tracks outstanding access_bytes initialized by access_bytes at the start of an iteration. When a write transaction is generated, the register value reduces the number of valid bytes in this transaction.
[0082] The byte-granular filtering feature can be used in conjunction with the TTU of compute tile 101 to support any arbitrary slicing of one or more tensors, such as a slice of elements along a particular dimension of a tensor. For example, system 100 can utilize the filtering feature to move a slice of a tensor to a different tile, where the slice can be based on any selected dimension. In addition, this can then be extended to any re- layout change, where data associated with a tensor is stored in memory 108 (or 110) based on a first layout and then undergoes a re-layout to change the way (or order) the data is stored in the tile memory.
[0083] Figure 5 Aspects of an example byte addressing operation are shown. This example operation can be associated with a read operation performed using the TTU of compute tile 101.
[0084] In Figure 5In the example of memory 108, the partition includes n number of physical memory banks 502, where n is an integer greater than or equal to 1. For example, the memory of compute tile 101 can include 32 physical memory banks 502, which can be identified as banks 0 through 31. Each bank can include a number of rows, such as 4, 16, 24, etc. Each row has a width that can be defined in bytes. For example, each row can be 16 bytes (16B) wide, such that compute tile 101 accesses data from memory 108 in blocks of 16B (e.g., 128 bits). In some implementations, memory 108 can include more or fewer physical memory banks, and each row of a bank can have a width that is greater than 16B or less than 16B (e.g., IB).
[0085] As described above, previous instruction set packing and alignment requirements limited, which imposed storage and bandwidth overheads that ultimately resulted in adverse effects on performance and power. In particular, these previous instruction sets can be linked with row addressing schemes that unnecessarily constrain memory access operations. In some implementations, for certain sizes and structures of memory, when reading data from the memory, the row addressing scheme can be limited to four byte granularity.
[0086] In Figure 5 In implementations of the present disclosure, although the example rows of memory 108 are shown as 4 bytes (4B) wide, the improved ISA of the present disclosure allows compute tile 101 to access data from tile memory in blocks of IB, 2B, 3B, 4B, or greater than 4B. Control logic (e.g., a controller) of compute tile 101 can identify an opcode and data fields in an instruction to perform a memory operation at one or more banks 502 of memory 108. The opcode and data fields can be processed to generate an example byte addressing sequence 504 for accessing consecutive bytes of different lengths stored at one or more banks 502.
[0087] The instruction can include a “byte count” 506 field that specifies a number of consecutive bytes to fetch with reference to a starting address of a memory location for memory 108, 110. As shown at Figure 5 As shown at, the address can be indicated as “0” or “2”. Compute tile 101 can reference a particular starting address in the byte addressing sequence and then perform a plurality of byte addressable memory accesses at memory 108, 110 based on this starting address. The byte accesses can be performed as part of a tensor read operation to obtain an input vector from one or more banks of memory 108.
[0088] In Figure 5In the example of read operation 508, a read operation is performed (508) to obtain distinct / different portions of byte-level data stored at respective address locations of bank 502 of memory 108. This particular read operation (508) is performed to obtain two bytes (byte count = 2). The operation starts at address "2" in row 0 of memory bank 502 and includes obtaining a first byte-level data stored at address "2" and a second different byte-level data stored at address "3". In some implementations, the distinct portions or length of byte-level data is defined based on the byte count integer. In these implementations, a byte addressing sequence is used to access a number of consecutive bytes, the number (e.g., length) of which is variable (e.g., defined by the instruction).
[0089] Another read operation (510) is performed to obtain distinct portions of byte-level data stored at respective address locations of bank 502. This particular read operation (510) is performed to obtain four bytes (byte count = 4). The operation starts at address "2" in row 0 of memory bank 502 and includes obtaining respective byte-level data stored at addresses 2 and 3 of row 0 and addresses 4 and 5 of row 1.
[0090] Another read operation (512) is performed to obtain distinct portions of byte-level data stored at respective address locations of bank 502. This particular read operation (512) is performed to obtain 10 bytes (byte count = 10). The operation starts at address "0" in row 0 of memory bank 502 and ends at address "9" in row 2 of memory bank 502. Thus, this particular read operation (512) includes obtaining respective byte-level data stored at addresses 0, 1, 2, 3, 4, 5, 6, 7, 8, and 9, where the sequence of addresses spans row 0, row 1, and row 2.
[0091] Read operations 508, 510, 512 can be directed to the same memory bank 502 or different memory banks 502. Notably, read or write operations can start at any valid byte address and need not start at row-aligned addresses 514 such as addresses 4, 8, 12, or 16. As discussed above, the integrated field of the improved ISA allows for automatic zero padding when accessing data from the memory 108, 110 of a compute tile. As an example, for read operations 508 and 512, system 100 can automatically zero one or more bytes based on the byte count accessed for a given read operation.
[0092] Each compute tile 101 includes a TTU 520 for performing access operations at the tile memory. For Figure 5For example read operations, based on the improved ISA's opcode and data fields, TTU 520 is operable to generate an address that can indicate any one of the one or more valid byte addresses. In particular, the streamlined configuration of the improved ISA and associated byte address mode allows TTU 520 to generate addresses at byte-level granularity. TTU 520 is also decoupled from conventional packet-boundary data processing.
[0093] ISA 304 can be configured to provide explicit hints for contiguous data size and its synchronization timing. For example, an "access_bytes" field can be used to define the minimum synchronization granularity, its corresponding timing, and the amount of contiguous data. ISA 304 can enable certain software-agnostic optimizations, such as exploiting TTU byte-level addressing without the overhead of any TTU reprogramming for the new instruction set. In some implementations, these features of the improved ISA can be utilized to simplify architecture design with respect to aspects such as synchronization, coherency, and tag matching.
[0094] In addition to the ISA, compute tile 101 can include specific hardware features and configuration of data pipelines that enable the tile to operate on data more efficiently at those granularities. In some implementations, compute tile 101 can include a coalesce buffer 522 that is used to increase SRAM1 (e.g., narrow memory) write granularity when writing the output of NLU 210 to memory 108. In some implementations, system 100 can implement a new byte-addressing ISA in cooperation with coalesce buffer 522 to perform gather-by-one.
[0095] The byte-addressing ISA and coalesce buffer 522 can be used at compute tile 101 to coalesce data from multiple packets so that the coalesced data can be written in an example contiguous address space at one time (e.g., concurrently). For example, based on coalesce buffer 522, instead of writing data to memory 108 using multiple discrete 4B transactions, a set of data can be written as a single 16B write transaction (524), saving clock cycles for the particular write operation. The 16B value can be indicated by the access_bytes field of an instruction issued via the byte-addressing ISA.
[0096] This coalesce and write operation can increase SRAM1 memory bandwidth and provide a lower likelihood of experiencing bank conflicts at memory 108, 110 relative to previous approaches. In some cases, the multiple packets from which data is coalesced can be data packets generated locally at compute tile 101 or external incoming data packets received at the compute tile via bus 109. In some implementations, this coalesce operation can be used to enhance NLU output granularity and bandwidth when writing the output to memory 108. For example, this can allow for byte-level NLU output granularity.
[0097] Figure 6 is an example process 600 for byte stream processing using example byte addressing operations of an instruction set. The process 600 can be implemented or performed using the system 100 described above. Thus, the description of the process 600 can refer to the above-mentioned computing resources of the system 100. In some examples, the steps or acts of the process 600 are implemented by programmed firmware instructions, software instructions, or both. Each type of instruction can be stored in a non-transitory machine-readable storage and executable by one or more processors of the apparatus and resources described in this document, such as a compute tile of a hardware accelerator or neural network processor.
[0098] In some implementations, the steps of the process 600 are performed at a hardware circuit to generate a layer output of a neural network layer. The output can be part of a computation of a generated image processing or image recognition output of a machine learning task or inference workload. As indicated above, the integrated circuit can be a specialized neural network processor or hardware machine learning accelerator configured to accelerate computations for generating various types of data processing outputs.
[0099] Referring again to the process 600, the system 100 receives an instruction for a compute tile of a hardware circuit (602). The system 100 can identify an opcode in the instruction for an operation involving accessing one or more banks of a first memory of the compute tile (604). For example, the instruction can be a byte address mode instruction that includes an opcode indicating a tensorop read operation or a tensorop write operation. In some implementations, the instruction itself or one or more opcodes of the instruction include a byte count field that indicates a number of bytes to access from a memory location of the compute tile 101.
[0100] The system 100 can determine and / or generate a byte addressing sequence based on the opcode of the instruction and any associated data fields (606). The byte addressing sequence is for accessing consecutive bytes of a distinct length stored at one or more memory banks 502 of the memory 108, 110. For example, the byte addressing sequence can operate to access consecutive bytes of different lengths starting at any byte address from any of a plurality of rows of a memory bank and from any of a plurality of memory banks of the memory 108, 110. In some implementations, the byte addressing sequence can be generated by a TTU of the compute tile to obtain or store an input of an input tensor or a parameter tensor. For example, each address of the byte addressing sequence identifies a memory location for reading (or writing) one byte of data. One or more bytes of data can represent an input or a vector of inputs for processing at a layer of a neural network.
[0101] Based on the sequence of byte addresses, the system 100 can perform a plurality of byte-addressable memory accesses to obtain an input vector from the memory bank (608). For example, the compute tile 101 processes the opcode and parameter values of the byte count field (e.g., access_bytes) in the byte address mode instruction. In response to this processing, the compute tile 101 can generate one or more byte access requests based on the fields or opcode of the instruction. Each byte access request provides byte-granularity access to data for at least one byte at any given address at any given line in the memory bank of the tile.
[0102] The byte access requests can be for write transactions, read transactions, or a combination of these transactions. For example, the byte access requests can be write transactions that include or correspond to a sequence of byte addresses. In some implementations, the byte access requests can be read transactions that incorporate auto-zero padding. An example sequence of byte addresses for a read transaction can include respective addresses of memory locations in the tile memory that store byte-level data for distinct portions. Each address location can correspond to an element of an input tensor that is processed as part of a machine learning workload. For example, each address location and the byte-level data stored at this location can be mapped to a corresponding element of an input tensor.
[0103] As indicated above, the byte-level data can be an input or a vector of inputs, and the system 100 can process the input vector at the hardware circuit based on the instruction (610). For example, the compute tile 101 can generate an output of a neural network layer in response to processing the input vector through the neural network layer. This processing can involve arithmetic operations performed at the compute unit 112. The arithmetic operations can produce a set of outputs that are written (stored) to the memory 108. The outputs can be activation values written to the memory 108 using byte-addressable memory accesses. These byte-level accesses can be performed based on the opcode in the byte address mode instruction. For example, the opcode can indicate a tensorop write transaction involving the TTU 520 of the compute tile 101 and the coalesced buffer 522 coupled to the memory 108 at the tile.
[0104] Figure 7 An example of a tensor or multi-dimensional matrix 700 is shown that includes an input tensor 704, a variant of a parameter tensor 706, and an output tensor 708. In the example of FIG. 7, the input tensor 704 includes a plurality of elements 702 that each correspond to an input value for a respective data value (or operand) of a computation performed at a given layer of a neural network. As indicated above, the data or operands can be inputs, activations, weights (or parameters) for a computation operation involving one or more neural network layers or other types of ML constructs implemented at the compute tile 101. Figure 7
[0105] For example, each input of the input tensor 704 can correspond to a respective element along a given dimension of the input tensor 704, each weight of the parameter tensor 706 can correspond to a respective element along a given dimension of the parameter tensor 706, and each output value or activation in the output set can correspond to a respective element along a given dimension of the output tensor 708. Relatedly, each element can correspond to a respective memory location or address in memory of the compute tile 101 assigned for operating on one or more dimensions of a given tensor 704, 706, 708.
[0106] The computations performed at a given neural network layer can include multiplying the input / activation tensor 704 with the parameter / weight tensor 706 over one or more processor clock cycles to produce a layer output that can include output activations. Multiplying the activation tensor 704 with the weight tensor 706 includes multiplying an activation from an element of the tensor 704 with a weight from an element of the tensor 706 to produce one or more partial sums. Figure 7 An example tensor 706 can be an unmodified parameter tensor, a modified parameter tensor, or a combination thereof. In some implementations, each parameter tensor 706 corresponds to a modified parameter tensor that includes non-zero CSP values derived based on the sparsity utilization techniques described above.
[0107] The processor cores of the system 100 can operate on: i) a scalar corresponding to a discrete element in a certain multi-dimensional tensor 704, 706; ii) a vector (e.g., the input vector 102) including values of multiple discrete elements 707 along the same or different dimensions of a certain multi-dimensional tensor 704, 706; or iii) a combination of these. Depending on the dimensionality of the tensor, each of the discrete element 707 or the multiple discrete elements 707 in a certain multi-dimensional tensor can be represented using X, Y coordinates (2D) or using X, Y, Z coordinates (3D).
[0108] The system 100 can compute multiple partial sums corresponding to products generated by multiplying a batch of inputs with corresponding weight values. As described above, the system 100 can perform accumulation of the products (e.g., partial sums) over multiple clock cycles. For example, accumulation of the products can be performed in the random access memory, shared memory, or scratchpad memory (e.g., the memory 108) of one or more compute tiles 101 based on the techniques described in this document. In some implementations, the accumulation of the products to the memory of the compute tile is performed using the byte addressing mode instructions of the ISA 304.
[0109] In some implementations, the input-weight multiplication can be written as a sum of products of each weight element multiplied by a discrete input of the input vector 102, such as a row or slice of the input tensor 704. This row or slice can represent a given dimension, such as the first dimension 710 of the input tensor 704 or a second, different dimension 715 of the input tensor 704. The byte-granularity filtering features of the improved ISA 304 can be used to support multiplication operations, including any arbitrary slice of one or more tensors, such as a slice of elements along a particular dimension of a tensor.
[0110] In some implementations, the output of a convolutional neural network layer can be computed using an example set of computations. The computations for a CNN layer can involve performing a 2D spatial convolution between a 3D input tensor 704 and at least one 3D filter (weight tensor 706). For example, convolving one 3D filter 706 onto a 3D input tensor 704 can result in a 2D spatial plane 720 or 725. The computations can involve computing a sum of products for a particular dimension of the input volume including the input vector 102.
[0111] For example, the spatial plane 720 can include output values computed from a sum of products of inputs along the dimension 710, while the spatial plane 725 can include output values computed from a sum of products of inputs along the dimension 715. The computations for generating the sum of products of output values in each of the spatial planes 720 and 725 can be performed: i) at the computation cells 114 a / b / c, ii) directly at the memory 110 using arithmetic operators coupled to the shared library of the memory 110, iii) or both. In some implementations, the reduction operations can be simplified and performed directly at the memory cells (or locations) of the memory 110 using various techniques for reduction of values for accumulation.
[0112] Embodiments of the subject matter and the functional operations described in this specification can be implemented in digital electronic circuitry, in tangibly-embodied computer software or firmware, in computer hardware, including the structures disclosed in this specification and their structural equivalents, or in combinations of one or more of them. Embodiments of the subject matter described in this specification can be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a tangible non-transitory program carrier for execution by, or to control the operation of, data processing apparatus.
[0113] Alternatively or additionally, the program instructions can be encoded on an artificially generated propagated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal that is generated to encode information for transmission to suitable receiver apparatus for execution by a data processing apparatus. The computer storage medium can be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or a combination of one or more of them.
[0114] The term“computing system” encompasses all kinds of apparatus, devices, and machines for processing data, including by way of example a programmable processor, a computer, or multiple processors or computers. The apparatus can include special purpose logic circuitry, e.g., an FPGA (field programmable gate array) or an ASIC (application specific integrated circuit). The apparatus can also include, in addition to hardware, code that creates an execution environment for computer programs, e.g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them.
[0115] A computer program, which can also be referred to or described as a program, software, a software application, a module, a software module, a script, or code, can be written in any form of programming language, including compiled or interpreted languages, or declarative or procedural languages, and it can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A program may, but need not, correspond to a file in a file system. A program can be stored in a portion of a file that holds other programs or data, e.g., one or more scripts stored in a markup language document, in a single file dedicated to the program in question, or in multiple coordinated files, e.g., files that store one or more modules, sub programs, or code portions. A computer program can be deployed to be executed on one computer or on multiple computers that are located at one site or distributed across multiple sites and
[0116] A computer program may, but need not, correspond to a file in a file system. A program can be stored in a portion of a file that holds other programs or data, e.g., one or more scripts stored in a markup language document, in a single file dedicated to the program in question, or in multiple coordinated files, e.g., files that store one or more modules, sub programs, or code portions. A computer program can be deployed to be executed on one computer or on multiple computers that are located at one site or distributed across multiple sites and
[0117] The processes and logic flows described in this specification can be performed by one or more programmable computers executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows can also be performed by special purpose logic circuitry, e.g., an FPGA (field programmable gate array) or an ASIC (application specific integrated circuit), and / or devices can be implemented as special purpose logic circuitry, e.g., an FPGA (field programmable gate array), an ASIC (application specific integrated circuit), or a GPGPU (general purpose graphics processing unit).
[0118] Computers suitable for the execution of a computer program include, by way of example, can be based on general or special purpose microprocessors or both, or any other kind of central processing unit. Generally, a central processing unit will receive instructions and data from a read-only memory or a random access memory or both. The elements of a computer generally are a central processing unit for performing or executing instructions and one or more memory devices for storing instructions and data. Generally, a computer will also include, or be operatively coupled to receive data from or transfer data to, or both, one or more mass storage devices for storing data, e.g., magnetic, magneto-optical disks, or optical disks. However, a computer need not have such devices. Moreover, a computer can be embedded in another device, e.g., a mobile telephone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a Global Positioning System (GPS) receiver, or a portable storage device, e.g., a universal serial bus (USB) flash drive, to name just a few.
[0119] Computer readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media and memory devices, including by way of example semiconductor memory devices, e.g., EPROM, EEPROM, and flash memory devices; magnetic disks, e.g., internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The processor and the memory can be supplemented by, or incorporated in, special purpose logic circuitry.
[0120] To provide for interaction with a user, embodiments of the subject matter described in this specification can be implemented on a computer having a display device, e.g., an LCD (liquid crystal display) monitor, for displaying information to the user and a keyboard and a pointing device, e.g., a mouse or a trackball, by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback, e.g., visual feedback, auditory feedback, or tactile feedback; and input from the user can be received in any form, including acoustic, speech, or tactile input. In addition, a computer can interact with a user by sending documents to and receiving documents from a device of the user; for example, by sending web pages to a web browser on a user’s client device in response to requests received from the web browser.
[0121] Embodiments of the subject matter described in this specification can be implemented in a computing system that includes a back-end component, e.g., as a data server, or that includes a middleware component, e.g., an application server, or that includes a front-end component, e.g., a client computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the subject matter described in this specification, or any combination of one or more such back-end, middleware, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication, e.g., a communication network. Examples of communication networks include a local area network ("LAN") and a wide area network ("WAN"), e.g., the Internet.
[0122] The computing system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other.
[0123] While this specification contains many specific implementation details, these should not be construed as limitations on the scope of any inventions, but rather as descriptions of features that can be specific to particular embodiments of particular inventions. Certain features that are described in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features that are described in the context of a single embodiment can also be implemented in multiple embodiments separately or in any suitable sub-combination. Moreover, although features can be described above as acting in certain combinations and even initially claimed as such, one or more features from a claimed combination can in some cases be excised from the combination and the claimed combination can be directed to a sub-combination or variation of a sub-combination.
[0124] Similarly, while operations are depicted in the drawings in a particular order, this should not be understood as requiring such an order, nor that all illustrated operations be performed, to achieve desirable results. In certain circumstances, multitasking and parallel processing can be advantageous. Moreover, the separation of various system modules and components in the embodiments described above should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.
[0125] Particular embodiments of the subject matter have been described. Other embodiments are within the scope of the following claims. For example, the acts recited in the claims can be performed in a different order and still achieve desirable results. As one example, the processes depicted in the accompanying figures do not necessarily require the particular order shown or sequential order to achieve the desired results. In certain implementations, multitasking and parallel processing can be advantageous.
[0126] In the following sections, additional examples are provided:
[0127] Example 1 : A computer-implemented method for accessing a memory bank of a hardware accelerator, the method comprising: receiving an instruction for a compute tile of the hardware accelerator, the instruction being executable at the compute tile to cause performance of operations comprising: identifying, in the instruction, an opcode for an operation involving access to one or more banks of a first memory of the compute tile; generating, based on the opcode, a byte addressing sequence for accessing consecutive bytes of different lengths stored at the one or more banks; performing, based on the byte addressing sequence, a plurality of byte-addressable memory accesses at the first memory to obtain a plurality of input vectors from the one or more banks of the first memory; and processing, based on the instruction, the plurality of input vectors at the hardware accelerator.
[0128] Example 2: The method of example 1, wherein the byte addressing sequence is operable to access the consecutive bytes of different lengths starting from any given byte address across the one or more banks of the first memory.
[0129] Example 3 : The method of example 1 or 2, further comprising: generating, based on the opcode, a byte access request that provides byte-granularity access for obtaining data for at least one byte stored at a given row of a bank among the one or more banks of the first memory.
[0130] Example 4: The method of any preceding example, wherein generating the byte addressing sequence comprises: generating a byte addressing sequence corresponding to a plurality of byte access requests, each byte access request being for accessing a distinct portion of byte-level data stored at a respective address location across the one or more banks of the first memory.
[0131] Example 5 : The method of any preceding example, wherein the opcode is for a tensor operation executable at the compute tile to iterate over respective elements of an input tensor.
[0132] Example 6: The method of example 5 as from example 4, further comprising: executing the plurality of byte access requests using the byte addressing sequence; based on the plurality of byte access requests executed, performing a plurality of byte-granularity accesses at the first memory; and traversing respective elements of the input tensor by performing the plurality of byte-granularity accesses at the first memory.
[0133] Example 7: The method of example 6 or example 5 as from example 4, wherein: each distinct portion of byte-level data stored at respective address locations across the one or more banks of the first memory corresponds to a respective element of the input tensor; and each respective address location storing a distinct portion of byte-level data is mapped to a corresponding element of the input tensor.
[0134] Example 8: The method of example 6 or example 7 as from example 6, wherein: a portion of byte-level data corresponds to an input vector of the plurality of input vectors; and the plurality of byte-addressable memory accesses to obtain a plurality of input vectors corresponds to at least one of: i) the plurality of byte-granularity accesses and ii) the plurality of byte access requests.
[0135] Example 9: The method of any preceding example, wherein processing the plurality of input vectors comprises: processing the plurality of input vectors by a neural network layer of a neural network implemented at the hardware accelerator; and generating an output of the neural network layer in response to processing the plurality of input vectors by the neural network layer.
[0136] Example 10: A system comprising a hardware accelerator; a processing device; and a non-transitory machine-readable storage medium for storing instructions for accessing memory banks of a hardware accelerator, the instructions executable by the processing device to cause performance of operations comprising: receiving instructions for a compute tile of the hardware accelerator, the instructions executable at the compute tile to cause performance of operations comprising: identifying, in the instructions, an opcode for an operation involving accessing one or more banks of a first memory of the compute tile; generating, based on the opcode, a byte addressing sequence for accessing consecutive bytes of different lengths stored at the one or more banks; based on the byte addressing sequence, performing a plurality of byte-addressable memory accesses at the first memory to obtain a plurality of input vectors from the one or more banks of the first memory; and based on the instructions, processing the plurality of input vectors at the hardware accelerator.
[0137] Example 11: The system of example 10, wherein the byte addressing sequence is operable to access the consecutive bytes of different lengths starting from any given byte address across the one or more banks of the first memory.
[0138] Example 12: The system of example 11, wherein the operations further comprise: generating a byte access request based on the operation code, the byte access request providing byte-granularity access for obtaining at least one byte of data stored at a given row of a bank among the one or more banks of the first memory.
[0139] Example 13: The system of example 11 or 12, wherein generating the byte addressing sequence comprises: generating a byte addressing sequence corresponding to a plurality of byte access requests, each byte access request for accessing a distinct portion of byte-level data stored at a respective address location across the one or more banks of the first memory.
[0140] Example 14: The system of any one of examples 10 to 13, wherein the operation code is a tensor operation for execution at the compute tile to iterate over respective elements of an input tensor.
[0141] Example 15: The system of example 14 when dependent on example 13, wherein the operations further comprise: executing the plurality of byte access requests using the byte addressing sequence; performing a plurality of byte-granularity accesses at the first memory based on the plurality of byte access requests that are executed; and iterating over respective elements of the input tensor by performing the plurality of byte-granularity accesses at the first memory.
[0142] Example 16: The system of example 15 or example 14 when dependent on example 13, wherein: each distinct portion of byte-level data stored at a respective address location across the one or more banks of the first memory corresponds to a respective element of the input tensor; and each respective address location storing a distinct portion of byte-level data is mapped to a corresponding element of the input tensor.
[0143] Example 17: The system of example 15 or example 16 when dependent on example 15, wherein: a portion of byte-level data corresponds to an input vector of the plurality of input vectors; and the plurality of byte-addressable memory accesses for obtaining the plurality of input vectors corresponds to at least one of: i) the plurality of byte-granularity accesses and ii) the plurality of byte access requests.
[0144] Example 18: The system of any one of examples 10 to 17, wherein processing the plurality of input vectors comprises: processing the plurality of input vectors by a neural network layer of a neural network implemented at the hardware accelerator; and generating an output of the neural network layer in response to processing the plurality of input vectors by the neural network layer.
[0145] Example 19: A non-transitory machine-readable storage medium for storing instructions for accessing a memory bank of a hardware accelerator, the instructions executable by a processing device to cause performance of operations comprising: receiving an instruction for a compute tile of the hardware accelerator, the instruction executable at the compute tile to cause performance of operations comprising: identifying, in the instruction, an opcode for an operation involving access to one or more banks of a first memory of the compute tile; generating, based on the opcode, a byte addressing sequence for accessing consecutive bytes of different lengths stored at the one or more banks; performing, based on the byte addressing sequence, a plurality of byte-addressable memory accesses at the first memory to obtain a plurality of input vectors from the one or more banks of the first memory; and processing, based on the instruction, the plurality of input vectors at the hardware accelerator.
[0146] Example 20: The non-transitory machine-readable storage medium of example 19, wherein: the byte addressing sequence is operable to access the consecutive bytes of different lengths starting from any given byte address across the one or more banks of the first memory; and the operations further comprise: generating, based on the opcode, a byte access request that provides byte-granularity access for obtaining data for at least one byte stored at a given row of a bank among the one or more banks of the first memory.
Claims
1. A computer-implemented method for accessing a memory bank of a hardware accelerator, the method comprising: receiving an instruction for a compute tile of the hardware accelerator, the instruction being executable at the compute tile to cause performance of operations comprising: identifying, in the instruction, an opcode for an operation involving access to one or more banks of a first memory of the compute tile; generating, based on the opcode, a byte addressing sequence for accessing consecutive bytes of different lengths stored at the one or more banks; performing, based on the byte addressing sequence, a plurality of byte-addressable memory accesses at the first memory to obtain a plurality of input vectors from the one or more banks of the first memory; and processing, based on the instruction, the plurality of input vectors at the hardware accelerator.
2. The method of claim 1, wherein the byte addressing sequence is operable to access the consecutive bytes of different lengths starting from any given byte address across the one or more banks of the first memory.
3. The method of claim 1 or 2, further comprising: generating, based on the opcode, a byte access request providing byte-granularity access for obtaining data for at least one byte stored at a given row of a bank among the one or more banks of the first memory.
4. The method of any preceding claim, wherein generating the byte addressing sequence comprises: generating a byte addressing sequence corresponding to a plurality of byte access requests, each byte access request being for accessing one or more distinct portions of byte-level data stored at respective address locations across the one or more banks of the first memory.
5. The method of any preceding claim, wherein the opcode is a tensor operation for execution at the compute tile to iterate over respective elements of an input tensor.
6. The method of claim 5 when dependent on claim 4, further comprising: performing the plurality of byte access requests using the byte addressing sequence; performing a plurality of byte-granularity accesses at the first memory based on the performed plurality of byte access requests; and iterating over respective elements of the input tensor by performing the plurality of byte-granularity accesses at the first memory.
7. The method of claim 6 or claim 5 when dependent on claim 4, wherein: each distinct portion of byte-level data stored at respective address locations across the one or more banks of the first memory corresponds to a respective element of the input tensor; and each respective address location storing a distinct portion of byte-level data is mapped to a corresponding element of the input tensor.
8. The method of claim 6 or claim 7 when dependent on claim 6, wherein: a portion of byte-level data corresponds to an input vector of the plurality of input vectors; and the plurality of byte-addressable memory accesses for obtaining a plurality of input vectors corresponds to at least one of: i) the plurality of byte-granularity accesses and ii) the plurality of byte access requests. 9. The method of any preceding claim, wherein processing the plurality of input vectors comprises: processing the plurality of input vectors by a neural network layer of a neural network implemented at the hardware accelerator; and generating an output of the neural network layer in response to processing the plurality of input vectors by the neural network layer.
10. A system comprising a hardware accelerator; a processing device; and a non-transitory machine-readable storage medium for storing instructions for accessing a memory bank of a hardware accelerator, the instructions executable by the processing device to cause performance of operations comprising: receiving instructions for a compute tile of the hardware accelerator, the instructions executable at the compute tile to cause performance of operations comprising: identifying, in the instructions, an opcode for an operation involving access to one or more banks of a first memory of the compute tile; generating, based on the opcode, a byte addressing sequence for accessing consecutive bytes of different lengths stored at the one or more banks; performing, based on the byte addressing sequence, a plurality of byte-addressable memory accesses at the first memory to obtain a plurality of input vectors from the one or more banks of the first memory; and processing, based on the instructions, the plurality of input vectors at the hardware accelerator.
11. The system of claim 10, wherein the byte addressing sequence is operable to access the consecutive bytes of different lengths starting from any given byte address across the one or more banks of the first memory.
12. The system of claim 11, wherein the operations further comprise: generating, based on the opcode, a byte access request that provides byte-granularity access for obtaining at least one byte of data stored at a given row of a bank among the one or more banks of the first memory.
13. The system of claim 11 or 12, wherein generating the byte addressing sequence comprises: generating a byte addressing sequence corresponding to a plurality of byte access requests, each byte access request for accessing a distinct portion of byte-level data stored at a respective address location across the one or more banks of the first memory.
14. The system of any of claims 10 to 13, wherein the opcode is a tensor operation for execution at the compute tile to iterate over respective elements of an input tensor.
15. The system of claim 14 when dependent on claim 13, wherein the operations further comprise: executing the plurality of byte access requests using the byte addressing sequence; performing a plurality of byte-granularity accesses at the first memory based on the executed plurality of byte access requests; and iterating over respective elements of the input tensor by performing the plurality of byte-granularity accesses at the first memory.
16. The system of claim 15 or claim 14 when dependent on claim 13, wherein: each distinct portion of byte-level data stored at a respective address location across the one or more banks of the first memory corresponds to a respective element of the input tensor; and and Each respective address location storing a distinct portion of byte-level data is mapped to a corresponding element of the input tensor.
17. The system of claim 15 or claim 16 when dependent on claim 15, wherein: a portion of byte-level data corresponds to an input vector of the plurality of input vectors; and the plurality of byte-addressable memory accesses to obtain the plurality of input vectors correspond to at least one of: i) the plurality of byte-granularity accesses and ii) the plurality of byte access requests.
18. The system of any one of claims 10 to 17, wherein processing the plurality of input vectors comprises: processing the plurality of input vectors by a neural network layer of a neural network implemented at the hardware accelerator; and generating an output of the neural network layer in response to processing the plurality of input vectors by the neural network layer.
19. A non-transitory machine-readable storage medium for storing instructions for accessing a memory bank of a hardware accelerator, the instructions executable by a processing device to cause performance of operations comprising: receiving instructions for a compute tile of the hardware accelerator, the instructions executable at the compute tile to cause performance of operations comprising: identifying, in the instructions, an opcode for an operation involving accessing one or more banks of a first memory of the compute tile; generating, based on the opcode, a byte-addressing sequence for accessing consecutive bytes of different lengths stored at the one or more banks; performing, based on the byte-addressing sequence, a plurality of byte-addressable memory accesses at the first memory to obtain a plurality of input vectors from the one or more banks of the first memory; and processing, based on the instructions, the plurality of input vectors at the hardware accelerator.
20. The non-transitory machine-readable storage medium of claim 19, wherein: the byte-addressing sequence is operable to access the consecutive bytes of different lengths starting from any given byte address across the one or more banks of the first memory; and the operations further comprise: generating, based on the opcode, a byte access request, the byte access request providing byte-granularity access for obtaining data for at least one byte stored at a given row of a bank among the one or more banks of the first memory.