Utilization of Data Sparsity in Machine Learning Hardware Accelerators
By employing compressed sparse parameters and mapping vectors to eliminate zero-valued operations, the hardware accelerator architecture enhances efficiency and reduces power consumption in neural network computations.
Patent Information
- Application Number
- JP2024568259
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2022-05-18
- Publication Date
- 2025-05-30
AI Technical Summary
Existing hardware accelerators for neural networks face inefficiencies in processing due to the presence of zero-valued operands in weight tensors, which leads to wasted computational cycles and increased power consumption.
The implementation of a hardware accelerator architecture that utilizes compressed sparse parameters and mapping vectors to perform sparse computations. This involves deriving compressed sparse parameters from parameter tensors, generating mapping vectors, and using these to process input vectors through neural network layers, thereby eliminating zero-valued operations.
This approach significantly reduces computational workload and power consumption by eliminating operations with zero-valued operands, leading to improved efficiency and performance in neural network computations.
Smart Images

Figure 2025516768000001_ABST
Abstract
Description
Background Art
[0001] This specification generally relates to using hardware integrated circuits to perform group convolutions of convolutional neural networks.
[0002] A neural network is a machine learning model that utilizes one or more layers of nodes to generate an output, such as a classification, for received inputs. Some neural networks include one or more hidden layers in addition to an output layer. Some neural networks can be convolutional neural networks (CNNs) configured for image processing, or recurrent neural networks (RNNs) configured for audio and language processing. Various types of neural network architectures can be used to perform various tasks related to classification or pattern recognition, prediction involving data modeling, and information clustering.
[0003] A neural network layer can have a corresponding set of parameters or weights. The weights are used to process an input (e.g., a batch of inputs) through the neural network layer and generate a corresponding output of the layer for calculating neural network inference. A batch of inputs and a set of kernels can be represented as tensors of inputs and weights, i.e., multi-dimensional arrays. A hardware accelerator is a dedicated integrated circuit for implementing a neural network. This circuit includes memory having positions corresponding to elements of the tensors that may be traversed or accessed using the control logic of the circuit.
Summary of the Invention
[0004] This document describes an integrated circuit architecture improved for a hardware accelerator, and corresponding techniques for processing an input vector using a set of mapping vectors and compressed sparse parameters in a neural network layer. Each of the set of mapping vectors and compressed sparse parameters can be generated based on an operation code ("opcode") that indicates a uniform sparsity format of a plurality of parameter tensors. The parameter tensors are associated with neural network layers of an artificial neural network, such as a CNN. The disclosed techniques can be used to accelerate tensor operations that support neural network computations involving processing an input of an input vector through one or more of the neural network layers.
[0005] One aspect of the subject matter described in this specification can be embodied in a computer-implemented method that includes a neural network implemented on a hardware accelerator. The method includes deriving a set of compressed sparse parameters from a parameter tensor, generating a mapping vector based on the set of compressed sparse parameters, processing instructions that direct a sparse computation to be performed using the compressed sparse parameters based on the sparsity of the parameter tensor, and based on the instructions, obtaining i) an input vector from a first memory of the hardware accelerator and ii) the compressed sparse parameters from a second memory of the hardware accelerator, and performing the sparse computation to process the input vector through a layer of the neural network using the mapping vector and the set of compressed sparse parameters.
[0006] These and other embodiments can each optionally include one or more of the following features. For example, in some embodiments, processing the input vector through a layer of the neural network includes performing a dot product matrix multiplication operation between the input of the input vector and corresponding weight values within the set of compressed sparse parameters.
[0007] In some embodiments, the dot product matrix multiplication operation is performed on one or more multiplication cells of the hardware accelerator based on the respective bit values of each bit in the mapping vector. In some embodiments, the method further includes accessing hardware selection logic coupled to a first memory and a second memory of the hardware accelerator, and using the hardware selection logic to select a particular input of the input vector and a corresponding weight value within a set of compressed sparse parameters based on the respective bit values of the bits in the mapping vector.
[0008] Deriving a set of compressed sparse parameters can include generating a modified parameter tensor that includes only non-zero elements along a particular dimension of the parameter tensor. Generating the modified parameter tensor can include, in the case of a particular column dimension of the parameter tensor, generating a compressed representation of the column dimension based on the non-zero elements of the column dimension and concatenating each non-zero element within the compressed representation of the column dimension.
[0009] In some embodiments, generating the modified parameter tensor includes holding the respective dimensional positions of each non-zero element within the parameter tensor prior to generating the modified parameter tensor. The parameter tensor can include a plurality of dimensions, and an opcode within an instruction indicates the sparsity of a particular one of these plurality of dimensions. The hardware accelerator is operable to process multi-dimensional parameter tensors, and an opcode within an instruction can indicate the uniform sparsity across each of these multi-dimensional parameter tensors.
[0010] The first memory can be the scratch - pad memory of the hardware accelerator and can be configured to store the inputs processed by the neural network layer and the activation functions. The second memory can include SIMD (single instruction, multiple data) registers, and the method includes storing a mapping vector at a first address of the SIMD registers and storing a set of compressed sparse parameters at different second addresses of the SIMD registers.
[0011] Other aspects of the subject matter described herein can be embodied in a computer - implemented method executed using a hardware accelerator that implements a neural network including a plurality of neural network layers. The method includes receiving instructions for a compute tile of the hardware accelerator.
[0012] This instruction is executable on the compute tile and causes the execution of operations including identifying an opcode within the instruction that indicates the sparsity of a parameter tensor, loading a set of compressed sparse parameters based on weight values derived from a parameter tensor that specifies weights for a layer of the neural network, and loading a mapping vector generated based on the set of compressed sparse parameters.
[0013] The operations include, based on the opcode, i) obtaining an input vector from a first memory of the hardware accelerator and ii) obtaining a set of compressed sparse parameters from a second memory of the hardware accelerator. The operations further include using the set of compressed sparse parameters to process the input vector through a layer of the neural network based on the mapping vector.
[0014] Other embodiments of this and other aspects include corresponding systems, apparatus, and computer programs configured to perform the actions of a method, encoded on a computer storage device. One or more computer systems can be configured in such a way by software, firmware, hardware, or combinations thereof, installed in the system, that during operation cause the system to perform the actions. One or more computer programs can be configured in such a way by having instructions that, when executed by a data processing apparatus, cause the apparatus to perform the actions.
[0015] The subject matter described herein can be implemented in certain embodiments so as to realize one or more of the following advantages. Techniques are described for leveraging sparsity in data processed for machine learning computations. By utilizing compressed sparse parameters having only non-zero weight values, for example, when processing an input image using a CNN machine learning model implemented on a computing device such as a tablet or smartphone, certain hardware and computing efficiencies are realized.
[0016] Computing efficiency is realized by generating compressed sparse parameters and corresponding mapping vectors when leveraging sparsity to accelerate the execution of artificial neural networks. The system detects future sparsity patterns in the data set processed in the neural network layer and generates a set of compressed sparse parameters that contain only non-zero values. The mapping vector rationalizes the processing of the data set by utilizing a specific hardware architecture of an application-specific integrated circuit that accelerates the execution of the artificial neural network by mapping discrete inputs of the input vector to the non-zero values of the compressed sparse parameters.
[0017] Multiplication operations involving zero-valued operands are generally regarded as wasted computational cycles. By processing neural network inputs that have only non-zero values, at least using compressed sparse parameters, a machine learning system can reduce its overall computational workload. This reduction is achieved by removing zero values from the weight values of the parameter tensors being processed for a neural network layer. When the computational workload is reduced, the corresponding power consumption and resource requirements (e.g., memory allocation and processor cycles) are also reduced.
[0018] Details of one or more embodiments of the subject matter described in this specification are set forth in the accompanying drawings and the description below. Other potential features, aspects, and advantages of the subject matter will become apparent from the description, the drawings, and the claims.
Brief Description of the Drawings
[0019]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Modes for Carrying Out the Invention
[0020] Like reference symbols and designations in the various drawings refer to like elements.
[0021] FIG. 1 is a block diagram of an example of a computing system 100 for implementing a neural network model in a hardware integrated circuit, such as a machine learning hardware accelerator. The computing system 100 includes one or more computing tiles 101, a host 120, and a higher-level controller 125 (the "controller 125"). As will be described in more detail below, the host 120 and the controller 125 cooperate to provide a data set and instructions to one or more computing tiles 101 of the system 100.
[0022] In some embodiments, the host 120 and the controller 125 are the same device. Also, the host 120 and the controller 125 can perform separate functions but can also be integrated into a single device package. For example, the host 120 and the controller 125 can form a central processing unit (CPU) that interacts with or cooperates with a hardware accelerator that includes a plurality of computing tiles 101. In some embodiments, the host 120, the controller 125, and the plurality of computing tiles 101 are included or formed on a single integrated circuit die. For example, the host 120, the controller 125, and the plurality of computing tiles 101 can form a dedicated system-on-chip (SoC) optimized to execute a neural network model for processing machine learning workloads.
[0023] Each computing tile 101 generally includes a controller 103 that provides one or more control signals 105 for storing the input (or activation function) of the input vector 102 at a memory location in the first memory 108 (the "memory 108") or accessing the input therefrom. Similarly, the controller 103 can also provide one or more control signals 105 for storing the weights (or parameters) of the matrix structure of the weights 104 at a memory location in the second memory 110 (the "memory 110") or accessing the weights therefrom. In some embodiments, the input vector 102 is obtained from an input tensor, while the matrix structure of the weights 103 is obtained from a parameter tensor. Each of the input tensor and the parameter tensor may be a multi-dimensional data structure such as a multi-dimensional matrix or tensor. This will be described in more detail below with reference to FIG. 7.
[0024] Each memory location in memories 108, 110 can be identified by a corresponding memory address. Each of memories 108, 110 can be implemented as a series of banks, units, or any other relevant storage medium or device. Each of memories 108, 110 can include one or more registers, buffers, or both. Generally, the controller 103 mediates access to each of memories 108, 110. In some embodiments, the input or activation function is stored in memory 108, memory 110, or both, and the weights are stored in memory 110, memory 108, or both. For example, the input and weights can be transferred between memory 108 and memory 110 to facilitate a particular neural network calculation.
[0025] Each computing tile 101 also includes a computing unit 112 having an input activation function bus 106, an output activation function bus 107, and multiply-accumulate cells (MACs) 114a / b / c. The controller 103 can generate a control signal 105 for obtaining operands stored in the memory of the computing tile 101. For example, the controller 103 can generate a control signal 105 for obtaining i) an exemplary input vector 102 stored in the memory 108, and ii) weights 104 stored in the memory 110. Each input obtained from the memory 108 is provided to the input activation function bus 106 for routing (e.g., directly routing) to the computing cells 114a / b / c within the computing unit 112. Similarly, each weight obtained from the memory 110 is routed to the cells 114a / b / c of the computing unit 104.
[0026] As described below, each cell 114a / b / c performs a calculation to generate a partial sum or an accumulated value for generating the output of a given neural network layer. An activation function can be applied to the set of outputs to generate a set of output activation functions of the neural network layer. In some embodiments, the output or the output activation functions are routed for storage and / or transfer by the output activation function bus 107. For example, the set of output activation functions can be transferred from a first computing tile 101 to a different second computing tile 101 for processing in the second computing tile 101 as input activation functions of different layers of the neural network.
[0027] Generally, each computing tile 101 and system 100 can include additional hardware structures to perform computations associated with multi-dimensional data structures such as tensors, matrices, and / or data arrays. In some embodiments, the input of the input vector (or tensor) 102, and the weights 104 of the parameter tensor can be pre-loaded into the memories 108, 110 of the computing tile 101. The input and weights are received as a set of data values reaching a particular computing tile 101 from a host 120 (e.g., an external host), via a host interface, or from a higher-level control such as a controller 125.
[0028] Each of the computing tile 101 and the controller 103 can include one or more processors, processing devices, and various types of memories. In some embodiments, the processors of the computing tile 101 and the controller 103 include one or more devices such as a microprocessor or a central processing unit (CPU), a graphics processing unit (GPU), an application-specific integrated circuit (ASIC), or a combination of different processors. Each of the computing tile 101 and the controller 103 can also include other computing and storage resources such as buffers, registers, control circuits, etc. These resources cooperate to provide additional processing options for performing one or more of the determinations and computations described herein.
[0029] In some embodiments, the processing unit(s) of the controller 103 execute program instructions stored in a memory for the controller 103 and the compute tile 101 to perform one or more of the functions described herein. The memory of the controller 103 can include one or more non-transitory machine-readable storage media. Non-transitory machine-readable storage media can include solid state memory, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (e.g., EPROM, EEPROM, or flash memory), or any other tangible medium capable of storing information or instructions.
[0030] The system 100 receives instructions that define a particular compute operation for execution by the compute tile 101. In some embodiments, the host can generate a set of compression parameters (CSP) and a corresponding mapping vector, e.g., a non-zero map (NZM), for a given operation. For example, the host 120 can transmit the compression parameters to the compute tile 101 via the host interface for further processing at that tile. The controller 103 can execute instructions programmed to analyze the received weights and data stream associated with the input, including the compression parameters and the corresponding mapping vector.
[0031] The controller 103 stores the input of the data stream and the weights in the computing tile 101. For example, the controller 103 can store the mapping vector and the compressed sparse parameters in the memory of the computing tile 101. This will be described in more detail below. Also, the controller 102 can analyze the input data stream to detect the operation code ("opcode"). Based on the opcode, the controller 102 can activate the dedicated data path logic associated with one or more computing cells 114a / b / c to perform sparse calculations using the compressed sparse parameters and the corresponding mapping vectors. As used in this document, sparse calculations include neural network calculations performed on a neural network layer using non-zero weight values within a set of compressed sparse parameters generated from a set of weights of the neural network layer.
[0032] In some embodiments, the opcode indicates the sparsity of one or more parameter tensors based on the values of K and N (described below). The controller 103 detects the opcode, including any associated tensor sparsity information, and based on that opcode, uses local read logic to retrieve the compressed parameters from the tile memory and wire or route those compressed parameters to the MACs 114a / b / c of the computing tile 101.
[0033] As will be described in detail below, the controller 103 can also analyze an exemplary data stream and, based on that analysis, generate a set of compressed sparse parameters and the corresponding mapping vectors that map the discrete inputs of the input vector to the compressed sparse parameters. To the extent that the operations and / or processes for generating the compressed sparse parameters and the corresponding mapping vectors are described with reference to the controller 103, each of these operations and processes can also be performed by the host 120, the controller 125, or both.
[0034] In some embodiments, by performing some (or all) of the operations on the host 120, such as analyzing tensor indices, performing direct memory access (DMA) operations to read an address space within the system memory (e.g., DRAM) to obtain input and weight values, generating compressed sparse parameters, and generating corresponding mapping vectors, it is possible to reduce the processing time in each computation tile 101 and improve the data throughput in the system 100. For example, by using the controller 125 to perform these operations on the host 120, if a set of already compressed parameters can be sent to a given tile computation 101, the size and amount of data that needs to be routed in the system 100 are reduced.
[0035] FIG. 2 shows an example of a parameter tensor 200 having a sparsity of K out of N, which can represent a uniform sparsity format indicated by the sparse tensor. Generally, in the case of K out of N sparsity, for every next N elements along a dimension of the tensor (e.g., the innermost dimension), K elements are non-zero.
[0036] One or more opcodes can indicate or specify the sparsity not only of the sparsity attributes of one or more parameter tensors, but also along a particular column (or row) dimension of a given tensor. For example, the opcode of a single instruction received at the computation tile 101 can specify the K out of N sparsity of the parameter tensor 200, including the K out of N sparsity of each column 202 or row 204 of the parameter tensor 200. In some embodiments, the tensor sparsity information specified by the opcode is based on the structure or configuration of the instruction set used in the system 100.
[0037] In the example of FIG. 2, K represents one or more non-zero values, and N is the number of elements of a given parameter tensor 200. In some examples, N is the number of elements for a given row or column of the parameter tensor. Each of K and N is an integer. N can be 1 or greater, while K can be 0 or greater. The sparsity of K / N can be a ratio or some other numerical value that is assigned to the sparsity parameter or communicated as the sparsity parameter.
[0038] The sparsity parameter characterizes a sparsity attribute or measure of sparsity in a dataset or tensor 200. For example, the sparsity parameter can represent the compression rate of a given {K,N} pair, equal to K / N. For instance, when K = 2 and N = 4, the compression rate is 50%. System 100 can support cases where the parameter is compressed in one (or more) dimension(s) along the column dimension corresponding to column 202, for example. For this particular type of reduction operation, column 202 can be described as the reduction dimension or the inner product dimension. In some embodiments, the sparsity in the dataset is based on the sparsity of one or more patterns that are detectable between the training phase of the neural network model, the deployment phase of the neural network model, or both.
[0039] The pattern of sparsity can be uniformly distributed among machine learning datasets, such as the parameter tensor 200 processed during the training and deployment phases of model execution. The uniformity of the sparsity pattern enables a specific measure of predictability that can be leveraged to realize the efficiency in the acceleration of neural network models. For example, as described below, a uniformly distributed sparsity pattern enables predicting, inferring, or otherwise detecting future patterns (e.g., sparsity attributes) of zero or non-zero weight values. In some embodiments, each of the controllers 103, 125 can be configured to learn, explore, and utilize different pattern options to achieve additional efficiency and optimization during model execution.
[0040] In the example of FIG. 2, one or more opcodes received at the compute tile 101 can indicate that each of the column 202 and row 204 includes a sparsity of K out of N of 1 / 2 (where K = 4 and N = 8 here). The controller 103 determines the value of the sparsity parameter based on the logical formula: % sparsity = K ÷ N. In this example, the controller 103 can assign a value of 1 / 2 to each of the sparsity parameters for each of the columns 202 and 204. In connection therewith, the opcode received at the compute tile 101 can also specify that the row 206, which can also be a column, includes a sparsity of K out of N of 5 / 8 (where K = 5 and N = 8 here). In some embodiments, K for a given K out of N sparsity is determined based on the hardware layout of the compute tile 101. For example, K can be determined based on the number of MAC circuits within the hardware compute cells of the compute units 112 in a given compute tile 101.
[0041] FIG. 3 shows an exemplary first architecture 300 for processing a parameter tensor to generate compressed sparse parameters, and FIG. 4 shows an exemplary second architecture 400 for processing a parameter tensor to generate compressed sparse parameters. In view of the similarity between architectures 300 and 400, each of FIGS. 3 and 4 is described in parallel by the following paragraphs.
[0042] The controller 103 processes the opcode and triggers one or more operations for leveraging sparsity in a dataset of a machine learning workload. The controller 103 can process the opcode to identify, or determine, for a given parameter tensor, respective measures of sparsity of the parameter tensor, such as rows of the tensor, or columns of the tensor (e.g., sparsity parameter values). This operation can also include an analysis of the parameter tensor. In some embodiments, the controller 103 triggers an operation for leveraging the sparsity of the parameter tensor in response to determining that a sparsity parameter value representing the sparsity of a measure of the parameter tensor exceeds a threshold parameter value.
[0043] Based on a comparison of not only the opcode, but also any associated sparsity threshold, the controller 103 triggers a determination of whether the weight values of the parameter tensor have zero values or non - zero values. For example, the controller 103 can analyze the discrete weight values of the parameter tensor to detect non - zero weight values. In response to detecting a non - zero weight value (which may be denoted as K, for example), the controller 103 then uses that non - zero weight value to generate a set or grouping of compressed parameters.
[0044] In some embodiments, controller 103 extracts the detected non-zero weight values and uses the extracted weights to generate a set of compression parameters. In some other embodiments, instead of extracting the weight values, controller 103 associates the detected non-zero weight values with a set of compression parameters, which is, for example, a set of compression parameters previously generated by host 120 and passed to controller 103 at compute tile 101 via an exemplary host interface. Also, controller 103 can use a combination of extraction and association to generate a grouping of compression parameters.
[0045] Controller 103 maps each detected non-zero weight value to mapping vectors 302, 402. Mapping vectors 302, 402 may be represented as bit vectors, bitmaps, or other related data structures for indicating correlations or mappings between discrete data items. In some embodiments, controller 103 determines the mapping of the mapping vectors with reference to corresponding input vectors 304, 404. For example, mapping vectors 302, 402 map the discrete inputs of input vectors 304, 404 to non-zero values of sets of compressed sparse parameters 305, 405, respectively.
[0046] In some embodiments, the mapping vector is a non-zero bitmap identified as parameter NZM. An exemplary CSP can correspond to a modified parameter tensor derived for the original, unmodified parameter tensor, and the mapping is configured to generate the modified parameter tensor after holding the respective dimensional positions of each non-zero element within the original, unmodified parameter tensor. For example, the mapping vector can have the same dimensions as the original matrix for which the mapping vector is determined, but the mapping vector has a 1-bit data type that is set to "1" for non-zero elements (e.g., non-zero weight values) within that position of the original matrix, or set to "0" for zero elements (e.g., zero weight values) within that position of the original matrix.
[0047] The input vectors 304, and the individual inputs of 304 can be represented as {a0, a1, a2, a3, aN}, and the individual weight values of the parameter tensors can be represented as {w0, w1, w2, w3, wN}. The mapping vectors 302, 402 use control values, such as binary values, to map the individual inputs (e.g., a0, a1, a2, etc.) of the input vectors 304, 404 to the non-zero weights within the sets 305, 405 of compressed sparse parameters. The compute tile 101 includes selection logic 314 for selecting the individual inputs of the input vectors 304, 404 with reference to the non-zero weights of the sets 305, 405 of compressed sparse parameters. The selection logic 314 aligns the extraction of the inputs at the input vectors with the corresponding non-zero weight values in the sets of compressed sparse parameters with reference to the mapping vectors.
[0048] In some embodiments, the selection logic 314 is implemented in hardware, software, or both. For example, the controller 103 can access hardware selection logic 314 coupled to a first memory 108 and a second memory 110 of a hardware accelerator. The controller 103 can use the selection logic 314 to select a particular input of the input vector and the corresponding weight value within the set of compressed sparse parameters based on each bit value of the bits within the mapping vector.
[0049] Controller 103 can generate a mapping vector that maps individual inputs of a multi-dimensional (3D) input tensor to non-zero weights within a multi-dimensional (3D) compressed sparse parameter tensor. In some embodiments, for a given multi-dimensional tensor, computational tile 101 is configured such that different cells or groups of cells within computational unit 112 operate on different columns or dimensions of the parameter tensor / weight matrix. Thus, computational tile 101 can generate a different bitmap or mapping vector for each cell, or for each grouping of cells. In this way, each computational tile 101 can include respective selection logic that is uniquely configured for each cell, for each grouping of cells, or for both.
[0050] Controller 103 generates control signals for storing a set of compressed sparse parameters and a corresponding mapping vector at a memory location of computational tile 101. For example, memory 110 can include SIMD (single instruction, multiple data) registers, each configured to store mapping vectors 302, 402 at corresponding first addresses of SIMD register 310, respectively. Similarly, the SIMD register can also store a set of compressed sparse parameters 305, 405 at different corresponding second addresses of SIMD register 310.
[0051] SIMD register 310 can include parallel cells that compute multiple dot products in parallel. The dot product computations can be performed in numerical formats or data types (dt) such as INT8, BFLOAT, HALF, and INT16. In some embodiments, the data type of the computation is specified or indicated using one or more data fields of the opcode and / or instruction received at computational tile 101.
[0052] Each computing cell can use 4B for each operand, that is, 4 INT8 elements for each operand, or 2 BFLOAT / HALF / INT16 elements for each operand. One operand is the input, and this input is read from an exemplary scratchpad memory (described below) and broadcast to the entire computing cell. The broadcast function is described below with reference to FIG. 5. The second operand is a weight value read from the SIMD register 310. In some embodiments, the weight values are different for each cell.
[0053] As shown in Table 1 below, the techniques and architectures described herein can be used to accelerate computations for various combinations of K-out-of-N sparsity and data types (dt).
Table 1
[0054] Each cell can execute a multiply and accumulate function. The accumulate function is executed on the partial sums generated from the multiplication operation. The accumulate function can be expressed as: (MAC(operand1(4B), operand4(4B), partial_sum(32B) -> partial_sum’(32B)). In some embodiments, each computing cell or MAC includes a cell accumulator, and the partial sums are stored in the cell accumulator as partial results.
[0055] In the examples of FIGS. 3 and 4, the memory 108 can be implemented as scratchpad memories 308 and 408, respectively. In some embodiments, one or more memory structures in the compute tile 101 can be implemented as a scratchpad memory, such as a shared scratchpad memory. Instead of simply overwriting data, the memory structure can be configured to support a direct memory access (DMA) mode that atomically reduces incoming vector data to a memory location, for example, to support atomic floating-point reduction. In some embodiments, the resources of the shared memories 308 and 408 are configured to function as a software-controlled scratchpad rather than, for example, a hardware-managed cache.
[0056] In each of FIGS. 3 and 4, Boolean “0” corresponds to a detected zero weight value and Boolean “1” corresponds to a detected non-zero weight value. In the example of FIG. 3, the mapping vector 302 is a 4-bit vector, and in the example of FIG. 4, the mapping vector 402 is an 8-bit vector. Each of the mapping vectors 302 and 402 can also have more or fewer bits. For example, each of the mapping vectors 302 and 402 can vary in size (or bits) based on design preferences, processing constraints, system configuration, or combinations thereof.
[0057] In some embodiments, one or more of the mapping vectors 302 and 402 enable rationalization of the processing of neural network operands of a machine learning dataset in the compute tile 101. To optimize the degree to which neural network operations can be rationalized over conventional approaches, the system 100 utilizes a specific hardware architecture of an application-specific integrated circuit that uses each pairing of a broadcast input bus coupled to a grouping of compute cells across multiple compute tiles to accelerate the execution of artificial neural networks. This is described in detail below with reference to the example of FIG. 5.
[0058] FIG. 5 shows an example of a processing pipeline 500 for routing inputs obtained from memory locations of memory 108 to one or more compute cells 114. Pipeline 500 utilizes a hardware architecture in which input bus 106 is coupled (e.g., directly coupled) to each of a plurality of groupings of hardware compute cells of an application specific integrated circuit. System 100 provides inputs of an input feature map or activation functions (e.g., a0, a1, a2, etc.) to a subset of MACs 114.
[0059] For example, each input of input vector 102 is provided to each MAC within the subset via input bus 106 of compute tile 101. System 100 can perform this broadcast operation across multiple compute tiles 101 to compute the product of a given neural network layer using each grouping of inputs and corresponding weights at each compute tile 101. In a given compute tile 101, the product is computed by multiplying each input (e.g., a1) and corresponding weight (e.g., w1) at each MAC within the subset using the multiplication circuitry of the MAC.
[0060] System 100 can generate the output of the layer based on the accumulation of the plurality of respective products computed at each MAC 114 within the subset. As will be described below with reference to FIG. 7, the multiplication operations performed within compute tile 101 can include i) a first operand (e.g., an input or activation function) stored at a memory location of memory 108 corresponding to each element of the input tensor, and ii) a second operand (e.g., a weight) stored at a memory location of memory 110 corresponding to each element of the parameter tensor.
[0061] In the example of FIG. 5, the shift register 502 can provide a shift function, in which the input of the operand 504 is broadcast on the input bus 106 and routed to one or more MACs 114. In some embodiments, the shift register 502 enables one or more input broadcast modes to the compute tile 101. For example, the shift register 502 can be used to broadcast inputs sequentially (one by one) from the memory 108, in parallel from the memory 108, or using some combination of these broadcast modes. The shift register 502 can be an integrated function of the memory 108 and can be implemented in hardware, software, or both.
[0062] As shown, in one embodiment, the weight (w2) of the operand 506 may have a zero weight value. To conserve processing resources, if the controller 103 determines that the weight (w2) has a zero value, the multiplication between the input (a2) and the weight (w2) can be skipped, so that those operands are not routed to or consumed by the cells 114a / b / c. The decision to skip that particular multiplication operation can be based on a mapping vector that maps the discrete inputs (an) of the input vector to the individual weights (wn) of the compressed sparse parameters, as described above.
[0063] FIG. 6 is a flowchart of an example of a process 600 for leveraging data sparsity during the calculation of a neural network implemented on a hardware accelerator. In some examples, the calculation is performed using a dedicated hardware integrated circuit to process neural network inputs, such as images or voice utterances.
[0064] For example, a hardware integrated circuit can be configured to implement a CNN that includes a plurality of neural network layers. In some cases, the neural network layer can include a group of convolutional layers. The input can be an example of an image as described above, including various other types of digital images or related graphic data. In at least one example, the integrated circuit implements an RNN to process an input derived from speech or other audio content. In some embodiments, process 600 is part of a technique used to accelerate neural network computations and also enables an improvement in the accuracy of image or audio processing output compared to other data processing techniques.
[0065] Process 600 can be implemented or executed using system 100 described above. Accordingly, the description of process 600 may refer to the computing resources of system 100 described above. In some examples, the steps or actions of process 600 are enabled by programmed firmware instructions, software instructions, or both. Each type of instruction is executable by one or more processors of the devices and resources described in this document.
[0066] In some embodiments, the steps of process 600 are executed in a hardware circuit to generate a layer output of a neural network layer. The output can be part of the calculations for a machine learning task or inference workload to generate an image processing or image recognition output. The integrated circuit can be a dedicated neural network processor or a hardware machine learning accelerator configured to accelerate the calculations for generating various types of data processing outputs.
[0067] Referring again to process 600, the input vector is retrieved from the first memory of compute tile 101 of the hardware accelerator (602). For example, the controller 103 of compute tile 101 generates a control signal to retrieve the input vector from an address location in the first memory 102. In some embodiments, the input vector corresponds to an input feature map of an image and can be a matrix structure of neural network inputs, such as an activation function generated by a previous neural network layer.
[0068] Compute tile 101 processes an opcode indicative of the sparsity among data used in neural network computations (604). For example, the controller 103 can process the opcode in response to scanning, analyzing, or otherwise reading one or more instructions received at the compute tile. The opcode can indicate the sparsity of an exemplary parameter tensor that includes the weights of one or more neural network layers. The parameter tensor and its corresponding weight (parameter) values are stored in and accessed from the memory of the hardware accelerator, such as the second memory 110. For example, the parameter tensor can be an 8×8 matrix that includes one or more columns or rows having a sparsity of K out of N (where K = 2 and N = 4 here) (i.e., a sparsity of 2 out of 4).
[0069] As described below, the controller 103 reads the opcode and configures the compute tile 101 to perform computations on a neural network layer for leveraging the sparsity indicated by the opcode. The opcode can be determined in advance by a compiler (e.g., a hardware accelerator) of the integrated circuit that generates an instruction set for execution on one or more compute tiles of the integrated circuit. For example, the compiler can be operable to generate the instruction set in response to compiling source code for performing neural network computations on a machine learning workload such as an inference or training workload. The instruction set can include a single instruction that is broadcast to one or more compute tiles 101.
[0070] Each single instruction consumed by a given compute tile can specify an opcode that indicates the sparsity of the parameter tensor assigned to that compute tile. The opcode can also instruct the operations to be performed on the compute tile 101 to leverage the tensor sparsity. In some embodiments, the opcode instructs the compute tile 101 to generate and execute a localized instruction or control signal for leveraging the tensor sparsity based on the compressed sparse parameters and corresponding mapping vectors that are received at the compute tile, locally generated at the compute tile 101, or both.
[0071] In some embodiments, the opcode is determined dynamically at runtime, e.g., by a higher-level controller 125 of the hardware accelerator or by the local controller 103 of the compute tile 101. The opcode can be a unique operation code included among a plurality of opcodes that are broadcast by the higher-level controller to a plurality of compute tiles of the hardware accelerator. As shown above, in some embodiments, the plurality of opcodes are broadcast in a single instruction that is provided to each of the plurality of tiles 101 of the hardware accelerator.
[0072] The set of compressed sparse parameters is derived from a parameter tensor obtained based on an opcode (606). The opcode in a single instruction received at compute tile 101 can indicate a sparsity of two quarters with respect to a given column or row dimension of one or more parameter tensors. For example, the opcode can indicate that multiple columns 202 of parameter tensor 200 have a pattern (e.g., a sparsity pattern), and along a given dimension, for every four elements, zero-valued weights are assigned to two of the four elements, and non-zero-valued weights are assigned to two of the four elements. The four elements can be {w0, w1, w2, w4}, and weights w1, w4 have non-zero weight values. Based on this, controller 103 can detect these non-zero weight values and then access (or generate) the set of compressed sparse parameters represented as {w1, w4}.
[0073] The mapping vector is generated based on the set of compressed sparse parameters and an opcode (608). In some embodiments, the mapping vector is generated outside compute tile 101 and then provided to and stored in compute tile 101. Based on the parameter tensor including the four elements {w0, w1, w2, w4} and the non-zero weight values of weights w1, w4, controller 125 can then generate the mapping vector with reference to the input vector for which the weights of the four elements {w0, w1, w2, w4} need to be processed. The mapping vector can be based on an encoding scheme indicating the positions of the zero and non-zero weight values of the corresponding parameter tensor. In some other embodiments, controller 103 can perform operations to locally generate the mapping vector at a given compute tile 101.
[0074] The opcode can include one or more fields within an instruction (e.g., a single instruction) received at the compute tile 101. In some embodiments, the opcode can include respective values for a first data item K and a different second data item N. The compute tile 101 is operable to support various options for implementing sparse computations as at least indicated by the data values of K and N. For example, in an instruction received at the compute tile 101, the compute tile 101 enumerates through these options based on the data values of the opcode, including the value(s) of one or more fields of the opcode.
[0075] For example, an opcode having a data value N = 8 and specifying a data type of int8(1B) for integer data, for example, notifies the controller 103 that a particular sparse computation requires reading or fetching 1B of activation function from the memory 108. Similarly, the opcode can instruct the controller 103 to perform a parameter read from the SIMD register 310. For example, the controller 103 fetches N bits from the non-zero mapping vector (e.g., NZM) 302 and K * 1B elements from the CSP 305, activates appropriate selection logic, and performs the sparse computation. Fetching the required NZM bits and corresponding CSP elements for the parameters involves reading respective address spaces within the SIMD register 310 that store some (or all) of these data items. * Exemplary operations can be described with reference to FIG. 3. Considering arithmetic operations on an uncompressed tensor from the perspective of "cell 0" within the compute unit 112, the compute tile 101 performs a0
[0076] w0 + a1 * w0 + a1 * w1 + a2 * w2 + a3 *Required to execute an example of the dot product of w3, where {a0, a1, a2, a3} is an example of an activation function vector and {w0, w1, w2, w3} is an example of a weight vector corresponding to a sparse parameter tensor. Computation tile 101 analyzes each value of the weights within the weight vector and generates a mapping vector based on those values.
[0077] For example, computation tile 101 can access, obtain, or otherwise generate a mapping vector corresponding to the weights {w0, w1, w2, w3} using the bitmap
[0101] , where each bit of the bitmap corresponds to each value of the weights within the weight vector. In this example, the bitmap
[0101] indicates that each of the weights w0 and w2 has a zero value. Also, computation tile 101 can access, obtain, or otherwise generate a compressed sparse parameter based on each value of the weights within the weight vector. For example, controller 125 can derive a set of compressed parameters {w1, w3} from the weights {w0, w1, w2, w3}, indicating that each of the weights w1 and w3 has a non-zero value. Alternatively, or in addition, computation tile 101 can also locally derive a set of compressed parameters based on the operations performed by its local controller 103.
[0078] In some embodiments, computation tile 101 obtains or generates a compressed sparse parameter CSP0 and associates a set of compressed parameters {w1, w3} with the compressed sparse parameter CSP0. Each of the mapping vector
[0101] and CSP0 can be stored and then accessed from respective address locations in SIMD register 310 of memory 110. For example, an exemplary mapping vector {0101} is stored at the NZM address of SIMD register 310, and the parameter CSP0 is stored at the CSP address of SIMD register 310.
[0079] Using the mapping vector
[0101] and CSP0{w1, w3}, calculation tile 101 can perform a dot product calculation at cell 0. For example, calculation tile 101 can perform the dot product calculation, a1 * w1 + a3 * w3. In some embodiments, to rationalize this calculation, calculation tile 101 can automatically initialize the multiplier of cell 0 based on non-zero weight values associated with parameter CSP0. Next, calculation tile 101 can extract inputs a1 and a3 from the {a0, a1, a2, a3} activation function vector based on the bitmap of the mapping vector. For example, the selection logic 314 of calculation tile 101 refers to the mapping vector and aligns that extraction of inputs a1 and a3 with the corresponding non-zero weight values {w1, w3} of CSP0.
[0080] An input vector is processed through a layer of a neural network using a set of mapping vectors and compressed sparse parameters (610). In some embodiments, calculation tile 101 performs a dot product matrix multiplication operation to process the input vector through a layer of the neural network. Calculation tile 101 performs a dot product matrix multiplication operation between the input of the input vector and the corresponding weight values within the set of compressed sparse parameters.
[0081] For example, the operation includes multiplying weight values by an input or activation function value in one or more cycles to generate a plurality of products (e.g., partial sums), and then performing an accumulation of those products over many cycles. For each input, the dot product matrix multiplication operation can be performed on one or more multiplication cells of the hardware accelerator based on the respective bit values of each bit in the mapping vector.
[0082] In some embodiments, the compute tile 101 either alone or in cooperation with other compute tiles executes a convolution operation to process an input vector through a layer of a neural network. For example, the system 100 can execute a convolution in a neural network layer by using compute cells of a hardware integrated circuit to process an input (e.g., an input vector) of an input feature map. The cells can be hardware multiply-accumulate cells of a hardware compute unit in a hardware integrated circuit.
[0083] Furthermore, processing the input also includes providing weights of compressed sparse parameters to a subset of the multiply-accumulate cells of the hardware accelerator. In some embodiments, the controller 103 determines a mapping of the input of the input vector to the cells of the compute tile based on one or more opcodes within an instruction (e.g., a single instruction) received at the compute tile 101. The controller 103 generates a control signal 105 for providing the weights of the compressed sparse parameters to a subset of the multiply-accumulate cells based on the determined mapping. As described above, the weights of the compressed sparse parameters are provided from an exemplary SIMD register.
[0084] When executing a neural network computation with N = 4, K = 2, and dt = half / bfloat / int16 (2 bytes), the compute tile 101: i) reads 4 bits of a mapping vector stored at the NZM address, e.g., {0101}, and two compressed parameters (4B), e.g., {w1, w3}, stored at the CSP address; ii) reads four elements (8B) of an activation function operand from the scratchpad memory 308; iii) uses the 4 bits of the mapping vector, e.g., {0101}, to select the correct two compressed parameters (4B), e.g., {w1, w3}, as corresponding weight operands; iv) feeds the appropriate activation function operand and the corresponding weight operand to the compute cells to perform a multiply-accumulate operation; and v) can be configured to increment the relevant addresses in the memory to read the next activation function operand and weight operand from the memory.
[0085] When performing neural network calculations with N = 8, K = 4, and dt = int, calculation tile 101 reads: i) 8 bits of the mapping vector stored at the NZM address, e.g., {01011001}, and 4 compressed parameters (4B) stored at the CSP address, e.g., {w1, w3, w11, w7}; ii) 8 elements (8B) of the activation function operand from the scratch pad memory 408; iii) uses the 8 bits of the mapping vector, e.g., {01011001}, to select the correct 4 compressed parameters (4B), e.g., {w1, w3, w11, w7}, as the corresponding weight operands; iv) feeds the appropriate 4 operands (activation function and corresponding weights) to the calculation cell to perform a multiply-accumulate operation; and v) can be configured to increment the relevant address in the memory to read the next activation function and weight operands from the memory.
[0086] When performing neural network calculations with N = 8, K = 2, and dt = half / bfloat / int16, calculation tile 101 reads: i) 8 bits of the mapping vector stored at the NZM address, e.g., {01000001}, and 2 compressed parameters (4B) stored at the CSP address, e.g., {w1, w3}; ii) 8 elements (16B) of the activation function operand from the scratch pad memory 408; iii) uses the 8 bits of the mapping vector, e.g., {01000001}, to select the correct 2 compressed parameters (4B), e.g., {w1, w3}, as the corresponding weight operands; iv) feeds the appropriate 2 operands (activation function and corresponding weights) to the calculation cell to perform a multiply-accumulate operation; and v) can be configured to increment the relevant address in the memory to read the next activation function and weight operands from the memory.
[0087] When performing neural network calculations with N = 16, K = 4, and dt = int, calculation tile 101 reads: i) 16 bits of the mapping vector stored at the NZM address, e.g., {01000001 00001100}, and 4 compressed parameters (4B) stored at the CSP address, e.g., {w1, w3}; ii) 16 elements (32B) of the activation function operand from the scratchpad memory 408; iii) uses the 16 bits of the mapping vector, e.g., {01000001 00001100}, to select the correct 4 compressed parameters (4B), e.g., {w1, w3, w11, w7}, as the corresponding weight operands; iv) feeds the appropriate 4 operands (activation function and corresponding weights) to the calculation cell to perform a multiply-accumulate operation; and v) can be configured to increment the relevant address in the memory to read the next activation function and weight operands from the memory.
[0088] When performing neural network calculations with N = 16, K = 2, and dt = half / bfloat / int16, calculation tile 101 reads: i) 16 bits of the mapping vector stored at the NZM address, e.g., {01000000 00000100}, and 2 compressed parameters (4B) stored at the CSP address, e.g., {w1, w3}; ii) 16 elements (32B) of the activation function operand from the scratchpad memory 408; iii) uses the 16 bits of the mapping vector, e.g., {01000001 00001100}, to select the correct 2 compressed parameters (4B), e.g., {w1, w3}, as the corresponding weight operands; iv) feeds the appropriate 2 operands (activation function and corresponding weights) to the calculation cell to perform a multiply-accumulate operation; and v) can be configured to increment the relevant address in the memory to read the next activation function and weight operands from the memory.
[0089] When performing neural network calculations with N = 32, K = 4, and dt = int, calculation tile 101 reads i) 32 bits of the mapping vector stored at the NZM address and 4 compression parameters (4B) stored at the CSP address, ii) 32 elements (32B) of the activation function operand from the scratch pad memory 408, iii) uses the 32 bits of the mapping vector to select the correct 4 compression parameters (4B) as the corresponding weight operand, iv) feeds the appropriate 4 operands (activation function and corresponding weights) to the calculation cell to perform a multiply-accumulate operation, and v) can be configured to increment the relevant addresses in memory to read the next activation function and weight operands from memory.
[0090] FIG. 7 shows an example of a tensor or multi-dimensional matrix 700 that includes an input tensor 704, a deformed form of the parameter tensor 706, and an output tensor 708. In the example of FIG. 7, each tensor 700 includes respective elements, and each element can correspond to respective data values (or operands) for calculations performed at a given layer of the neural network.
[0091] For example, each input of the input tensor 704 can correspond to respective elements along a given dimension of the input tensor 704, each weight of the parameter tensor 706 can correspond to respective elements along a given dimension of the parameter tensor 706, and each output value or activation function within the set of outputs can correspond to respective elements along a given dimension of the output tensor 708. In relation, each element can correspond to respective memory locations or addresses within the memory of the calculation tile 101 that are assigned to operate in one or more dimensions of a given tensor 704, 706, 708.
[0092] The computation performed on a given neural network layer can include multiplying an input / activation function tensor 704 and a parameter / weight tensor 706 in one or more processor clock cycles to generate a layer output, which may include an output activation function. Multiplying the activation function tensor 704 by the weight tensor 706 includes multiplying the weights from the elements of the tensor 706 by the activation function from the elements of the tensor 704 to generate one or more partial sums. The exemplary tensor 706 of FIG. 7 can be an unmodified parameter tensor, a modified parameter tensor, or a combination thereof. In some embodiments, each parameter tensor 706 corresponds to a modified parameter tensor that includes non-zero CSP values derived based on the sparsity exploitation techniques described above.
[0093] The processor core of the system 100 can operate on i) a scalar corresponding to a discrete element within a multi-dimensional tensor 704, 706, ii) a vector of values including a plurality of discrete elements 707 along the same or different dimensions of a multi-dimensional tensor 704, 706 (e.g., an input vector 102), or iii) a combination thereof. In a multi-dimensional tensor, each discrete element 707, or each of the plurality of discrete elements 707, can be represented using X, Y coordinates (2D) or X, Y, Z coordinates (3D), depending on the number of dimensions of the tensor.
[0094] System 100 can calculate a plurality of partial sums corresponding to the products generated by multiplying with weight values that support batch input. As described above, System 100 can perform the accumulation of products (e.g., partial sums) over many clock cycles. For example, the accumulation of products can be performed in the random access memory, shared memory, or scratchpad memory of one or more computing tiles based on the techniques described in this document. In some embodiments, the input weight multiplication can be written as the sum of the products of each weight element multiplied by discrete inputs of the input vector 102, such as a row or slice of the input tensor 704. This row or slice can represent a given dimension, such as the first dimension 710 of the input tensor 704 or a different second dimension 715 of the input tensor 704.
[0095] In some embodiments, an exemplary set of calculations can be used to calculate the output of a convolutional neural network layer. The calculations of the CNN layer can involve performing a 2D spatial convolution between a 3D input tensor 704 and at least one 3D filter (weight tensor 706). For example, convolving one 3D filter 706 over the 3D input tensor 704 can generate 2D spatial planes 720 or 725. These calculations can involve calculating the sum of dot products for a particular dimension of the input volume that includes the input vector 102.
[0096] For example, the spatial plane 720 can include the output values of the sum of products calculated from the inputs along the dimension 710, and the spatial plane 725 can include the output values of the sum of products calculated from the inputs along the dimension 715. The calculations for generating the sum of products of the output values in each of the spatial planes 720 and 725 can be performed i) in the computing cells 114a / b / c, ii) directly in the memory 110 using arithmetic operators coupled to the shared bank of the memory 110, iii) or both. In some embodiments, the reduction operation can be rationalized or performed directly in the memory cells (or locations) of the memory 110 using various techniques for reducing the accumulated values.
[0097] Embodiments of the subject matter and the functional operations described in this specification can be implemented in digital electronic circuitry, tangibly embodied computer software or firmware, computer hardware, such as the structures disclosed in this specification and their structural equivalents, or combinations of one or more of them. Embodiments of the subject matter described in this specification can be implemented as one or more computer programs, i.e., as one or more modules of computer program instructions, encoded on a tangible non-transitory program carrier for execution by, or to control the operation of, a data processing apparatus.
[0098] Alternatively, or in addition, the program instructions can be encoded in an artificially generated propagated signal, e.g., a mechanically generated electrical, optical, or electromagnetic signal, generated to encode information for transmission to a receiver device suitable for execution by a data processing apparatus. A computer storage medium may be, or include, a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or a combination of one or more of them.
[0099] The term "computing system" includes, by way of example, all kinds of devices, devices, and machines for processing data, including programmable processors, computers, or multiple processors or computers. The device can include dedicated logic circuitry, e.g., an FPGA (Field Programmable Gate Array) or an ASIC (Application Specific Integrated Circuit). The device can also include, in addition to hardware, code that creates an execution environment for the computer program in question, e.g., processor firmware, protocol stack, database management system, operating system, or code constituting a combination of one or more of them.
[0100] A computer program (which may be called or described as a program, software, software application, module, software module, script, or code) can be written in any form of programming language, including a compiled or interpreted language, or a declarative or procedural language, and it can be deployed in any form, including as a stand-alone program or in the form of a module, component, subroutine, or other unit suitable for use in a computing environment.
[0101] A computer program may correspond to a file in a file system, but it does not necessarily have to. The program can be stored as part of a file that holds other programs or data (e.g., one or more scripts stored in a document in a markup language), or in a single file dedicated to the program in question, or in multiple related files (e.g., files that hold one or more modules, subprograms, or portions of code). A computer program can be deployed to be executed on one computer or on multiple computers located at one or more locations and interconnected by a communication network.
[0102] The processes and logical flows described herein can be executed by one or more programmable computers executing one or more computer programs to perform functions by operating on input data and generating output. The process and logical flows can also be executed by special-purpose logic circuits, such as FPGAs (field-programmable gate arrays), or ASICs (application-specific integrated circuits), or GPGPUs (general-purpose graphics processing units), and the apparatus can also be implemented as those special-purpose logic circuits.
[0103] A computer suitable for the execution of a computer program includes a general-purpose or special-purpose microprocessor or both, or any other kind of central processing unit, and can be based on them, for example. Generally, the central processing unit receives instructions and data from read-only memory, random access memory, or both. Some elements of the computer are a central processing unit for executing and running instructions and one or more memory devices for storing instructions and data. Generally, the computer also includes one or more mass storage devices for storing data, such as magnetic disks, magneto-optical disks, or optical disks, or is operably connected to receive data from them, transmit data to them, or do both. However, the computer does not necessarily have such devices. Further, to give some examples, the computer can be embedded in other devices, such as mobile phones, personal digital assistants (PDAs), mobile audio or video players, game consoles, global positioning system (GPS) receivers, or portable storage devices, such as universal serial bus (USB) flash drives.
[0104] Computer-readable media suitable for storing computer program instructions and data include, by way of example, semiconductor memory devices (such as EPROM, EEPROM, and flash memory devices), magnetic disks (such as internal hard disks or removable disks), magneto-optical disks, and all forms of non-volatile memory, media, and memory devices such as CD-ROM and DVD-ROM disks. The processor and memory can be complemented by, or incorporated in, dedicated logic circuitry.
[0105] To interact with a user, embodiments of the subject matter described herein can be implemented on a computer having a display device for displaying information to the user, such as an LCD (liquid crystal display) monitor, and a keyboard and a pointing device, such as a mouse or a trackball, by which the user can input to the computer. Other types of devices can also be used to provide interaction with the user. For example, the feedback provided to the user may be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback, and the input from the user can be received in any form, such as acoustic input, voice input, or tactile input. Further, the computer can interact with the user by sending and receiving documents to and from the devices used by the user, for example, by sending a web page to a web browser on the user's client device in response to a request received from a web browser.
[0106] Embodiments of the subject matter described herein can be implemented in a computing system that includes, for example, a backend component as a data server, or a middleware component, such as an application server, or a frontend component, such as a client computer having a graphical user interface or a web browser by which the user can interact with embodiments of the subject matter described herein, or a computing system that includes any combination of one or more such backend, middleware, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication, such as, for example, a communication network. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), such as the Internet.
[0107] A computing system can include a client and a server. The client and the server are generally far apart from each other and typically communicate through a communication network. The client-server relationship is created by computer programs that operate on respective computers and have a client-server relationship with each other.
[0108] Although this specification contains many details of specific embodiments, these should not be construed as limiting the scope of any invention or what can be claimed, but rather as an explanation of features that may be specific to particular embodiments of a particular invention. Specific features described in the context of individual embodiments herein can also be implemented in combination within a single embodiment. Conversely, the various features of the present invention described in the context of a single embodiment can also be implemented separately or in any suitable sub-combination in multiple embodiments. Further, even if a feature has been described above as functioning in a particular combination and was initially claimed as such, one or more features from the claimed combination may in some cases be deleted from the combination, and the claimed combination may be directed to a sub-combination or variation of a sub-combination.
[0109] Similarly, although operations are shown in the drawings in a particular order, this should not be understood as requiring that such operations be performed in the particular order or sequence shown, or that all of the operations shown be performed, in order to obtain a desirable result. In certain circumstances, multitasking and parallel processing may be advantageous. Further, the separation of the various system modules and components in the above embodiments should not be understood as requiring such separation in all embodiments, and the described program components and systems can generally be integrated into a single software product or packaged into multiple software products.
[0110] Particular embodiments of the subject matter have been described. Other embodiments are within the scope of the following claims. For example, even if the actions recited in the claims are performed in a different order, desirable results can still be obtained. As an example, the processes illustrated in the accompanying figures do not necessarily require to be in the particular order shown, or in sequential order, to obtain desirable results. In certain embodiments, multitasking and parallel processing may be advantageous.
Claims
1. A method executed by a computer, comprising a neural network implemented on a hardware accelerator, the method comprising: Deriving a set of compressed sparse parameters from a parameter tensor; Generating a mapping vector based on the set of compressed sparse parameters; Processing instructions for instructing sparse computations to be performed using the compressed sparse parameters based on the sparsity of the parameter tensor; Based on the instructions, i) obtaining an input vector from a first memory of the hardware accelerator, and ii) obtaining the compressed sparse parameters from a second memory of the hardware accelerator; Performing the sparse computation to process the input vector through a layer of the neural network using the mapping vector and the set of compressed sparse parameters. A method comprising the above steps.
2. The method according to claim 1, wherein processing the input vector through the layer of the neural network includes performing a dot product matrix multiplication operation between the input of the input vector and corresponding weight values within the set of compressed sparse parameters.
3. The method according to claim 2, wherein the dot product matrix multiplication operation is performed on one or more multiplication cells of the hardware accelerator based on respective bit values of each bit within the mapping vector.
4. Accessing hardware selection logic coupled to the first memory and the second memory of the hardware accelerator; Using the hardware selection logic to select a specific input of the input vector and corresponding weight values within the set of compressed sparse parameters based on respective bit values of bits within the mapping vector. The method according to claim 2, further comprising the above steps.
5. Deriving the set of compressed sparse parameters includes: Generating a modified parameter tensor that includes only non-zero elements along a specific dimension of the parameter tensor. The method according to claim 1.
6. Generating the modified parameter tensor includes: In the case of a specific column dimension of the parameter tensor, Generating a compressed representation of the column dimension based on non-zero elements of the column dimension. Concatenating each non-zero element within the compressed representation of the column dimension; The method according to claim 5, comprising: **Claim 7** The method according to claim 6, wherein generating the modified parameter tensor includes retaining the respective dimensional positions of each non-zero element within the parameter tensor prior to generating the modified parameter tensor. **Claim 8** The parameter tensor includes a plurality of dimensions; The method according to claim 1, wherein the opcode within the instruction indicates sparsity of a specific one of the plurality of dimensions. **Claim 9** The hardware accelerator is operable to process a plurality of multi-dimensional parameter tensors; The method according to claim 1, wherein the opcode within the instruction indicates uniform sparsity across each of the plurality of multi-dimensional parameter tensors. **Claim 10** The first memory is a scratchpad memory of the hardware accelerator and is configured to store inputs and activation functions processed by the neural network layer. The method according to claim 1. **Claim 11** The second memory includes a plurality of SIMD (single instruction, multiple data) registers; The method includes storing the mapping vector at a first address of the SIMD register; and storing the set of compressed sparse parameters at a different second address of the SIMD register. The method according to claim 10, comprising: **Claim 12** A hardware accelerator configured to implement a neural network, a processing device, and a non-transitory machine-readable storage medium for storing instructions, wherein the instructions are executable by the processing device to derive a set of compressed sparse parameters from a parameter tensor; generate a mapping vector based on the set of compressed sparse parameters; process instructions that direct sparse computations to be performed using the compressed sparse parameters based on sparsity of the parameter tensor; and acquire, based on the instructions, i) an input vector from a first memory of the hardware accelerator and ii) the compressed sparse parameters from a second memory of the hardware accelerator. Performing the sparse calculation to process the input vector through the layer of the neural network using the set of the mapping vector and the compressed sparse parameters; A system that causes the execution of operations including.
13. Processing the input vector through the layer of the neural network includes performing a dot product matrix multiplication operation between the input of the input vector and the corresponding weight values within the set of the compressed sparse parameters. The system according to claim 12.
14. The dot product matrix multiplication operation is executed on one or more multiplication cells of the hardware accelerator based on the respective bit values of each bit within the mapping vector. The system according to claim 13.
15. The operations include accessing the hardware selection logic coupled to the first memory and the second memory of the hardware accelerator; using the hardware selection logic to select a specific input of the input vector and the corresponding weight values within the set of the compressed sparse parameters based on the respective bit values of the bits within the mapping vector; The system according to claim 13, further including.
16. Deriving the set of the compressed sparse parameters includes generating a modified parameter tensor including only non-zero elements along a specific dimension of the parameter tensor. The system according to claim 12.
17. Generating the modified parameter tensor includes in the case of a specific column dimension of the parameter tensor, generating a compressed representation of the column dimension based on the non-zero elements of the column dimension; concatenating each non-zero element within the compressed representation of the column dimension; The system according to claim 16, including.
18. Generating the modified parameter tensor includes holding the respective dimensional positions of each non-zero element within the parameter tensor before generating the modified parameter tensor. The system according to claim 17.
19. The parameter tensor includes a plurality of dimensions, The opcode within the instruction indicates the sparsity of a specific dimension among the plurality of dimensions. The system according to claim 12.
20. The hardware accelerator is operable to process a plurality of multi-dimensional parameter tensors, The system according to claim 12, wherein the opcode in the instruction exhibits uniform sparsity across each of the plurality of multi-dimensional parameter tensors.
21. The first memory is The scratchpad memory of the hardware accelerator, The system according to claim 12, configured to store inputs and activation functions processed by the neural network layer.
22. The second memory includes a plurality of SIMD (single instruction, multiple data) registers, The operation is Storing the mapping vector at a first address of the SIMD register, Storing the set of compressed sparse parameters at a different second address of the SIMD register, The system according to claim 21, further comprising.
23. A non-transitory machine-readable storage medium for storing instructions executable by a processing device of a hardware accelerator configured to implement a neural network, The execution of the instructions is Deriving a set of compressed sparse parameters from a parameter tensor, Generating a mapping vector based on the set of compressed sparse parameters, Processing instructions that direct sparse calculations to be performed using the compressed sparse parameters based on the sparsity of the parameter tensor, Based on the instructions, i) obtaining an input vector from a first memory of the hardware accelerator, and ii) obtaining the compressed sparse parameters from a second memory of the hardware accelerator, Executing the sparse calculation to process the input vector through the layer of the neural network using the mapping vector and the set of compressed sparse parameters, A non-transitory machine-readable storage medium that causes the execution of operations including.
24. A method of being executed using a hardware accelerator implementing a neural network including a plurality of neural network layers, Receiving instructions for a compute tile of the hardware accelerator, The instructions are Executable on the compute tile, Identifying an opcode within the instruction that indicates sparsity of the parameter tensor; Loading a set of compressed sparse parameters based on weight values derived from a parameter tensor that specifies a plurality of weights for a layer of the neural network; Loading a mapping vector generated based on the set of compressed sparse parameters; Based on the opcode, obtaining (i) an input vector from a first memory of the hardware accelerator and (ii) the set of compressed sparse parameters from a second memory of the hardware accelerator; Based on the mapping vector, using the set of compressed sparse parameters to process the input vector through the layer of the neural network; A method that causes execution of operations including the above.
Citation Information
Patent Citations
Neural network accelerator
US20210004668A1