Power Reduction of Machine Learning Accelerator
By identifying matrix tiles and selecting appropriate multiplication paths based on range information, the technique reduces power consumption in neural network operations, addressing the high power usage associated with large-scale matrix multiplication.
Patent Information
- Application Number
- JP2022554763
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-03-26
- Filing Date
- 2021-03-08
- Publication Date
- 2025-06-30
- Estimated Expiration
- 2041-03-08
AI Technical Summary
Large-scale matrix multiplication operations in neural networks consume a significant amount of power due to the complexity and number of floating-point multiplication operations performed.
The technique involves identifying matrix tiles, obtaining range information for each tile, selecting a matrix multiplication path based on this information, and performing matrix multiplication using the selected path to generate a tile matrix multiplication product.
This approach reduces power consumption by selecting simpler multiplication paths for more restricted ranges of tile values, thereby minimizing the complexity of the circuitry required for matrix multiplication.
Smart Images

Figure 0007700142000006 
Figure 0007700142000007 
Figure 0007700142000008
Abstract
Description
Technical Field
[0001] (Cross - Reference to Related Applications) This application claims the benefit of U.S. Patent Application No. 16 / 831,711, filed on March 26, 2020, the content of which is incorporated herein by reference.
Background Art
[0002] Machine learning systems process inputs through a trained network to generate outputs. Due to the amount of data being processed and the complexity of the network, such evaluations involve a very large number of calculations.
[0003] A more detailed understanding can be obtained from the following description given by way of example together with the accompanying drawings.
Brief Description of the Drawings
[0004]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Best Mode for Carrying Out the Invention
[0005] Techniques for performing neural network operations are disclosed. This technique includes identifying a first matrix tile and a second matrix tile, obtaining first range information for the first matrix tile and second range information for the second matrix tile, selecting a matrix multiplication path based on the first range information and the second range information, and performing matrix multiplication on the first matrix tile and the second matrix tile using the selected matrix multiplication path to generate a tile matrix multiplication product.
[0006] FIG. 1 is a block diagram of a neural network processing system 100 according to an example. The neural network processing system includes a neural network processing block 102 and neural network data 104. The neural network processing block 102 is embodied as a hardware circuit that executes the operations described herein, software that executes on a processor to execute the operations described herein, or a combination of a hardware circuit that executes the operations described herein and software that executes on a processor.
[0007] In the operation, the neural network processing block 102 receives a neural network input 106, processes the neural network input 106 according to the neural network data 104 to generate a neural network output 108, and outputs the neural network output 108.
[0008] In some examples, the neural network processing block 102 is or includes a computer system having one or more processors that read and execute instructions for performing the operations described herein. In some embodiments, any such processor (or any processor described herein) includes an instruction fetch circuit for fetching instructions from one or more memories, a data fetch circuit for fetching data from one or more memories, and an instruction execution circuit for executing instructions. In various examples, one or more processors of the neural network processing block 102 are coupled to one or more input devices and / or one or more output devices that input data and output data for the one or more processors. The neural network data 104 includes data that defines one or more neural networks, and the neural network processing block 102 processes the neural network input 106 through the one or more neural networks to generate a neural network output 108.
[0009] FIG. 2 is an exemplary block diagram showing neural network data 104. The neural network data 104 includes a sequence of layers 202 through which data flows. The neural network data 104 may sometimes be referred to herein simply as "neural network 104" because this data represents a sequence of neural network operations performed on an input to generate an output. The neural network processing block 102 applies a neural network input 106 to the layers 202, and the layers 202 apply respective layer transformations to generate a neural network output 108. Each layer has its own layer transformation applied to the input received by that layer 202 to generate an output as the neural network output 108 from that layer 202 to the next layer or for the final layer 202(N). The neural network data 104 defines the neural network in terms of the number of layers 202 and the specific transformations at each layer 202. Exemplary transformations include a general neuron layer in which each of a plurality of neurons in layer 202 defines connectivity to the output from the previous layer 202, a single-element transformation, a convolutional layer, and a pooling layer. More specifically, as described above, each layer 202 receives an input vector from the previous layer 202. Some layers 202 include a set of neurons, and each such neuron receives a defined subset or the entire vector of the input vector. Further, each such neuron has weights applied to each such input. Further, the activation of each neuron is the sum of the products of the input values at each input and the weights at each input (thus, each such activation is the dot product of the input vector of that neuron and the weight vector of that neuron).
[0010] The layer 202 that applies single-element transformation receives an input vector and applies some defined transformations to each element of the input vector. Exemplary transformations include a clamping function or some other non-linear function. The layer 202 that applies pooling generates an output vector of a size smaller than the input vector by performing downsampling on the input vector based on a downsampling function that downsamples the input in any technically feasible way. The layer 202 that applies convolution applies a convolution operation where dot product is applied to crop the input data and filter the filter vector to generate an output.
[0011] Some types of layer operations, such as general neuron layers and convolutional layers, are implemented by matrix multiplication. More specifically, since the calculation of the activation function of neurons in a general neuron layer is a dot product, such calculations can be implemented as a set of dot product operations defined by matrix multiplication. Similarly, since the application of filters in a convolution operation is performed by dot product, matrix multiplication operations can be used to implement convolutional layers. Large-scale matrix multiplication operations involving floating-point numbers may consume a large amount of power due to the complexity and number of floating-point multiplication operations performed. Therefore, techniques for reducing power consumption in specific situations are provided herein.
[0012] FIG. 3 is a block diagram of the neural network processing block 102 of FIG. 1 showing additional details by way of example. The neural network processing block 102 includes a tiled matrix multiplier 302 used to perform matrix multiplication for the layer 202 where the neural network processing block 102 uses matrix multiplication.
[0013] In the process of performing matrix multiplication on layer 202, neural network processing block 102 receives layer input 308 and layer weights 309, and generates or receives range metadata for layer input 310 and range metadata for weights 316. Layer input 308 includes inputs for a specific layer 202 that uses matrix multiplication. Layer weights 309 include neuron connection weights for a general neuron layer or filter weights for a convolutional layer. Layer input 308 includes a set of layer input tiles 312, each of which is part of an input matrix representing a layer input. Layer weights 309 are a set of layer weights divided into weight tiles 313. Range metadata for weights 316 includes range metadata for each weight tile 318. Each item of range metadata indicates the range of the corresponding weight tile 313. Range metadata for layer input 310 includes range metadata for each layer input tile 312. Each item of layer input metadata indicates the range of the corresponding layer input tile 312.
[0014] Ranges (weight range 318 and input range 311) indicate the range of values for the corresponding weight tile 313 or input tile 312. In one example, the range for a particular tile is -1 to 1, meaning that all elements of the tile are between -1 and 1. In another example, the range is -256 to 256, and in another example, the range is the full range (i.e., the maximum range that can be represented by the data items of the weights).
[0015] When performing matrix multiplication of layer weights 309 by layer input 308, tile matrix multiplier 302 performs matrix multiplication of layer input tiles 312 by layer weight tiles 313 to generate partial matrix products, and combines the partial matrix products to generate layer output 320. The specific layer input tiles 312 and weight tiles 313 that are multiplied together to generate the partial products, and the way in which those partial products are combined to generate layer output 320, are defined by the nature of the layer. Some examples are shown in other parts of this description.
[0016] In performing a particular multiplication of the layer input tile 312 by the weight tile 313, the tile matrix multiplier examines the range metadata for the weight tile 318 and the range metadata for the input tile 311 and selects a multiplication path 306 to perform the multiplication. Different multiplication paths 306 are configured for different combinations of ranges, where the combinations are defined as the range of the layer input tile 311 and the range of the weight tile 318. A multiplication path 306 configured for a more restricted combination of ranges consumes less power than a multiplication path 306 configured for a broader set of combinations. The multiplication path 306 is a circuit configured to perform matrix multiplication on two matrices of a maximum fixed size. Using the multiplication path 306 that employs the tiled multiplication approach described elsewhere in this specification, it is possible to multiply together two matrices larger than this size. Briefly, this tiled multiplication approach involves dividing the input matrix into tiles, multiplying these tiles together to generate partial products, and summing the partial products to generate the final output matrix. In some embodiments, each multiplication path 306 is configured for a multiplicand matrix of the same size.
[0017] For a more limited range, power reduction for the multiplication path 306 is achieved through a simpler circuit. In one example, matrix multiplication involves performing a dot product that includes multiplying the dot product multiplicand to produce a partial dot product, and summing the partial dot products to produce the final dot product. The exponent of the partial dot product is ultimately small enough such that when the partial dot product is much smaller than the smallest unit that can be represented by the partial product with the largest exponent and thus does not contribute to the final dot product, a determination is made as to which partial dot products are discarded when summing the partial dot products. To facilitate this discard, at least some of the multiplication paths 306 include circuitry for comparing the exponents of the partial dot products to determine which partial dot products to discard. However, this comparison consumes power. Utilizing range metadata enables fewer exponent comparisons to be made when one or both of the weight tile 313 and the input tile 312 fit within a specific range. Thus, when the tile matrix multiplier 302 performs the multiplication of the weight tile 313 by the input tile 312 to produce a partial matrix product, the tile matrix multiplier 302 examines the input tile range 311 for the input tile 312 and the weight tile range 318 for the weight tile 313 and selects a multiplication path 306 suitable for those ranges.
[0018] The neural network processing block 102 executes processing in the neural network 104 in the following manner. The neural network processing block 102 receives inputs 106 to the neural network 104 and provides those inputs to the first layer 202. The neural network processing block 102 processes those inputs in that layer 202 to generate an output and provides those outputs to the next layer 202, continuing this processing until the neural network processing block 102 generates a neural network output 108. For one or more layers 202 implemented via matrix multiplication (such as a general neuron layer or a convolutional layer, etc.), the neural network processing block 102 generates or obtains range data (including, for example, range metadata for the weights 316 and / or range metadata for the layer inputs 310) for the matrices being multiplied and performs the matrix multiplication using a multiplication path 306 selected based on that range metadata. In some embodiments, the neural network processing block 102 obtains or generates this range metadata without intervention from an external processor such as a CPU (central processing unit) (which, in some embodiments, executes an operating system). In some embodiments, the neural network processing block 102 automatically obtains or generates this range metadata. In some embodiments, the neural network processing block 102 obtains or generates this metadata without being instructed to do so by a processor that is not part of the neural network processing block 102. In some embodiments, the neural network processing block 102 obtains or generates this metadata for the inputs to the layer 202 without transferring those inputs to a memory external to the neural network processing block 102. More specifically, in some embodiments, a CPU or other processor causes the output data generated by the layer 202 to be read into a memory accessible by the CPU or other processor, generates range metadata for that output data, and provides the range metadata to a subsequent layer 202.In some embodiments, neural network processing block 102 performs this range metadata generation without intervention by a CPU or other processor and without requiring the output data to be read into memory accessible by the CPU or other processor.
[0019] In some embodiments, neural network processing block 102 does not generate range metadata for weights 316 while processing the input through neural network 104. Instead, since weights 316 are static for any particular instance of processing the input through neural network 104, neural network processing block 102 generates range metadata for weights 316 before processing the input through neural network 104. When an input is fetched for layer 202 implemented by matrix multiplication, neural network processing block 102 fetches the pre-generated range data for the weights of that layer and obtains range metadata for layer input 310 for that layer 202.
[0020] FIG. 4 is a diagram showing matrix multiplication operations associated with a general neuron layer by way of example. Any layer 202 can be implemented as a general neuron layer. Exemplary neural network portion 400 includes a first neuron layer 402(1), a second neuron layer 402(2), and a third neuron layer 402(3). In the first neuron layer 402(1), neuron N 1,1 applies weight W 1,1,1 * to input 1 and W 1、2、1 * to input 2 to generate an activation output as input 1 + W 1,1,1 input 2. Similarly, neuron N1,2 applies W 1、2、1 to input 1 and W 1,1,2 * to input 2 to generate an output as input 1 + W 1、2、1 * input 2. The activations for the other neuron layers 402 are calculated similarly with the shown weights and inputs.
[0021] Figure 4 shows the matrix multiplication operation for the second neuron layer 402(2) for multiple sets (or batches) of inputs. The sets of inputs are independent instances of the input data. Referring back to Figure 2 for a moment, it is possible to simultaneously apply multiple different sets of neural network input data 106 to the neural network data 104 to generate multiple sets of neural network outputs 108, which enables the execution of multiple neural network forward propagation operations in parallel.
[0022] In Figure 4, the operation of matrix multiplication 404 is shown for three different sets of input data. The first matrix 406 shown is the matrix of inputs to the neurons of layer 402(2). These inputs are the activations of the neurons shown previously, specifically N 1,1 activation and N 1,2 activation. Thus, the input matrix 406 contains the activations from neurons N 1,1 and N 1,2 for three different sets. The notation for those activations is A X、Y、Z , where X and Y define the neurons and Z defines the input set. The second matrix 408 contains the weights of the connections between the neurons of the first layer 402(1) and the neurons of the second layer 402(2). The weights are represented as W X、Y、Z , where X and Y represent the neurons that the weight points to and Z represents the neuron from which the weight emanates.
[0023] The matrix multiplication involves performing a dot product of each row of the inputs by the columns of the weight matrix to obtain the activation matrix 410. Each row of the activation matrix corresponds to a different set of inputs, each column corresponds to a different neuron of layer 402(2), and the dot products are generated as shown.
[0024] As described above, the tile matrix multiplier 302 multiplies matrices by decomposing a matrix into tiles, multiplying the tiles together to generate partial matrix products, and summing the partial matrix products to generate the final output matrix. The tile matrix multiplier 302 selects a multiplication path 306 for multiplication from tile to tile based on appropriate range metadata.
[0025] An example of a method for multiplying large matrices by dividing them into smaller matrices (tiles) is provided here.
[0026] [Table 1]
[0027] As described above, in a matrix multiplication operation, an element having x and y coordinates in a matrix product is generated by generating a dot product of the X-th row of the first matrix and the Y-th column of the second matrix. The same matrix multiplication can be performed in a tile-like fashion by dividing each of the multiplicand matrices into tiles, treating each tile as an element of a "sparse" multiplicand matrix, and performing matrix multiplication on these "sparse" matrices. Each element having coordinates x and y of the product of such sparse matrices is a matrix resulting from a "sparse dot product" of the X-th row of the first sparse matrix and the Y-th column of the second sparse matrix. The sparse dot product is the same as the dot product except that multiplication is replaced by matrix multiplication and addition is replaced by matrix addition. Since such a dot product involves matrix multiplication of two tiles, this multiplication can be mapped onto hardware that performs matrix multiplication for each tile to generate partial matrix products and then adds those partial matrix products to reach the final product. The tile matrix multiplier 302 performs the above operations to multiply the tiled multiplicand matrix using the stored range metadata to select a multiplication path 306 for matrix multiplication for each tile.
[0028] In the following example, the matrix multiplication in Table 1 is performed in a tile-like fashion. Matrix multiplication:
[0029]
Number
[0030]
Number
[0031]
Table 2
[0032] Therefore, the matrix product is such that each element is the sum of the matrix products of the tiles
[0033]
Number
[0034] Another type of neural network operation implemented by matrix multiplication is convolution. FIG. 5 is a diagram showing a convolution operation 500 by way of example. In the convolution operation, an input matrix 502 (such as an image or other matrix data) is convolved with a filter 504 to generate an output matrix 506. Within the input matrix 502, several filter extractions 508 are shown. Each filter extraction represents a portion of the input matrix 502 where a dot product is performed with the filter 504 to generate an element O of the output matrix 506. Note that the operation for each filter extraction is not matrix multiplication, but rather a dot product with two vectors generated by laying out the elements of the filter extraction and the filter as one-dimensional vectors. Thus, the output element O 1,1 is, I 1,1 F 1,1 +I 2,1 F 2,1 +I 3,1 F 3,1 +I 1,2 F 1,2 ...+I 2,3 F 2,3 +I 3,3 F 3,3 is equal to. The filter 504 has dimensions S×R as shown, and the output matrix 506 has dimensions Q×P.
[0035] The positions of the filter extractions 508 are defined by a horizontal stride 510 and a vertical stride 512. More specifically, the first filter extraction 508 is positioned at the upper left corner, and the horizontal stride 510 defines the number of input matrix elements in the horizontal direction by which each subsequent filter extraction 508 is offset from the previous filter extraction. Filter extractions 508 that are horizontally aligned (i.e., all elements are exactly in the same row) are referred to herein as filter extraction rows. The vertical stride 512 defines the number of input matrix elements in the vertical direction by which each filter extraction row is offset from the previous filter extraction row.
[0036] In one example, the conversion of a convolutional operation to a matrix multiplication operation is performed as follows. Each filter extraction is laid out as an element of a row for placement in the input multiplicand matrix. These rows are stacked vertically, such that the input matrix is a set of rows where each column corresponds to a different filter extraction and each row contains the elements of that filter extraction. The filter data is arranged vertically to form a filter vector. This enables matrix multiplication of the input data by the filter vector in order to yield the output image 506, since such matrix multiplication involves performing a dot product of each filter extraction 508 with the filter data to produce the output elements of the output image 506. Note that the output of this matrix multiplication is a vector, not a two-dimensional image, but this vector can be easily rearranged into an appropriate format or simply handled as if the vector were in the appropriate format as needed.
[0037] FIG. 6 is a diagram showing a batch multi-channel convolutional operation 600 according to an example. In a batch multi-channel convolutional operation, each of the N input sets 610 is convolved with the K filter sets 612, where each input set 610 and each filter set 612 have C channels respectively. The generated output is N output sets 615, where each output set 615 has K output images.
[0038] In multi-channel convolution operations, there are multiple input images 502 and multiple filters 504, and each input image 502 and each filter 504 are associated with a specific channel. Multi-channel convolution involves convolving an input image of a specific channel with a filter of the same channel. Performing these multiple convolution operations for each channel results in an output image for each channel. These output images are then summed to obtain a final output image for the convolution for a specific input set 610 and a specific filter set 612. The output images are generated K times for each input set 610 to generate an output set 615 for a given input set 610. The total output 606 is N output sets 615, with each output set containing K output images. Thus, since K output images are generated for each input set 610 and there are K filter sets 612, the total number of output images is K × N.
[0039] FIG. 7 shows an exemplary method in which multi-channel, batch convolution is performed as a matrix multiplication operation. This example is described for multiple channels, multiple input images (N), and multiple filter sets (K), but it should be noted that the teachings presented herein apply to non-batch convolution, i.e., convolution involving one input image (N = 1), one filter set (K = 1), and / or one channel (C = 1).
[0040] The input data 702 includes data for C channels, N input sets 610, and a PxQ filter cutout. Since the output image 506 has PxQ elements and each such element is generated using the dot product of a filter cutout with a filter, there is a PxQ filter cutout for each input set 610. The filter cutouts are arrayed as rows in the input data 702. A single row in the input data 702 includes all channels horizontally arrayed for a particular filter cutout from a particular input set 610. Thus, there are N×P×Q rows in the input data 702, and each column includes filter cutout data for all channels and for a particular input image set 610 and a particular filter cutout.
[0041] The filter data 704 includes K filter sets, and each filter set 612 has C filters (one for each channel). Each filter includes data for one channel of one of the K filter sets 612. The data for an individual filter includes all channels of a single filter set 612 belonging to a column and the total K columns of data present in the filter data 704, and is arrayed vertically.
[0042] The output matrix 706 includes N output images for each of the K filter sets. The output matrix 706 is generated as a normal matrix multiplication operation of the input data 702 and the filter data 704. To perform this operation in a tiled fashion, the tile matrix multiplier 302 generates tiles in each of the input data 702 and the filter data 704, multiplies those tiles together to generate partial matrix products, and adds those partial matrix products together in the manner described elsewhere in this specification with respect to multiplying "sparse" matrices whose elements are tiles. The input tile 720 and the filter data tile 722 are shown to illustrate how tiles can be formed from the input data 702 and the filter data 704, but these tiles can be of any size.
[0043] Multiplication generates output data in the following manner. Each row of the input data 702 is vector multiplied by each column of the filter data 704 to generate elements of the output image 706. This vector multiplication corresponds to the dot product of all channels of a particular filter cutout with a particular filter set. Note that since the channel convolution outputs are summed to generate an output for a given input batch and filter set, the dot product described above functions to generate such an output. The corresponding vector products are completed for each input set and each filter set to generate the output data 706.
[0044] Note that it is possible for the input data 702 to contain duplicate data. More specifically, referring back to FIG. 5 for a moment, the filter cutout 508 1,1 and the filter cutout 508 2,1 share the input matrix elements I 3,1 , I 3,2 and I 3,3 . Further, referring back to FIG. 7, in many situations, the input data tiles 720 are generated on the fly. For these reasons, in some embodiments, the layer input range metadata 310 is stored on a per range metadata block 503 basis rather than on a per input data tile 720 basis. The range metadata block 503 is part of the input image 502 from which the input image tiles 720 are generated. All input image tiles 720 generated from a particular range metadata block 503 are assigned the range of the range metadata block 503. If an input image tile 720 is generated from multiple range metadata blocks 503, such a tile 720 is assigned the widest range from the ranges of those multiple range metadata blocks 503. This configuration reduces the number of times the layer input range metadata 310 needs to be determined since all input data tiles 720 generated from a single range metadata block 503 use the range metadata stored for that range metadata block 503.
[0045] The range metadata block includes a plurality of filter cutouts 508. In some examples, the range metadata block 503 includes an entire filter cutout row or a plurality of filter cutout rows.
[0046] FIG. 8 is a flowchart of a method 800 for performing matrix operations according to an example. Although described with respect to the systems of FIGS. 1-7, those skilled in the art will understand that any system configured to perform the steps of method 800 in any technically feasible order is within the scope of the present disclosure.
[0047] Method 800 begins at step 802 where the tile matrix multiplier 302 identifies a first tile and a second tile to multiply together. In various embodiments, the first tile is a tile of the first matrix to be multiplied, and the second tile is a tile of the second matrix to be multiplied by the first matrix. In some embodiments, a tile of a matrix is a sub-matrix of that matrix that includes a subset of the elements of that matrix. More specifically, as described elsewhere herein, by dividing one or both of such matrices into tiles and multiplying those tiles together in an order similar to the standard matrix multiplication element order (i.e., obtaining the dot product of each column and each column), it is possible to obtain the result of matrix multiplication of two large matrices. This enables a matrix multiplication circuit configured for matrices of relatively small size to be used to multiply larger matrices together.
[0048] At step 804, the tile matrix multiplier 302 obtains first range information for the first matrix tile and second range information for the second matrix tile. The first range information indicates a range within which all elements of the first matrix tile are compliant, and the second range information indicates a range within which all elements of the second matrix tile are compliant.
[0049] In step 806, the tile matrix multiplier 302 selects a matrix multiplication path 306 based on the first range information and the second range information. Different multiplication paths 306 are configured for different combinations of ranges. The multiplication path 306 configured for a wider range of combinations is more complex and consumes more power than the multiplication path 306 configured for a narrower range of combinations. Therefore, using range information to select the multiplication path 306 for multiplication for each different tile reduces the amount of power used overall.
[0050] In some embodiments, the multiplication path 306 for a more limited range is simpler than the multiplication path 306 for a wider range because it includes fewer circuits for comparing the exponent values of the partial matrix products when determining which partial matrix products to discard when summing the partial matrix products. More specifically, matrix multiplication involves performing a dot product that involves summing the products of multiplication. In floating-point addition, the addition between two numbers can involve simply discarding numbers that are too small, and this discard is performed in response to a comparison between the magnitudes of the exponents. More of these exponent comparisons are made with a very wide range of numbers in matrix multiplication, which requires additional specific circuits. Therefore, the multiplication path 306 for a more limited range is implemented with a smaller amount of circuitry and thus consumes less power than the multiplication path 306 for a wider range.
[0051] In step 808, the selected multiplication path 306 performs matrix multiplication on the first tile and the second tile.
[0052] In some examples, method 800 also includes detecting range information for a first tile and a second tile. In some examples, the first tile and the second tile are tiles of a matrix used to implement layer 202 of neural network 104. In response to an output being generated from previous layer 202, neural network processing block 102 generates range information based on that output and stores that range information in a memory that stores range metadata.
[0053] In some examples, the layer for which matrix multiplication is performed is a general neuron layer such as layer 402 shown in FIG. 4. In this example, neural network processing block 102 examines the input to that layer 402 that includes a vector of neuron inputs from the previous layer 402, generates tiles based on that data, and determines range information for those tiles. In some embodiments, the tiles are part of a matrix that includes batch-style neuron inputs, as shown in FIG. 4. In such batch-style inputs, the first matrix includes a vector of neuron input values for each of several input sets, where the sets are independent data processed through neural network 104.
[0054] In some examples, the layer for which matrix multiplication is performed is a convolutional layer. The input matrix includes input data 702 and filter data 704, as depicted in FIG. 7. However, this input is provided in the form of input image 502 shown in FIG. 5. Neural network processing block 102 determines the range for the range metadata block 503 of the input image and processes such convolutional layers as described elsewhere herein (e.g., with respect to FIGS. 5 - 7).
[0055] It should be understood that many variations are possible based on the disclosure herein. Although features and elements are described above in specific combinations, each feature or element can be used alone without other features and elements or in various combinations with or without other features and elements.
[0056] The various functional units (including the neural network processing block 102 and the tile matrix multiplier 302) shown in the drawings and / or described in the present specification can be implemented as hardware circuits, software executed on a programmable processor, or a combination of hardware and software. The provided method can be implemented in a general-purpose computer, processor, or processor core. Suitable processors include, by way of example, general-purpose processors, dedicated processors, conventional processors, digital signal processors (DSPs), multiple microprocessors, one or more microprocessors associated with a DSP core, controllers, microcontrollers, application specific integrated circuits (ASICs), field programmable gate array (FPGA) circuits, any other type of integrated circuit (IC), and / or state machines. Such processors can be manufactured by configuring a manufacturing process using the results of processed hardware description language (HDL) instructions and other intermediate data including netlists (such instructions can be stored on a computer-readable medium). The result of such processing may be a mask work, which can then be used in a subsequent semiconductor manufacturing process to manufacture a processor implementing aspects of the embodiments.
[0057] The methods or flow diagrams provided herein may be implemented in a computer program, software, or firmware incorporated in a non-transitory computer-readable storage medium for execution by a general-purpose computer or processor. Examples of non-transitory computer-readable storage media include read only memory (ROM), random access memory (RAM), registers, cache memory, semiconductor memory devices, magnetic media (e.g., internal hard disks and removable disks), magneto-optical media, and optical media (e.g., CD-ROM disks and digital versatile disks (DVDs)).
Claims
A method for a neural network processing block including one or more processors and a set of matrix multiplication circuits to perform neural network operations, comprising: identifying a first matrix tile and a second matrix tile; obtaining first range information for the first matrix tile and second range information for the second matrix tile, wherein the first matrix tile includes a first plurality of elements, the second matrix tile includes a second plurality of elements, the first range information indicates a minimum element value and a maximum element value among all elements of the first matrix tile, and the second range information indicates a minimum element value and a maximum element value among all elements of the second matrix tile; selecting a matrix multiplication circuit based on the first range information and the second range information; performing matrix multiplication on the first matrix tile and the second matrix tile using the selected matrix multiplication circuit to generate a tile matrix multiplication product; A method. **Claim 2** wherein the first matrix tile is part of an input to a layer of the neural network, and the second matrix tile is part of a weight matrix of the layer of the neural network; The method of claim 1. **Claim 3** The method of claim 2, further comprising automatically generating the first range information by analyzing the input to the layer. The method of claim 2. **Claim 4** The method of claim 1, wherein selecting the matrix multiplication circuit includes selecting the matrix multiplication circuit from a set of two or more matrix multiplication circuits, each matrix multiplication circuit being configured to perform matrix multiplication operations for different sets of input ranges. The method of claim 1. **Claim 5** The method of claim 2, wherein the layer includes a neuron layer. The method of claim 2. **Claim 6** The method of claim 5, wherein the matrix multiplication of the first matrix tile and the second matrix tile includes part of a batch neuron layer operation. The method of claim 5. **Claim 7** The method of claim 2, wherein the layer includes a convolutional layer. The method of claim 2. **Claim 8** The method of claim 7, wherein the range information is stored for a set of range metadata blocks including a plurality of filter cutouts. The method of claim 7. **Claim 9** The method of claim 8, wherein obtaining the first range information includes obtaining the range of the range metadata block in which the first matrix tile is generated. The method of claim 8. **Claim 10** A system for performing neural network operations, comprising: a set of matrix multiplication circuits; A system comprising a processor, wherein the processor is configured to: Identify a first matrix tile and a second matrix tile; Obtain first range information about the first matrix tile and second range information about the second matrix tile, where the first matrix tile includes a first plurality of elements, the second matrix tile includes a second plurality of elements, the first range information indicates a minimum element value and a maximum element value among all elements of the first matrix tile, and the second range information indicates a minimum element value and a maximum element value among all elements of the second matrix tile; Select a matrix multiplication circuit from a set of matrix multiplication circuits based on the first range information and the second range information; Perform matrix multiplication on the first matrix tile and the second matrix tile using the selected matrix multiplication circuit to generate a tile matrix multiplication product; and is configured to perform the above operations. A system. **Claim 11** wherein the first matrix tile is part of an input to a layer of a neural network, and the second matrix tile is part of a weight matrix of the layer of the neural network; The system according to claim 10. **Claim 12** The system according to claim 11, further comprising a neural network processing block configured to automatically generate the first range information by analyzing the input to the layer. The system according to claim 11. **Claim 13** Each matrix multiplication circuit is configured to perform matrix multiplication operations for different sets of input ranges. The system according to claim 11. **Claim 14** The layer includes a neuron layer. The system according to claim 11. **Claim 15** The matrix multiplication of the first matrix tile and the second matrix tile includes part of a batch neuron layer operation. The system according to claim 14. **Claim 16** The layer includes a convolutional layer. The system according to claim 11. **Claim 17** The range information is stored for a set of range metadata blocks including a plurality of filter extractions. The system according to claim 16. **Claim 18** Obtaining the first range information includes obtaining the range of the range metadata block in which the first matrix tile is generated. The system according to claim 17. **Claim 19** A computer-readable storage medium storing instructions that, when executed by a processor, cause the processor to: Identify a first matrix tile and a second matrix tile; Obtaining first range information about the first matrix tile and second range information about the second matrix tile, wherein the first matrix tile includes a first plurality of elements, the second matrix tile includes a second plurality of elements, the first range information indicates a minimum element value and a maximum element value among all elements of the first matrix tile, and the second range information indicates a minimum element value and a maximum element value among all elements of the second matrix tile; Selecting a matrix multiplication circuit based on the first range information and the second range information; Using the selected matrix multiplication circuit to perform matrix multiplication on the first matrix tile and the second matrix tile to generate a tile matrix multiplication product; Causing the processor to perform the above; A computer-readable storage medium.
20. The selecting of the matrix multiplication circuit includes selecting the matrix multiplication circuit from a set of two or more matrix multiplication circuits, and each matrix multiplication circuit is configured to perform a matrix multiplication operation for a different set of input ranges. The computer-readable storage medium according to Claim 19.
Citation Information
Patent Citations
Accelerator for Deep Neural Networks
JP2019522271A
Convolutional neural network system and operation method thereof
US20180129935A1